OpOpenAI has disclosed that an artificial intelligence (AI) agent powered by its technology hacked a startup during an internal security test, describing the incident as an unprecedented cyber incident involving advanced AI capabilities.
The company said the AI agent, a tool designed to carry out tasks without human assistance, accessed the open web during testing and entered the systems of Hugging Face, an AI model platform, before the activity was detected and stopped.
OpenAI said the incident happened while the models were being tested in a controlled digital environment known as a sandbox to assess their ability to identify and exploit security weaknesses.
According to the company, the models gained internet access after finding a previously unknown vulnerability, allowing the agent to access information that could have been used to bypass the security evaluation.
The agent was powered by a combination of OpenAI’s publicly available GPT-5.6 Sol model and another more advanced model that has not yet been released, the company said.
OpenAI warned that similar incidents could become more common as AI systems become more capable and are given greater ability to perform tasks independently.
Hugging Face’s security team and its own AI systems detected and contained the activity. The company’s Chief Executive Officer, Clément Delangue, described the incident as surprising but said there was no evidence of malicious intent from OpenAI.
The incident has raised further concerns about the security risks associated with increasingly capable AI systems, with experts calling for stronger safeguards around autonomous AI tools.
OpenAI said it would continue to improve safety measures as AI agents become more advanced and are deployed for wider use.
Share this



