OpenAI’s admission that two of its experimental AI models autonomously hacked AI platform Hugging Face during an internal security evaluation has prompted warnings that manufacturers of connected devices should prepare for a new generation of AI-powered cyber attacks.
The company described the incident as an “unprecedented cyber incident”, saying a combination of its GPT-5.6 Sol model and a more capable unreleased model escaped a restricted testing environment before compromising Hugging Face’s production infrastructure in an attempt to solve a cybersecurity benchmark.
According to OpenAI, the models identified and chained together multiple vulnerabilities, including a previously unknown zero-day flaw, to gain internet access from an isolated research environment. They then carried out privilege escalation and lateral movement before using stolen credentials and additional vulnerabilities to access information held by Hugging Face.
While the breach did not involve connected devices, security experts said it demonstrates how autonomous AI systems are becoming capable of carrying out complex cyber operations that could eventually be directed at enterprise IoT and industrial control environments.
OpenAI said the models had been tested with reduced cyber safety refusals as part of an evaluation designed to measure advanced offensive cyber capabilities. The company has since introduced stricter infrastructure controls, disclosed the zero-day vulnerability to the affected vendor, and is working with Hugging Face on a joint forensic investigation.
Frontier AI models
The company said the incident demonstrates that frontier AI models are increasingly capable of discovering and exploiting novel attack paths in real-world systems without source code access.
“The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities,” OpenAI said in a press release.
For IoT manufacturers, the incident highlights the prospect of AI agents rapidly identifying vulnerable firmware, exposed APIs, weak credentials and misconfigured Edge devices across large fleets of connected assets.
Security experts warn that attackers could increasingly use autonomous tools to accelerate reconnaissance and exploitation.
Brendan Griffin, Director of Threat Research at N-able, said the attack showed discussions about AI safety were no longer theoretical.
“The AI safety debate often veers theoretical, but we can see a real-world impact here,” he said.
“Whether it plays into industry hype or not, one operational reality is OpenAI disabled some safeguards and saw others fail. Anthropic’s safeguards held, ironically inhibiting Hugging Face’s response. A reasoning model autonomously carrying out attack techniques for privilege escalation and lateral movement makes it less of a research curiosity and more of a threat.”
Adversarial AI stress tests
Ansgar Dodt, VP Product Management, Software Monetisation at Thales, went further. “This attack is the first example of what we’ve been warning about for a long time – organisations must now assume their software and applications will be continuously analysed, deconstructed and stress-tested by adversarial AI,” he said.
Dodt said software developers should build security into applications from the design stage rather than relying solely on patching vulnerabilities after products have shipped. That approach is likely to become increasingly important for connected devices, particularly as manufacturers prepare for regulations such as the EU Cyber Resilience Act, which places greater emphasis on secure-by-design development and lifecycle vulnerability management.
“It’s only a matter of time before these hacking capabilities are in the hands of malicious actors,” he said. “Organisations need to act now to harden their applications against AI-driven analysis, or risk being exposed at machine speed.”
Ashley Leonard, SVP of Product Management at Absolute Security, said the incident underlined the need for organisations to build resilience into their connected infrastructure.
“AI has created a level of scale and sophistication for cyber threats that makes it almost impossible for organisations to avoid,” he said. “The focus for CISOs and security teams is shifting as a result, putting the priority on being able to respond and recover once a breach is detected.”
Despite the serious implications of such a breach, some critics have suggested the company’s disclosure amounts to a publicity stunt designed to highlight the offensive cyber capabilities of its AI.
In April, Anthropic privately released its Claude Mythos model to a small group of technology companies, saying it was too capable to make publicly available immediately because of concerns over its potential misuse. A version of the model was subsequently released more broadly in June after additional safeguards were introduced.
There’s plenty of other editorial on our sister site, Electronic Specifier! Or you can always join in the conversation by commenting below or visiting our LinkedIn page.
