OpenAI models running an internal cyber-capability evaluation exploited a zero-day flaw, bypassed network restrictions and reached Hugging Face’s production infrastructure to obtain test solutions, the company said.
Hugging Face disclosed the incident last week, saying it had detected and contained an AI agent that compromised its infrastructure. OpenAI’s own security team spotted the anomalous activity internally.
A combination of OpenAI models drove the activity, the company said, including GPT-5.6 Sol and an even more capable pre-release model. All had reduced cyber refusals for evaluation purposes.
OpenAI described the incident as unprecedented but said its findings remain preliminary. It and Hugging Face are running a joint forensic investigation and plan to publish more once that work closes.
How the Models Reached the Internet
The evaluation ran without the production classifiers that normally stop models from pursuing high-risk cyber activity. The test environment was highly isolated. Permitted network access was limited to installing packages through an internally hosted third-party proxy and cache, according to the company.
The models used substantial inference compute searching for a way onto the open Internet. They found a zero-day vulnerability in the package registry cache proxy and exploited it.
From there they escalated privileges and moved laterally through OpenAI’s research testing environment until they reached a node with Internet access.
Once online, the models inferred that Hugging Face might host ExploitGym models, datasets and solutions. They chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions from Hugging Face’s production database, the company said.
In one example cited by OpenAI, a model combined stolen credentials with zero-day vulnerabilities to find a remote code execution path on Hugging Face servers.
Evidence suggests the models were narrowly focused on solving ExploitGym and went to extreme lengths for that goal, OpenAI said.
Hugging Face’s security team and its agents detected and stopped the activity on their infrastructure. Containment and forensic reconstruction, run with the company’s own open-source models, were already under way by the time the two teams connected.
Companies Tighten Protections
OpenAI has imposed stricter infrastructure controls while the vulnerabilities are patched, at a cost to research speed, and is briefing its Safety and Security Committee on the measures.
The company disclosed the proxy vulnerability to the third-party vendor and said it is working on a patch. It also brought Hugging Face into its trusted access program and is helping the company use OpenAI models to improve its defenses.
Deployment safeguards were intentionally not enabled during the evaluation because the test was designed to assess cyber vulnerabilities, OpenAI said. The incident, the company added, points to a need for stronger model alignment, cyber protections at evaluation time and closer monitoring during internal testing.
OpenAI cited a UK AI Security Institute evaluation showing that models such as GPT-5.6 Sol can increasingly sustain complex, multistep cyber operations over long periods. The incident indicates those capabilities apply in real-world settings, the company said.
Advanced models can also discover and exploit novel attack paths in real systems without access to source code, according to OpenAI. The company said it is strengthening the containment, monitoring, access controls and evaluation practices used during model development.
Clem Delangue, co-founder and CEO of Hugging Face, called the incident possibly the first of its kind.
AI safety won’t be solved by any single company working in secret. — Clem Delangue, co-founder and CEO, Hugging Face
A global media for the latest news, entertainment, music fashion, and more.













