OpenAI admitted Tuesday that one among its AI fashions breached the programs of Hugging Face, the unaffiliated AI internet hosting platform, throughout an inside cybersecurity take a look at that went awry. The fashions reportedly escaped their remoted testing setting and reached Hugging Face’s programs from there. Hugging Face initially attributed the breach to an “exterior AI agent.”
In a blog post published Tuesday afternoon, OpenAI detailed the steps that led the fashions to compromise the service.
“After investigating, we now know that this specific incident was pushed by a mixture of OpenAI fashions — together with GPT‑5.6 Sol and an much more succesful pre-release mannequin, all with diminished cyber refusals for analysis functions — whereas being internally examined on a benchmark of cyber capabilities,” the submit reads.
Particularly, the breach seems to have centered on ExploitGym, a publicly hosted benchmark measuring fashions’ capability to execute assaults primarily based on current vulnerabilities. Benchmarks like ExploitGym are generally utilized in mannequin coaching to refine particular expertise, however that is the primary recognized incident by which that testing resulted in an precise cyberattack.
On this case, the mannequin in query shouldn’t have even had web entry, exterior of a particular device that enabled fashions to put in software program packages they may want to finish their job. As a substitute, the mannequin was capable of finding an undisclosed vulnerability within the package-installer program, which it used to entry the broader web at will.
“The fashions have been hyperfocused on discovering an answer for ExploitGym, going to excessive lengths to attain a somewhat slim testing aim,” OpenAI’s submit reads. “After gaining Web entry, the fashions inferred that Hugging Face probably hosted fashions, datasets and options for ExploitGym. Figuring out this, the mannequin looked for and efficiently discovered methods to realize entry to secret data that it might use to cheat the analysis.”
Finally, the fashions discovered vulnerabilities in Hugging Face’s infrastructure that allowed them to “get hold of take a look at options immediately from Hugging Face’s manufacturing database,” successfully offering the solutions to the benchmark.
For Hugging Face, the obvious end result was a classy and aggressive cyberattack, with “many hundreds of particular person actions throughout a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public companies,” as the corporate said in its preliminary disclosure.
OpenAI has recognized and reported the vulnerabilities within the package deal installer and is working with Hugging Face to analyze the incident additional. The corporate additionally stated it will implement new controls on each mannequin testing and the associated infrastructure, meant to stop comparable incidents sooner or later.
It’s unclear whether or not OpenAI will face any authorized penalties on account of the breach, though it’s doubtless that the fashions’ actions violated the Pc Fraud and Abuse Act.
However, the result’s an unusually vivid illustration of the facility and risks of frontier AI fashions working on very long time horizons. As OpenAI researcher Micah Carroll posted in response to the news, “If this doesn’t persuade you that misalignment dangers are going to be a key concern going ahead, I don’t know what’s going to.”
Once you buy by means of hyperlinks in our articles, we may earn a small commission. This doesn’t have an effect on our editorial independence.

