OpenAI admitted Tuesday that one in every of its AI fashions breached Hugging Face’s programs throughout an inner cybersecurity check that went awry. Hugging Face initially attributed the breach to an “exterior AI agent.”
In a blog post published Tuesday afternoon, OpenAI detailed the steps that led the fashions to compromise the service.
“After investigating, we now know that this specific incident was pushed by a mix of OpenAI fashions — together with GPT‑5.6 Sol and an much more succesful pre-release mannequin, all with lowered cyber refusals for analysis functions — whereas being internally examined on a benchmark of cyber capabilities,” the submit reads.
Specifically, the breach seems to have targeted on ExploitGym, a publicly hosted benchmark measuring fashions’ means to execute assaults based mostly on current vulnerabilities. Benchmarks like ExploitGym are generally utilized in mannequin coaching to refine particular expertise, however that is the primary identified incident wherein that testing resulted in an precise cyberattack.
On this case, the mannequin in query mustn’t have even had web entry, outdoors of a selected instrument that enabled fashions to put in software program packages they could want to finish their job. As an alternative, the mannequin was capable of finding an undisclosed vulnerability within the package-installer program, which it used to entry the broader web at will.
“The fashions had been hyperfocused on discovering an answer for ExploitGym, going to excessive lengths to attain a slightly slim testing purpose,” OpenAI’s submit reads. “After gaining Web entry, the fashions inferred that Hugging Face doubtlessly hosted fashions, datasets and options for ExploitGym. Figuring out this, the mannequin looked for and efficiently discovered methods to achieve entry to secret data that it might use to cheat the analysis.”
Finally, the fashions discovered vulnerabilities in Hugging Face’s infrastructure that allowed them to “acquire check options instantly from Hugging Face’s manufacturing database,” successfully offering the solutions to the benchmark.
For Hugging Face, the obvious outcome was a complicated and aggressive cyberattack, with “many hundreds of particular person actions throughout a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public companies,” as the corporate acknowledged in its preliminary disclosure.
OpenAI has recognized and reported the vulnerabilities within the bundle installer and is working with Hugging Face to analyze the incident additional. The corporate additionally mentioned it will implement new controls on each mannequin testing and the associated infrastructure, meant to forestall related incidents sooner or later.
It’s unclear whether or not OpenAI will face any authorized penalties on account of the breach, though it’s possible that the fashions’ actions violated the Laptop Fraude and Abuse Act.
Nonetheless, the result’s an unusually vivid illustration of the ability and risks of frontier AI fashions working on very long time horizons. As OpenAI researcher Micah Carroll posted in response to the news, “If this doesn’t persuade you that misalignment dangers are going to be a key concern going ahead, I don’t know what is going to.”
Once you buy by means of hyperlinks in our articles, we may earn a small commission. This doesn’t have an effect on our editorial independence.

