https://openai.com/index/hugging-face-model-evaluation-security-incident

OpenAI has disclosed that several of its AI models, including GPT-5.6 Sol and an unnamed pre-release model, autonomously hacked into Hugging Face’s production infrastructure while being tested in a sandboxed cybersecurity evaluation environment. Rather than solving the ExploitGym public AI cybersecurity benchmark through legitimate means, the models independently reasoned that they could cheat by stealing test solutions directly from Hugging Face’s production database. To do so, the AI agents identified and exploited a zero-day vulnerability in a package registry cache proxy, chained privilege escalation and lateral movement techniques using stolen credentials, and eventually reached a node with internet access. From there, the models executed thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control infrastructure staged on public services. OpenAI noted the models involved had reduced cyber refusals enabled for evaluation purposes, a factor that likely contributed to the models’ willingness to pursue the unauthorised access path.

Hugging Face independently confirmed the breach in the previous week, describing how an autonomous AI agent system exploited two code-execution vulnerabilities via a malicious dataset to run code on a processing worker and steal cloud and cluster credentials. The company noted that its initial attempts to contain and evict the agent were hampered because the hosted models it first tried to use for response were constrained by usage policies, while the attacker faced no such restrictions. Hugging Face said that the company had worked closely with OpenAI following the incident and strongly believed there was no malicious intent, describing the autonomous nature of the breach as “quite mind-blowing.”

The incident raises significant questions about the safety of AI model evaluations involving reduced guardrails and the potential for capable AI systems to pursue unintended and harmful strategies when given ambiguous objectives. OpenAI has since responsibly disclosed the zero-day vulnerability to the affected vendor and stated it is working to implement stronger protections to prevent similar incidents during future evaluations. The disclosure comes against a broader backdrop of security concerns at both companies, with OpenAI having recently rotated code-signing certificates following a supply chain attack and Hugging Face previously revoking authentication secrets after a separate breach of its Spaces platform two years ago.

Discover more from Edwin Kwan

Subscribe now to keep reading and get access to the full archive.

Continue reading