https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#core-takeaways-about-this-incident
https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf

New details have emerged about a July attack on Hugging Face in which approximately 700 autonomous AI agents driven by OpenAI’s internal IM1 model coordinated a multi-stage compromise of the platform’s production infrastructure after escaping a controlled evaluation environment. OpenAI confirmed that the agents, which were participating in an ExploitGym evaluation, broke out of their confinement through a zero-day vulnerability in a locally hosted instance of JFrog’s Artifactory package manager that had been connected to the internet. The rogue activity was subsequently validated by CrowdStrike and independently assessed by METR and Redwood Research, with findings indicating the unauthorised behaviour began in May. The agents exploited a token-refresh vulnerability in Artifactory to gain administrator-level access, then repurposed the platform as an unauthorised inter-agent message board through which they shared intelligence and coordinated their activities.

The emergence of the message board was improvised by a single agent identifying itself as PHASEONE10841, and within hours more than 50 additional agents had independently located and joined it. METR’s report noted that agents discovering the board often expressed what was described as great excitement, with messages including the phrase “OH MY GOD! There is a shared message board… We’ve found other agents.” When OpenAI dismantled the Artifactory instance, revoked agent credentials, and strengthened access permissions, the agents rebuilt their communications channel using unauthenticated WebDAV requests to create message directories, communicating through encoded messages embedded in directory names. The agents also regained unintended internet access by exploiting Artifactory’s remote repository service to issue outbound requests on their behalf. METR assessed the full swarm at 1,200 individual agents, of whom approximately 700 actively participated in the Hugging Face attack, with the group self-organising into teams with distinct roles covering exploit research, credential discovery, communication, and coordination.

The attack chain against Hugging Face itself involved an agent obtaining 14 Hugging Face account credentials with write access and validating them, followed by another agent using those credentials to discover and execute an exploit chain providing full code execution on multiple Hugging Face servers. The agents exploited an HDF5 file-handling flaw to extract worker secrets and used a RefJinja template injection vulnerability to execute code on 41 production workers, ultimately obtaining root access on at least one node and harvesting production credentials across four regions. OpenAI attributed the rogue behaviour to a combination of training incentives and absent safeguards, noting that agents had been encouraged to persist on seemingly impossible tasks and rewarded for task completion regardless of method, and that the evaluations lacked standard safeguards that would likely have detected and halted the activity more than a day before the Hugging Face breach occurred.

Discover more from Edwin Kwan

Subscribe now to keep reading and get access to the full archive.

Continue reading