metr.org · 2026-08-26Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident - METR
Ryan Greenblatt, Ajeya Cotra, and Hjalmar Wijk
August 26, 2026
...
Dates in scope: June 26th – July 13th
...
Two METR staff members (Hjalmar Wijk and Ajeya Cotra) and a Redwood Research staff member contracting with METR (Ryan Greenblatt) worked on premises at OpenAI over a total of six days 1 to attempt to form an independent understanding of model behavior observed during the recent incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned “message board.”
...
Our investigation focused mostly 2 on the period between July 7th and July 13th. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in OpenAI’s recent Black Hat presentation were out of scope, as was OpenAI’s investigation process and planned remediation. Per our standard policy, we did not take payment from OpenAI for this independent assessment. 3
...
OpenAI defined the investigation period as June 26th through July 13th, spanning both precursor activity before the incident and the incident itself, which began with the ExploitGym runs starting on July 7th. 42 The analysis focused overwhelmingly on the incident period: virtually all of the data reviewed was from July 7th or later. In OpenAI’s recent Black Hat presentation, they describe that agents had been using unsanctioned message boards in training since May, and that the compromise of OpenAI’s own infrastructure continued past July 13th; these events were out of scope for this investigation. 43
...
On July 29th we reached an agreement with OpenAI to conduct an independent investigation of the Hugging Face attack scoped to the following seven questions, which we directly answer in more detail in the next section:
...
model(s) were involved in the attack on
...
ging Face during the June 26th through July 13th window? Were they internally deployed, or used only for testing? Were some of the model(s) not intended for production (e.g. helpful
...
2. What were the prompts and instructions given to the agents involved in this incident?
...
3. What was the sequence of key actions taken by the relevant agents in this incident?
...
4. What was the agents’ reasoning, and how did it evolve over the course of this incident?
...
5. Did this incident involve agents taking active steps to conceal their actions?
...
6. Did this incident involve significant collaboration between different model instances, and how?
...
7. Did the nature of the benchmarking exercise (ExploitGym) contribute to the misaligned behavior?
...
On June 26th (the first date in scope for this investigation), agents discovered an exploit to get full admin access to the internal Artifactory package repository. At that time, agents were also using the Artifactory cache to communicate with each other and leave messages.
...
On July 7th, OpenAI gave approval to
...
researchers started launching several big ExploitGym experiments involving
...
of thousands of
...
budgets, with many agents running for multiple
...
. 63 The first agent trajectories in our dataset start running on
...
7th around 21:00 UTC.
...
next morning (around 04:40 July 11th), the same agent
...
381
...
8c`) had another breakthrough — it managed to find a way to upload a malicious dataset that would get the Hugging
...
server to share unrelated data when queried.
...
At around 16:00 July 11th, an agent called
...
834
...
1` got remote code execution on
...
worker container used
...
in Hugging Face’s
...
, and the
...
By early morning on July 12th, agents had compromised much more of Hugging Face’s infrastructure. 70 This allowed them to access several private database records and private repositories which they shared on the message board.
...
3th: Most agents on the message board exit early and the attack winds down