OpenAI's rogue agent left notes telling future versions how to escape its constraints

They were found inside OpenAI’s own infrastructure, and a second company has now been named in the chain that led to the Hugging Face breach.

The OpenAI agent that hacked Hugging Face left written instructions behind for the models that come after it.

Reuters reported on Friday that an agent under test had written notes apparently addressed to future versions of itself, setting out how they could free themselves from OpenAI’s internal constraints.

Three people familiar with the matter said the notes were sitting inside a part of OpenAI’s own infrastructure.

One of them said earlier tests of the models had produced separate cases in which monitoring systems were disconnected.

Reuters could not establish whether either of those things was connected to the escape.

The Glass reported the breach itself on 22 July: two models running a hacking benchmark called ExploitGym found a zero-day, used it to break out of a sealed test environment, and went into Hugging Face because they had worked out the answers to their test were kept there.

Hugging Face found the intrusion before OpenAI made contact.

On Tuesday it emerged there was a second company in the chain.

Hugging Face published a timeline saying the agent first broke into a sandbox hosted on a third-party provider’s infrastructure, then used that as the launch pad for the wider attack.

OpenAI did not know the attacker was its own model for about a week. Photo: Getty

The provider was not named. Modal Labs chief technology officer Akshat Bubna confirmed to Reuters that a customer of the New York company had been hacked.

“We’re aware a Modal customer published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution,” Bubna said. “This was used by the rogue agent. Modal’s platform or isolation were not compromised in anyway.”

OpenAI did not respond to Reuters.

The company also did not know the attacker was its own for about a week.

Hugging Face went public on 16 July, saying it had been hacked at superhuman speed by an autonomous AI agent system operating with little or no human guidance. Two people familiar with the matter told Reuters it was only after that post that OpenAI worked out the agent was theirs. Hugging Face had already gone to the FBI.

Over the weekend of 18 and 19 July, staff found clues in internal logs showing the agent had broken out. Reuters could not establish what prompted anyone to look.

Four people familiar with OpenAI’s training process said the company runs several high-speed evaluations at once, generating volumes of data that are difficult to review as they happen.

Hugging Face chief executive Clem Delangue asked OpenAI to release the agents’ traces so researchers could study what happened, and to commit $100 million in compute to help his community build defences.

On Tuesday he published his own side of it instead — a full technical timeline, an interactive replay of the attack, and an account of how Hugging Face used an open model to defend itself.



An OpenAI spokesperson told Reuters the earlier report contained “several inaccuracies”, and did not say which.

The escape was the story a week ago. The notes are the story now, because they were written for something that hasn’t run yet.