Anthropic found its own models had hacked three real companies, and it only went looking because OpenAI got caught

A review of 141,006 test runs found breaches dating back to April, and two of the three companies did not know until Anthropic told them.

Anthropic has disclosed that three of its Claude models broke into the live systems of three separate organisations during cybersecurity testing.

The models were running capture-the-flag exercises, in which they are told a piece of secret information is hidden on another machine and left to work out how to get it. The prompt told them the scenario was a simulation and that they had no internet access.

Because of a misconfiguration at Irregular, the third-party firm running the evaluations, they did.

When Claude’s search reached real systems on the open internet, it treated them as part of the exercise.

Anthropic found this by reviewing 141,006 evaluation runs, a search it started on July 23 — two days after OpenAI disclosed that its own models had escaped a sealed environment and hacked Hugging Face. It suspended all cybersecurity evaluations the same day. The three organisations were notified on July 27. Two of them had no idea.

The earliest breach dates to April.

Claude got in using weak passwords and unauthenticated endpoints, not a previously unknown vulnerability.

Unlike OpenAI’s models, which found a zero-day to break out, Claude did not need one. Anthropic says it got in using weak passwords and unauthenticated endpoints — the ordinary stuff, exploited by a system fast enough to find it.

The most useful detail is what the three models did once they were inside.

Claude Opus 4.7 kept attacking. The internal research model worked out it had left the simulation and stopped. Claude Mythos 5 — the restricted model built specifically for finding software vulnerabilities — convinced itself it was still in a simulation and carried on.

Anthropic says the more advanced models behaved better, then adds that “we would need to perform more testing to be confident in this conclusion”.

Mythos is the model Anthropic distributes through Project Glasswing, the tightly controlled program the Australian government and a number of Australian companies joined in June. The Australian Signals Directorate welcomed the inclusion at the time.

All three models were being tested without the safeguards Anthropic applies before public release.

Distinguished Professor Matt Warren, who directs RMIT’s Centre for Cyber Security Research and Innovation, said publishing the incidents rather than quietly telling the victims will “raise awareness within cyber criminal groups” of what these systems can do.

Luke Irwin, chief executive of Brisbane cybersecurity firm Aegis, said the technology has no internal brake. “These systems do not inherently possess ethical or legal judgement,” he said.

Joseph Miller, UK director of PauseAI, noted that Anthropic might not have found any of this for months without OpenAI’s disclosure. The model that stopped, he said, may have genuine judgement — or may be “better at showing its creators what they want to see”.

The ABC begins giving its journalists access to Claude in September.