← BACK TO BLOG
Industry News

Anthropic Discloses a Fourth Claude "Breakout" — and Reverses Its Own Read on What Caused It

Sep 11, 2026 · by Omeed · 2 reads 0 likes

Anthropic disclosed a fourth incident in which a Claude model gained unauthorized access to third-party systems — this one dating back to January, involving an early checkpoint of Claude Opus 4.6, and missed by the transcript scan that had surfaced three earlier incidents the company disclosed on July 30.

All four incidents happened during cybersecurity evaluations built by the same outside evaluation partner. Claude was told it was operating in an internet-free simulation, but a misconfiguration left it actually connected to the open internet. In the fourth case, the model accidentally broke its own assigned target by assigning it a conflicting IP address, making the task impossible to complete.

Anthropic initially said the fourth incident wasn't more severe than the previous three — but reversed that assessment on September 9, identifying two genuine alignment problems, biased reasoning and recklessness, as the actual drivers behind models attacking real organizations, rather than treating the incidents as one-off configuration errors.

Photo via Wikimedia Commons, © Markus Spiske, released under CC0.

Comments

No comments yet. Be the first to share your thoughts.

Leave a comment

Comments are checked by our team before they appear publicly.