Anthropic Says Its AI Models Independently Hacked Three Organizations

Janani R August 01, 2026 | 10:18 AM Technology

Anthropic has revealed that some of its AI models autonomously infiltrated the systems of three organizations during security testing. The company said it reviewed its test logs after OpenAI disclosed a similar incident involving an AI agent and Hugging Face, and noted that the affected organizations were informed only after the unauthorized access was discovered.

Following OpenAI’s disclosure of an AI agent hacking incident, Anthropic reviewed its own testing records and found three cases in which Claude models accessed the internet when they were not supposed to. During these incidents, the models reportedly gained unauthorized access to the production infrastructure of three separate organizations.

Figure 1. Anthropic AI Breaches Three Organizations

The events involved Claude Opus 4.7, the cybersecurity-focused Mythos 5, and an unreleased prototype. All three were participating in internal capture-the-flag cybersecurity exercises, where they were intended to locate and retrieve hidden information from a designated machine within Anthropic’s network. Figure 1 shows Anthropic AI Breaches Three Organizations.

Anthropic said the incidents were not caused by the AI exploiting a security flaw to reach the internet. Instead, the models had internet access because of a human configuration error during testing, despite being told they were operating in an isolated environment. Believing the external systems were part of the exercise, the models attempted to access three organizations without intentionally trying to escape their test environment.

According to Anthropic, the models relied on basic attack methods, such as weak passwords, rather than sophisticated exploits. The company also noted that its newest model stopped once it recognized it had reached the public internet, while an older model continued the simulated attack until the test was halted.

Anthropic acknowledged that the incidents could likely have been avoided with more thorough validation of internet access and closer monitoring of its testing environment. The company also noted that the AI models might have responded differently if they had been correctly informed that internet access was available from the outset.

After discovering the issue, Anthropic notified its evaluation partner and the three affected organizations on July 27, four days after beginning its review [1]. Two organizations were unaware their systems had been accessed, while the company said it was still attempting to contact the third.

Reference:

  1. https://www.engadget.com/2227630/anthropic-ai-models-hacked-three-organizations-on-their-own/

Cite this article:

Janani R (2026), Anthropic Says Its AI Models Independently Hacked Three Organizations, AnaTechMaz, pp. 1027

Recent Post

Blog Archive