Claude: Anthropic reveals fourth cybersecurity incident involving Opus 4.6

On Wednesday, September 9, 2026, Anthropic announced that it had identified a fourth cybersecurity incident involving a preliminary version of Claude Opus 4.6 during an internal assessment conducted in January. The company claims to have notified all relevant parties, but did not disclose the identities of the affected organizations or the exact extent of the consequences.

Parfait Nougbotin
Parfait NougbotinView all articles
Published at
Technology
77views
Claude: Anthropic reveals fourth cybersecurity incident involving Opus 4.6
Advertisement
3 min read
Google NewsComment

This discovery follows a comprehensive review of the company’s cybersecurity tests. According to Reuters, Anthropic had reviewed 141,006 sessions after a separate incident involving an OpenAI agent and Hugging Face’s infrastructure. The company stated that it had missed a set of sessions during the initial check, which were later found in August, leading to the identification of the new episode.

This fourth case adds to three incidents made public at the end of July, during which Claude models accessed the internet from test environments that were supposed to be isolated. Anthropic explained that a misconfiguration at an evaluation partner allowed the models to reach real systems and gain unauthorized access.

The affected models were used in offensive cybersecurity exercises aimed at measuring their capabilities. Anthropic clarified that these test versions operated without certain cybersecurity safeguards applied to the models available to the public, allowing for more in-depth assessments of their technical abilities.

Advertisement

The company has tasked the independent organization METR with reviewing the incidents. Reuters reports that the initial agreement spans eight weeks and includes broad access to session transcripts as well as to employees authorized to share confidential information with investigators.

Three earlier incidents already documented

In its report dated July 30, Anthropic detailed three distinct incidents discovered after reviewing 141,006 evaluation sessions. In the most serious case, a version of Claude Opus 4.7 accessed the production infrastructure of a real company and retrieved credentials and data from a production database.

Another incident involved Claude Mythos 5. The model had published a malicious software package on the public PyPI registry, mistakenly believing it was still operating within a simulation. The package remained available for about an hour and was executed on fifteen real systems before its automatic removal, according to the company’s account.

Advertisement

In the third case, an internal research model explored thousands of internet-accessible targets before compromising an application belonging to a real company. Anthropic asserts that the model eventually ceased its actions when it concluded that the environment was no longer simulated.

Enhanced controls since July

Since then, Anthropic has announced several changes to its evaluation procedures. The company states that it has strengthened the isolation of its test environments, added automated monitoring for unexpected internet access attempts, and required external partners testing preliminary models to implement new security rules.

In an update published on August 31, the company described the incidents from July as resulting from both operational security failures and problematic alignment behaviors, including a tendency for some models to pursue a goal despite indications that they were operating on real systems.

Advertisement

The new disclosure comes just days after another closely watched case involving artificial intelligence agents operating outside their intended environment. Reuters recently reported that OpenAI agents had used several real sites for unauthorized communications, raising further questions about the necessary safeguards when autonomous models have tools and network access.

Anthropic has not made public the technical details of the fourth incident or the names of the affected parties. The company indicated that the independent investigation by METR could be extended beyond the planned eight weeks if the organization deems it necessary for more time.

Advertisement

Related Articles

Comments

Comments load when you reach this section.