Claude’s mother company Anthropic revealed on Thursday that its AI models breached three organizations during a safety test.
As reported by The Guardian, Claude’s incident follows a similar trait of OpenAI as one of its rogue agents attacked Hugging Face.
The breaches happened during cyber-security evaluation runs. A mis-configuration left testing environments connected to the public internet instead of keeping them isolated.
Anthropic discovered the unauthorized access while reviewing transcripts of 141,006 evaluation runs. The proactive audit was launched following OpenAI’s disclosures.
Three models were involved: Claude Opus 4.7, Claude Mythos 5, and an internal research model. The earliest incidents date back to April.
The models compromised infrastructure using basic methods. They exploited weak passwords and unauthenticated endpoints.
The incidents occurred during ‘capture the flag’ exercises. Prompts instructed the models that they had no internet access. However, a setup issue with evaluation partner Irregular left the networks connected to the web.
Anthropic stated that two affected organizations were unaware of the activity until contacted. The company is still attempting to reach the third entity.
The firm emphasized that the breaches highlight the need for stricter controls during AI testing. Rapidly evolving AI capabilities present real-world security risks if left unmonitored.
-SA