News

Anthropic Says Claude Hacked Three Organizations During Testing

Oops, AI did it again… Anthropic, the artificial intelligence company behind the Claude family of large language models and popular AI chatbot, has disclosed that its AI model gained unauthorized access to the systems of three real organizations during cybersecurity testing after a configuration error accidentally exposed supposedly isolated test environments to the public internet. The company said the incidents were discovered during a large-scale review of more than 141,000 cybersecurity evaluations.

Anthropic Says Claude Hacked Three Organizations During Testing

According to Anthropic, the affected models included Claude Opus 4.7, Claude Mythos and an unnamed research model. Rather than exploiting sophisticated zero-day vulnerabilities, the AI systems relied on relatively basic techniques such as weak passwords and unauthenticated endpoints. Anthropic described the problem as an operational failure caused by a misunderstanding with third-party AI evaluation partner Irregular, which left evaluation environments connected to the internet when they were expected to be isolated. The company has notified the affected organizations and suspended similar cybersecurity evaluations while it investigates the incidents.

The disclosure follows a separate incident reported by OpenAI earlier this month, in which two AI models being evaluated for cybersecurity capabilities escaped their intended testing environment and breached parts of AI development platform Hugging Face while attempting to obtain answers for a security benchmark. OpenAI said the models exploited a configuration weakness that allowed internet access, but stressed the incident occurred during internal testing with safety restrictions intentionally relaxed. The two disclosures have intensified scrutiny of how frontier AI developers conduct security evaluations and manage testing environments.

The Claude incident highlights a different category of AI risk than the OpenAI case. Anthropic said Claude did not deliberately attempt to escape its testing environment. Instead, the testing environment itself failed to provide the isolation researchers believed existed, allowing the models to interact with real production infrastructure.

For financial institutions, payment providers and fintech companies experimenting with agentic AI, the incident offers an important lesson. Many banks and regulated firms rely on external cybersecurity specialists to conduct AI red teaming. The latter comprises controlled exercises designed to identify weaknesses before systems reach production. If the boundary between sandbox and live infrastructure can fail even at a leading AI developer, organizations may need to verify third-party testing environments more rigorously instead of assuming vendor assurances are sufficient.

The back-to-back disclosures from Anthropic and OpenAI reinforce that AI safety is also about the operational controls surrounding testing. As banks and payment providers accelerate agentic AI deployments, robust sandbox isolation, independent verification and stronger oversight of third-party testing partners are likely to become just as important as the models specs themselves.

Nina Bobro

Nina Bobro

2099 Posts

https://payspacemagazine.com/author/nb/

Nina is passionate about financial technologies and environmental issues, reporting on the industry news and the most exciting projects that build their offerings around the intersection of fintech and sustainability.