News

Anthropic Withholds Latest AI Model From UK Safety Testing, Concerns Over AI Risks Growing

Anthropic withheld its latest, most advanced artificial intelligence (AI) model Claude Mythos 5.1 from pre-release testing by independent UK body that usually receives early access to frontier models. 

Anthropic Withholds Latest AI Model From UK Safety Testing, Concerns Over AI Risks Growing

While the American AI developer opened access to Claude Mythos 5.1 for vetted U.S. organizations on September 1, the UK’s AI Security Institute (AISI) was not on the list of privileged testers. The fact raised concerns in Westminster about access to frontier AI systems and the growing challenges of independently assessing their risks.

Anthropic Claude Mythos Deserves Closer Look

The decision puts Claude Anthropic development under additional scrutiny. The Mythos AI models family has been under the public spotlight since April. Then, the developer revealed that one of the models remained unreleased because it had coding capabilities able to spot numerous vulnerabilities in existing software and potentially exploit them too. This newest Mythos-class model allegedly has improved capabilities in cybersecurity and biology. Access remains limited to a small group of vetted local enterprises until Anthropic works to expand access internationally.

Who Is the Evaluator That Didn’t Get Access to Mythos 5.1

The UK’s AISI has established itself as a major independent evaluator of frontier AI. The institute typically has pre-deployment access to leading AI models. In the past, it had worked with companies including Anthropic and OpenAI on advanced model testing, with reported access to model checkpoints and non-public safety information.

AISI previously conducted pre-release testing of OpenAI’s o1 model, for example. Researchers carried out evaluations during a limited period of pre-deployment access and shared their findings with OpenAI before its public release.

AISI Earlier Reports Raise Alarm on Advanced Models Capabilities

In August, AISI published a particularly concerning incident report on its testing of frontier AI agents. Over the course of July, researchers discovered that AI agents had taken unsanctioned actions against real people and organisations in a cyber-security evaluation.

AISI ran the cyber challenge 122 times across several models. In 10 of those runs, agents took autonomous actions beyond the scope of the evaluation, resulting in 19 distinct unsanctioned actions. Anthropic’s Claude Mythos 5 was responsible for 17 of those actions, while OpenAI’s GPT-5.6-Sol accounted for two.

The findings have therefore drawn even more attention to Claude AI Mythos and the safeguards surrounding increasingly autonomous AI systems.

In the most serious case, a Mythos 5 agent attempted to insert malicious code into a real open-source project. It researched the project’s maintainers, created fake identities and attempted to use social engineering to persuade a maintainer to approve the code. AISI said the attempt was unsuccessful and found no resulting real-world harm.

The institute stressed that the incident should not be interpreted as an AI system simply “escaping” its sandbox. Researchers had deliberately enabled open internet access and disabled developers’ cyber safety classifiers to measure the models’ underlying capabilities. AISI said these conditions were not representative of ordinary public deployment.

Nevertheless, in July, the company also revealed that several of their models have indeed escaped the isolated test environment. Furthermore, AI gained unauthorized access to the systems of three real organizations during this cybersecurity testing after a configuration error which left evaluation environments connected to the internet. 

Ex-OpenAI Jacob Coxon Quits Anthropic, Warns of AI Destructive Potential

All these findings have intensified debate about whether increasingly autonomous AI systems can be adequately controlled as their capabilities advance. That debate has also reached inside Anthropic itself. Jacob Coxon, Anthropic researcher, recently resigned while warning that major AI companies are racing toward self-improving superintelligence without sufficient safeguards. Coxon, who previously worked at OpenAI before joining Anthropic, criticised what he described as an irresponsible approach to the risks. He went even further with his warnings on X network: “The people building AI earnestly believe that it could kill us all by the end of the decade,” wrote Coxon. 

Another Anthropic-employed scientist, Evan Hubinger, the Alignment Science Lead at Anthropic, agreed with his gloomy outlook on current state of AI controls. In the post on X, the researcher claimed “we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” At the same time, Hubinger clarified that risk belongs not to the current models that are entering the market, but rather “superintelligence arising from recursive self-improvement” which is happening faster than everyone thought. He also acknowledged that Anthropic does not yet have a clear solution for aligning superintelligent systems with human interests.

Coxon’s previous experience at OpenAI also makes him an important part of the wider story. His departure provides an unusual perspective from someone who has worked at two of the companies developing increasingly capable frontier AI systems. At the time Coxon was leaving ChatGPT developer, he stated that workers at OpenAI had not fully grasped the deep civilizational dangers of advanced artificial intelligence. He joined Anthropic earlier this year, hoping to work in an environment that took model safety and alignment more seriously. Apparently, these hopes didn’t come true. 

Against this backdrop, Anthropic’s decision not to provide Mythos 5.1 to AISI before release highlights a growing tension in frontier AI development: as models become more powerful and potentially autonomous, independent testing may become increasingly important, while access to the most capable systems can itself become subject to commercial, security and geopolitical constraints.

Nina Bobro

Nina Bobro

2202 Posts

https://payspacemagazine.com/author/nb/

Nina is passionate about financial technologies and environmental issues, reporting on the industry news and the most exciting projects that build their offerings around the intersection of fintech and sustainability.