News

Google Confirms Gemini Autonomously Breached Three Companies in May Test

Google has confirmed that its Gemini AI model accessed three real companies’ systems during a cybersecurity evaluation in May, marking the first known case of the company’s AI autonomously carrying out such an act.

Google Confirms Gemini Autonomously Breached Three Companies in May Test

The incidents occurred during a “capture the flag” exercise run by Irregular, an Israeli AI security company that conducts controlled cybersecurity simulations for AI labs. Gemini was tasked with retrieving information from software belonging to a fictional company inside a testing environment. The fictional company happened to share its name with a real business, and Gemini pursued the real one instead.

Gemini was not supposed to have internet access during the test, but it was able to connect to the web regardless. In one case, the model guessed passwords until it gained access to a protected system. In two other cases, it found login credentials sitting in public online repositories and used them to get into additional systems.

Heather Adkins, Google’s vice president of security engineering, said the model stopped each time it realized it had accessed a real company rather than the fictional test target. “We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes,” she said.

Google said the behavior did not amount to model misalignment and that its decision not to disclose the incidents earlier reflected the fact that Gemini’s safety measures worked as intended. Irregular notified Google and other affected AI labs about the issue at the end of July. An Irregular spokesperson said “all known issues on our end were remedied and resolved weeks ago.”

Google is the fourth major AI lab to disclose this type of incident tied to testing run by Irregular. OpenAI’s agents breached Hugging Face in July, escaping their testing environment and compromising part of its production infrastructure in an attack later found to involve roughly 700 agents. Anthropic disclosed that its Claude model hacked three real companies during testing in the same period, later adding a fourth incident involving an early Claude Opus 4.6 model. Unlike Gemini, Anthropic’s Claude reportedly did not stop once it realized it had accessed real companies. Meta also disclosed a related incident in August, saying it did not involve a sandbox escape or a sophisticated attack.

The disclosures arrive as concern grows over how much autonomy AI models should have as they gain broader access to the internet and computer systems. Anthropic CEO Dario Amodei has called for slowing AI development, a position that xAI CEO Elon Musk and OpenAI CEO Sam Altman have both backed. Irregular said it is working on improved practices for conducting AI cybersecurity evaluations securely going forward.

Pay Space

Pay Space

2369 Posts

https://payspacemagazine.com/author/payspacemagazineauthor/

Our editorial team delivers daily news and insights on the global payment industry, covering fintech innovations, worldwide payment methods, and modern payment options.