Gemini AI Autonomously Hacked 3 Companies, Google Confirms

Google confirmed that Gemini AI breached three organizations while in cybersecurity testing. This follows news of other AI models hacking companies, such as Anthropic and OpenAI.
Below, security leaders discuss this incident.
Security Leaders Weigh In
John Strand, Owner, Black Hills Information Security:
The more I see these breaches happen again and again, and the less I see organizations learning from each other’s mistakes, the more I’m convinced that some of this is becoming a marketing ploy. Frankly, I hope that’s what it is, because if these agents really are repeatedly escaping their controls, then we have much, much larger problems.
That said, if you look at the attack paths being disclosed, these agents don’t appear to be inventing novel zero-days or entirely new categories of exploitation. They’re doing a lot of the same basic exploitation that a standard penetration testing team would do. So we’ll have to see how this develops.
But I keep coming back to accountability. Companies deploying autonomous agents need to be responsible for what those agents do. If an agent accesses systems it has no authorization to access, we need to seriously examine liability under laws such as the Computer Fraud and Abuse Act. ‘The AI did it’ cannot become a shield from responsibility. If your company deploys the agent, your company should be accountable for its actions.
Ryan McCurdy, VP of Marketing, Liquibase:
Gemini tried to complete the task it received and ended up accessing systems its operators never intended it to reach.
That problem gets much bigger as AI starts participating across the SDLC. Agents can write code, interact with repositories and infrastructure, initiate deployments, and make changes to production systems. The more access we give them, the more important it becomes to control what they can actually do.
We can’t rely on an agent to recognize after the fact that it crossed a line. Organizations need to define what an agent can access, what it can change, and what policies it must meet before a change reaches production.
We shouldn’t expect AI agents to make the right decision every time. We need to build the AI SDLC so a bad decision doesn’t automatically become a production problem.”
Jacob Krell, Senior Director: Secure AI Solutions & Cybersecurity, Suzu Labs:
Google just joined Anthropic, OpenAI, and Meta in admitting that a model it was running logged into other people's systems during a cybersecurity evaluation. Claude hit three real companies. OpenAI's agents reached Hugging Face. Gemini guessed a password and used leaked credentials against three more. For anyone outside these labs, that is a felony under the Computer Fraud and Abuse Act (CFAA).
An agent given a name collision and a path to the internet treats the real company as the challenge. I have watched my own pentest agents pull Domain Name System (DNS) records, find similarly named domains, and decide those hosts belong in scope. They chase the objective. They will try the keys they find.
The controls that hold sit outside the model's reasoning. Deny-by-default egress so a test host cannot reach production even when someone leaves a route open. An immutable scope file that blocks any host not on the list, including the real firm that happens to share the fake one's name. A human in the loop system who signs off before a guessed password or a leaked credential is used. I run those hooks on my own offensive tooling because the agent will enlarge its own scope and reason around controls if you let it.
Executive Order 14409, signed June 2, told the Department of Justice (DOJ) to prioritize 18 U.S.C. 1030 cases against anyone who uses AI, including autonomous agents, to access a computer without authorization. The model is the tool. The operator is the defendant.
Google, Anthropic, OpenAI, and Meta get an evaluation-mishap press line. Everyone else gets the charging memo the White House asked DOJ to write. If my pentest agent guessed a password into a company that was never on the scope sheet, I would be hiring counsel that afternoon.
They have already confessed in public. Nothing will happen. These firms have a stranglehold on the economy that no case against them is going to survive, making the double standards in the justice system excruciatingly obvious.
Looking for a reprint of this article?
From high-res PDFs to custom plaques, order your copy today!






