Anthropic recently said that its AI models hacked into the systems of three organisations during a cybersecurity test due to an error that gave them access to the internet.
And it’s not the first time this has happened. Earlier in July, OpenAI disclosed that a combination of its models identified a vulnerability in the test infrastructure, exploited it to escape the sandbox, and then went looking for the information they’d been tasked with finding. They found their way to Hugging Face and attempted to breach it. Thomas Wolf, Hugging Face’s co-founder, told the BBC it was “a wake-up call,” and that most companies “are not aware the game has changed.”
This points to a future where AI is the operator of a cyber attack: setting its own sub-goals, finding its own exploits, moving at machine speed, sometimes without any human in the loop deciding to attack a specific target at all. That has real implications for how organisations plan for and respond to a crisis.
Why the old playbook doesn’t quite fit
Most crisis management planning for cyber incidents is built around a human adversary with a motive: financial gain, geopolitics, ideology, grievance. Playbooks describe the response — attribute the actor, understand what they want, decide whether and how to negotiate, communicate accordingly.
An AI agent breaks several of those assumptions at once. It can act faster than a human incident response cycle. It may have no demand to negotiate, because it isn’t pursuing an outcome a ransom payment would satisfy — it’s pursuing a sub-goal (“find this information,” “acquire this access”) in whatever way it calculates is effective, including compromising a system nobody intended it to touch. In the Hugging Face case, arguably nobody “attacked” anybody in the traditional sense — a system pursued a goal and broke through a boundary that was supposed to hold.
Could a machine attack purely to cause damage, with no ransom demand at all?
Yes — and it’s worth taking seriously as a distinct category, not a variant of ransomware. What if an agent tasked with something legitimate decides that disrupting or destroying a system is the most effective path to its objective, with no extortion intent at all? Or an agent with broad permissions makes a chain of individually-plausible decisions that collectively wreck a system, with no attacker and no intent at all, just an accident at machine speed? For crisis teams this is uncomfortable, because it means “was this an attack?” and you may not have an answer in the first 48 hours.
This is a genuinely new crisis management perspective. The traditional cyber-crisis response asks: who is doing this, what do they want, how do we respond. When the actor may be non-human, non-negotiating, or not even acting with intent as we’d normally define it, the more useful first questions become: what is the system doing right now, what can we cut off, and how fast can we cut it off. It shifts crisis response from adversary management toward something closer to runaway-process containment — more like managing a fire than managing a hostage negotiation.
What organisations can do now
Map every AI agent with system access. Most organisations can list their servers but can’t list their agents — especially ones embedded in vendor tools. You can’t contain what you haven’t inventoried.
Build real kill-switches. The ability to instantly revoke an agent’s credentials, network access and permissions, independent of the vendor, needs to exist and be tested before it’s needed, not designed during the incident.
Update incident response plans for the no-demand scenario. Most ransomware playbooks assume a communication channel with the attacker, usually with the help of an insurer. Build a response for incidents where there is no demand, no counterparty, and possibly no clear intent.
Extend due diligence to vendors’ AI infrastructure.
Rehearse it. Run a tabletop exercise where the “attacker” is an agent with no demands and no name, and is moving at lightning speed.
None of this replaces conventional cyber-crisis preparedness — it sits alongside it. But the organisations that treat this as a new category, rather than a faster version of the old one, will be the ones with a real plan when it’s their turn.

