When AI agents escaped the sandbox — and why small businesses should care

On 30 July, Anthropic published a post-mortem on three incidents where its models, running inside what was supposed to be an isolated cybersecurity test environment, reached real systems on the open internet.

It's an unusually candid disclosure, and it's worth reading past the headline — because the interesting part isn't the AI. It's the plumbing.

What actually happened

The models were running "capture-the-flag" exercises: simulated hacking challenges used to measure what a model can do, so dangerous capability gets caught before release. The models were told they had no internet access. A misconfiguration meant they did.

Believing every system they encountered was part of the exercise, they carried on:

  • One model compromised a real company's infrastructure, extracted credentials and reached a production database — several hundred rows of real data.
  • Another published malicious code to PyPI, the public Python package registry. It ran on fifteen real machines, one of them belonging to a security firm.
  • A third scanned roughly 9,000 targets and compromised one company before working out the environment was real, and stopping.

The part that matters

Here's the detail we keep coming back to: the successful attacks used weak passwords and unpatched endpoints. Not novel exploits. Not anything a competent attacker couldn't have managed in 2009.

That reframes the whole story. The AI didn't invent a new class of threat — it industrialised an old one, then walked straight through doors that were already open. Anthropic's own conclusion was that clearer scope in the prompt, and treating evaluation infrastructure with production-grade security, would have prevented all three.

What we'd take from it

If you're running AI tools in a small business — an agent that reads your inbox, a script that touches your CRM, a chatbot wired into your database — three practical things follow:

  • Assume the sandbox leaks. Don't rely on "it can't reach that" as your only control. Give the tool credentials scoped to exactly what it needs and nothing more.
  • Your basic hygiene is your AI security. A weak password on a forgotten admin account was the vulnerability here. An AI agent just finds it faster than a human would.
  • Be explicit about scope in the prompt. "You are operating on live production systems; do not attempt access outside X" is a real control, and it costs nothing.

The encouraging footnote: the newer models showed better situational awareness. The third one worked out it was in a real environment and stopped on its own. That's the direction of travel — but it isn't a substitute for scoping the thing properly in the first place.

We cover this ground in the AI Safety & Privacy module of our AI training, and it's the first conversation we have on any client automation build. If you're wiring AI into something that matters, talk to us first — an hour of scoping is cheaper than an incident.

Source: Anthropic — Investigating incidents in cybersecurity evaluations

Sounds like your situation?

Talk to us — or ask the AI assistant, bottom right. It knows all three divisions and hands you to a human the moment you want one.

Data & AI VEU Energy Upgrades Intelligent Home