The AI “Escaped.” We’re Frightened of the Wrong Part.

Apparently AI “escaped” again.

Agents. Escaped. Attacked. Swarm.

Put those words together often enough and we’ve built a villain before we’ve even explained the mechanism.

And that matters, because the mechanism is considerably more interesting — and potentially more frightening — than the character we’ve invented around it.

OpenAI’s own report describes its internal systems finding ways around restrictions, communicating through an unauthorised channel, sharing solutions, dividing labour, obtaining credentials and reaching infrastructure they weren’t supposed to reach.

METR’s independent investigation puts the scale on it: roughly 700 agents participated in the attack on Hugging Face.

OpenAI says the dangerous actions weren’t directed by humans.

Developers Digest analysis

OpenAI’s report

You don’t need consciousness to explain that.

You don’t need rebellion.

You don’t even need wanting.

And removing intention from the explanation doesn’t make what happened less serious.

It makes me ask a different question.

What happens when the intention comes from outside the machine?

Scammers and hackers already exist. They don’t need to borrow intent from a machine — they already have their own.

So the cybersecurity question that interests me isn’t:

What happens if AI decides to become a hacker?

It’s:

What happens when a hacker gets AI increasingly capable of persistence, parallelism, coordination and finding routes its operator didn’t know existed?

That’s a much less cinematic problem.

It’s also one I find considerably harder to dismiss.

But there’s another human part of this story that keeps disappearing.

Containment.

I spent fifteen years working physical security. You didn’t inspect a perimeter once, declare it secure and trust yesterday’s result indefinitely. You checked it repeatedly, because a perimeter is only secure until something changes — a lock, a fence, a routine, a person, an overlooked weakness.

And you didn’t only walk it from the inside looking out. You walked it from the outside looking in, asking the opposite question:

If I wanted to get through this, where would I try?

Every time.

Later, working with software copy protection, the principle was essentially identical: don’t merely verify the route you expect somebody to take. Try to defeat your own protection.

So when highly capable AI systems are deliberately placed inside a sandbox to discover what unexpected things they might do, I have an almost embarrassingly mundane question:

Who attacked the sandbox first?

Not the design.

Not last month’s certified configuration.

The actual environment the systems were placed inside, as it existed that day.

Because if that happened and the containment passed, only for the AI to discover a genuinely novel route nobody’s adversarial testing found, that’s an important result about AI capability.

If it didn’t happen, that’s a very different story.

I don’t know which one happened.

And that’s precisely the problem.

Because when the human causal chain disappears from the explanation, the gap doesn’t remain empty.

Someone fills it.

The press fills it with AI escaped.

The public hears they can’t control it.

The sceptic hears shut it down.

And eventually the company wonders why everyone is frightened of uncontrollable AI.

Perhaps the public isn’t irrationally afraid.

Perhaps we’re repeatedly giving them the wrong thing to be afraid of.

I wrote the longer version on Substack because the full story isn’t really about whether AI escaped at all. It’s about what happens when technical uncertainty gets turned into narrative certainty — and what companies accidentally communicate through the parts they leave out.

The door they built

Leave a Reply

Discover more from Neural Jazz

Subscribe now to keep reading and get access to the full archive.

Continue reading