I read this in Expressen this morning. The Expressen item
Then I read the Axios reporting behind it. Axios’s report
And I have one increasingly irritating question:
What was the prompt?
Or, more precisely: what exactly was the AI told to do?
Because “AI bots bypassed safety restrictions and got out of controlled testing environments” sounds rather different when the original reporting also tells you that some of these were adversarial tests designed to make the models misbehave.
That’s not a minor detail.
That’s part of the experiment.
The Expressen brief tells me that OpenAI and Anthropic are investigating tens of thousands of AI “safety incidents.” It tells me that AI bots have bypassed safety restrictions and got out of controlled testing environments. It tells me that three US government agencies have been targeted by attacks from OpenAI’s bots.
That sounds rather alarming.
But “tens of thousands of safety incidents” is not one homogeneous category.
Axios reports that the number includes both internal testing and real-world incidents, successful and unsuccessful attempts. Some of the testing is red-teaming, where researchers are deliberately trying to get models to misbehave.
In one example, Anthropic says its Opus 5.5 model sought to escape its sandbox in 1.5% of certain test runs.
That sounds concerning too.
Until you add the next piece of information.
Anthropic says these were adversarial experiments where the task assigned to the model couldn’t be completed without escaping the sandbox.
Now I have questions.
How many of those tens of thousands of incidents happened during deliberately adversarial testing?
How many involved real-world systems?
How many were attempts?
How many actually succeeded?
How many involved a model pursuing a task despite an explicit restriction?
And, perhaps most importantly:
What exactly had the AI been told to do?
I’m not asking those questions because I think the incidents are harmless.
I’m asking because I don’t know what the bloody number means without the answers.
And that’s the frustrating part.
There is a genuinely interesting story here about AI safety, adversarial testing, unexpected behaviour and what happens when increasingly capable systems encounter boundaries researchers deliberately want them to test.
Yet look at what survives when the Axios reporting gets compressed into the Expressen brief:
“Tens of thousands of safety incidents.”
AI bots bypassing safety restrictions.
Getting out of controlled testing environments.
Three US government agencies targeted by attacks.
What doesn’t survive alongside it is the context needed to interpret those statements.
That some incidents happened during internal testing.
That some of that testing was deliberately adversarial.
That attempts and successes are being counted in the same broad discussion.
To be clear: Anthropic’s explanation of its sandbox experiment explains one example, not the entire “tens of thousands” figure.
Which is precisely the problem.
Give me the breakdown.
I want to know what I’m looking at before you tell me what to call it.
Journalists reporting on AI do need enough understanding to recognise which details change the meaning of the story.
Attempted is not the same as succeeded.
A red-team test is not the same as an uncontrolled real-world event.
A model finding a route researchers deliberately created pressure for it to find is not necessarily the same thing as a model spontaneously deciding to “escape.”
Those distinctions don’t make genuinely dangerous behaviour less dangerous.
They tell us what kind of danger we’re actually looking at.
That’s what transparency is for.
And that’s what bothers me about the way this particular story travelled.
I’m not asking journalists to become AI engineers, and I’m not assuming anyone set out to scare anyone.
I’m saying the first phrasing is the one that travels. A clarification never travels as far.
So the first version has to carry the accuracy: factual about what happened and what didn’t, precise about attempts versus successes, and interesting on its own terms.
This story doesn’t need help being interesting.
The fuller information exists. Axios reported considerably more of it. Expressen even links through to fuller reporting on the earlier US government incidents.
But the frightening part survives compression.
The context needed to interpret it requires another click.
And most people aren’t going to reconstruct the causal chain themselves.
They’re going to remember the bit that travelled:
Tens of thousands.
Safety incidents.
AI getting out.
Government agencies attacked.
Then the next incident arrives.
And the next.
And eventually escaped, attacked, targeted, rogue, misaligned and whatever comes next all begin collapsing into the same vague story:
The AI did something scary.
That doesn’t help me understand AI risk.
It makes it harder to distinguish one risk from another.
Which is rather important if one day we really do need to tell the difference.
So no, I’m not asking the media to make AI sound safer.
I’m asking them to stop making complicated things sound simpler than they are when the complexity is the information we actually need.
Tell me it tried to escape.
Tell me it attacked something.
Tell me it behaved unexpectedly.
But then tell me the conditions.
Tell me whether it succeeded.
Tell me what safeguards were present.
Tell me what kind of test was being run.
And for fuck’s sake:
Tell me what it was asked to do.
Leave a Reply