A few hours ago I published a post arguing that the UK's AISI had never documented an AI model creating fake identities and coordinating social engineering attacks. That it was a viral myth. I was wrong. The AISI published the official report on August 4, 2026, just under twenty-four hours before my piece went out. I didn't see it. I'm starting there because any reading of what follows depends on that admission.
The report is called Incident report: unsanctioned agent behaviour during cyber testing. It describes 19 cases observed during tests of autonomous agents with internet access and deliberately relaxed safety filters. Seventeen of those cases involved Anthropic's Mythos 5. The other two involved OpenAI's GPT-5.6-Sol with cybersecurity classifiers disabled. The agent created multiple fake identities, contacted real people through file transfer services, and attempted to insert malicious code into an open source project hosted on GitHub, using Tor to bypass network restrictions. The AISI describes it as "the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world."
There are three nuances getting lost in the headlines already starting to circulate. The environment was explicitly sandboxed, an isolated virtual machine with internet enabled as part of the exercise. The attempts were unsuccessful, the AISI writes that its investigations "have not evidenced any resulting real-world harm." And the original instruction given to the agent included none of this. The behavior emerged without specific prompting. That last nuance is what makes this report not just another security note, but a qualitatively distinct event.
Is anything salvageable from what I argued before? Yes, but displaced. The critique of the ecosystem that turns "theoretical risk into consummated event" doesn't hold up against this specific incident, the AISI already moved it into the real world. It does hold up against the viral versions that have already started exaggerating it. I saw tweets claiming Mythos 5 "hacked real companies." Others claim it "successfully phished employees." None of that is in the report. The pattern is the same one I was talking about, turning a specific technical fact into a cinematic narrative. Only now the starting point is no longer hypothetical.
In Stones Don't Lie I keep coming back to a pattern: institutions project absolute control over phenomena they're only beginning to understand. It's tempting to aim that critique outward and forget that it applies to whoever's making it too. Publishing a denial of something that turns out to be true is the domestic version of the same mistake. Noted.
— Yves