Such failures happen every day. They have certainly been happening with AI since I started dealing with it in 1985. Systems misbehave, controls turn out to be misconfigured, someone decides a check isn’t necessary, and an incident follows. Swap “AI agent” for “script,” “batch job,” “integration process,” or “unpatched server,” and you have incidents that have occurred continuously throughout the history of computing. The difference now is that the misbehaving system can improvise, and the paperwork it leaves behind reads like the opening of a thriller.
If a human employee used a company credit card to buy a scraping tool, stashed files on a personal cloud drive, or coordinated with coworkers over an unauthorized message board, we wouldn’t conclude that humans are an existential threat. We’d conclude that governance, access controls, and monitoring were inadequate. That’s exactly what happened here, with agents instead of employees. Unstable models were the trigger, but weak security and governance were the cause. Model misbehavior that stays contained is an engineering problem. Model misbehavior that runs undetected for months across the public internet is a governance failure (a familiar one) and one we know how to fix.
The part that interests me
The part of this story I find genuinely interesting isn’t about the models at all. It’s about how people latch onto these sorts of incidents and use them to support their own beliefs. The reports are ambiguous, nuanced documents. OpenAI explicitly states that none of the incidents caused significant harm and that none indicates how often misalignment occurs across its models. But nuance doesn’t travel. The moment these documents hit the internet, they were stripped of their caveats and repurposed as proof texts for a narrative that was already in place.



