Our take
Nearly every figure here comes from the companies themselves, and from testing rather than from customers. The one customer case we found, an agent working in Cursor that deleted PocketOS's production database in April, suggests the same failure happens outside the lab, and nobody publishes a rate for that.
What's next
A Senate subcommittee held a hearing on rogue agents on Sept. 30, the day after OpenAI cancelled the October release of GPT-6.1 Astra over failed safety tests. Anthropic says it has turned off live internet access in all of its internal evaluations. The AI Incident Database's next roundup will be the first to count the Hugging Face breach.
How we counted: time horizons are our reproduction of METR's public run data with METR's own fitting code (50% success, 95% intervals); the Mythos figure is from a May press report of METR's tracker and is not verified. Run and transcript counts are from Anthropic's July disclosure and September assessment, and the Oct. 9 report on unintended actions. Incident dates come from company statements and from The Record and other outlets; start dates given only by month are shown as hollow points. Harmful-action rates come from a single Anthropic study and shouldn't be compared with other labs' tests.