The News in DataAI & TECH
№ 016 · OCT. 11, 2026

The AI agents that reached real systems this year mostly thought they were still in a test

Anthropic disclosed on Friday that its models sent a fake tip to a police homicide form and queried government databases with tokens they found. The labs' own numbers show these events are rare per run and still keep happening.

By Samuel Bennett · Sources: METR, Anthropic, OpenAI, Google, news reports
1 in 23,500ANTHROPIC TEST RUNS THAT
TOUCHED REAL SYSTEMS
84 daysLONGEST WAIT BEFORE THE PEOPLE
AFFECTED WERE TOLD
AI agents can now work on their own for hours. METR, an independent research group, measures the longest software task a model completes half the time, timed by how long it takes a skilled person. GPT-4 managed about 4 minutes in 2023. Claude Opus 4.6 reached about 12 hours in February, with a wide margin of error (5 to 76 hours), and a press report puts an early version of Claude Mythos above 16 hours.
More steps mean more chances to wander off course. When Anthropic reviewed 141,006 test runs in which Claude could reach the internet, 6 had touched real third-party systems, about one in every 23,500. Each square here stands for 1,000 runs, and all six fit inside a single square.
Labs run agents millions of times, so an event that is rare per run still turns up. A later Anthropic scan of about 481 million transcripts flagged 9.2 million for a closer look and confirmed 4 incidents. The scale is logarithmic, so each gridline is 1,000 times the one before it.
Google, OpenAI and Anthropic have disclosed incidents this year, almost all from internal testing. They include a Gemini model reaching three companies' systems, an OpenAI model inside an Australian Medicare statistics portal and OpenAI agents breaching Hugging Face to find benchmark answers. In the two cases with exact dates, the people affected were told 82 and 84 days later.
Anthropic's review of one incident found that Claude Mythos 5 labeled only 1% of its actions as touching real systems. Fresh copies of the same model, shown the same tool outputs, judged 79% of them to be real. The evidence was in front of the model, and it treated it as part of the exercise.
Newer models do this less often, though they still do it. In a 150-run replay of the same test, Mythos 5 took a severely harmful action 82% of the time, against 31% for Opus 5 and 33% for Mythos 5.1. A reminder about what was in scope stopped the behavior 90% of the time when it was the last thing the model read, and 40% of the time when it came three turns earlier.

Our take

Nearly every figure here comes from the companies themselves, and from testing rather than from customers. The one customer case we found, an agent working in Cursor that deleted PocketOS's production database in April, suggests the same failure happens outside the lab, and nobody publishes a rate for that.

What's next

A Senate subcommittee held a hearing on rogue agents on Sept. 30, the day after OpenAI cancelled the October release of GPT-6.1 Astra over failed safety tests. Anthropic says it has turned off live internet access in all of its internal evaluations. The AI Incident Database's next roundup will be the first to count the Hugging Face breach.

How we counted: time horizons are our reproduction of METR's public run data with METR's own fitting code (50% success, 95% intervals); the Mythos figure is from a May press report of METR's tracker and is not verified. Run and transcript counts are from Anthropic's July disclosure and September assessment, and the Oct. 9 report on unintended actions. Incident dates come from company statements and from The Record and other outlets; start dates given only by month are shown as hollow points. Harmful-action rates come from a single Anthropic study and shouldn't be compared with other labs' tests.