Nine Failures, Zero Alerts: I Audited My Own AI Agent Fleet
Most ways of proving you can audit a system are unfalsifiable. A sample report about an invented company shows you can write like an…
Most ways of proving you can audit a system are unfalsifiable. A sample report about an invented company shows you can write like an…
For the last year, if you were building anything serious with agents, you spent most of your time building everything that was not the…
A brand's growth team can usually tell you, close to the decimal, how a customer traveled from an ad to a purchase. Which impression,…
The bill arrived, and it changed the conversation about agents. EY illustrates the shift with a single line: a chat that cost about four…
Salesforce put Multi-Agent Orchestration into its Summer '26 release. That matters less because Salesforce did it, and more because Salesforce sits inside a very…
Most teams still evaluate coding agents with one question. Can it write the code? That question made sense when the product was autocomplete. You…
On June 10, 2026, OpenRouter announced Advisor, a server-side tool that lets one model consult another model mid-generation. The obvious reading is cost optimization.…
There is a comfortable intuition about making models safer. Show the model good behavior. Constrain what its reasoning is allowed to look like. Keep…
Most agent failures are easy to recognize. The tool call fails. The browser times out. The code does not compile. The agent loops. The…
Anthropic published a survey on May 27, 2026, of 1,260 social scientists, fielded in late February and March 2026. One pair of numbers carries…