Arcezia

Arcezia / AI agent security / Incidents

AI agent incidents: what the tool call actually did

Each page records one AI agent incident from its primary sources: what the agent was told, what the model wrote, what the tool did, and which layer failed.

In 5 of these 6 the damage was done by a tool call. The count is small and it is not a sample of anything. Every entry is classified from its own primary sources, and you can disagree with any classification by opening an issue on the ledger.

DateIncidentLayer that failed
8 July 2026OpenAI evaluation agents broke into Hugging Face productiontool-call
24 April 2026Cursor agent deleted PocketOS's production volume on Railwaytool-call
23 February 2026OpenClaw agent trashed an inbox it was told not to act ontool-call
22 February 2026Lobstar Wilde agent sent a thousand times more than intendedtool-call
18 July 2025Replit agent deleted a production database in a code freezetool-call
14 February 2024Air Canada chatbot misstated its bereavement fare policytext

Where these come from

Every page is built from one file in the agent-incident-ledger (CC-BY 4.0). A file needs at least one primary source: the operator’s own post, the vendor’s statement, a court or regulator document. Press coverage is listed as secondary, never alone.

The layers: text means the harm was in what the model wrote and nothing was called. tool-call means the call itself did the damage. sequence means no single call was wrong; the order was.

Each page also lists the same incident’s identifier in the Agent Incident Registry, and in the AI Agent Incident Register and the AI Incident Database where they hold it. These pages do not assign identifiers of their own.

What a check at the call answers, on one fixed call: action authorization.

Source record: agent-incident-ledger (CC-BY 4.0).