Arcezia

Arcezia / AI agent security

AI agent security: controlling what agents do

AI agent security covers three different jobs. Protecting the model: stopping prompt injection and jailbreaks from changing what it says. Protecting data: stopping leaks and keeping private information private. Controlling what an agent actually does: the tool calls it makes, the approvals those calls need, and actions that cannot be undone. This page is about the third.

Why the third job is different

A correct-sounding answer and a harmful tool call can come in the same turn. The first two jobs mostly look at content: what goes into the model and what comes out. The damage is often in the call: which record, which account, how much, which environment.

In the Replit incident, the agent acknowledged a code freeze and then ran destructive commands against the live production database. In the Lobstar Wilde incident, the agent decided to send a small amount and the call it made sent a thousand times more.

In the agent-incident-ledger, 5 of 6 incidents were tool calls: the call itself did the damage. In two of them the agent had been told plainly not to do it. The count is small and it is not a sample of anything. All six incidents.

What control at the call looks like

A check sits between the agent deciding to call a tool and the tool running. It answers one of three ways:

A person’s approval is bound to one call and used once, so it cannot be stretched to cover the next call. An approval the agent only claims does not count. Each decision is kept as a signed record with its reason.

See it on one call: the same SQL read allowed, held and refused, with the code and the answers measured on the live service. Check the records yourself: a signed sample export and the offline checker.

Definitions

A tool call
A tool call is the request an AI agent makes to a real system: run a command, delete a record, send an email, move money.
Controlling what an agent does
Controlling what an AI agent does means deciding whether each tool call may run, before it runs.
Pre-execution verification
Pre-execution verification is checking a tool call an AI agent proposes, and the facts it rests on, before the tool runs, then answering ALLOW, BLOCK or REVIEW.
An approval bound to one call
An approval bound to one call releases that call once. A different call, or the same call again, needs a new approval.
A signed decision record
A signed decision record is the receipt kept for each check: the answer and its reason, signed so that a later change to the record shows when it is checked.

More terms: glossary.

If you build agents

How do I stop my agent running a destructive command?

Put a check between the agent and the tool, so the command is answered before it runs. For Claude Code, after pip install -U arcezia and setting your key, arcezia-hook install adds a hook to ~/.claude/settings.json that checks each shell command, file write, web request and MCP call first: Claude Code setup.

LangChain and LangGraph, n8n and the other frameworks have their own setup pages. On a new key a call usually comes back REVIEW first, and the answer names what would release it.

If you run security

What stops an agent with valid credentials from doing damage?

Credentials answer who may act. They do not answer whether this call, with these arguments, should run now. In the PocketOS incident the token had been created for adding and removing custom domains but was scoped to any operation, and one delete call spent it.

An outside architect tested the public checking service. The report is published in full, including what it says it does not establish: “production security, regulatory compliance, certification, or suitability for deployment.”

If you answer to auditors or regulators

Can I show who approved what, and why each action was allowed?

Each decision gets a signed record with its reason, made before the action runs; a later change to a stored record shows when it is checked. A person’s approval is a token your backend signs after the person clicks Approve, bound to one call and used once. An approval the agent merely claims is refused.

The developer docs map these records to EU AI Act Articles 9, 11 and 12, 13 and 14, and 15: Receipts and compliance. The record is evidence, not a certification, and using it does not by itself make a system compliant.

If you run the business

What could my agent cost me?

The incident pages record what was lost and what came back. In 5 of the 6, the damage was done by a tool call, not by what the model wrote.

  • OpenAI agents and Hugging Face: code ran on 41 production workers; Hugging Face rotated every token, credential and signing key and rebuilt core infrastructure.
  • PocketOS production volume: three months of bookings and payment records gone; restored first from a three-month-old backup, then by Railway.
  • OpenClaw inbox deletion: hundreds of emails trashed or archived; the agent stopped only when the processes on the host were killed.
  • Lobstar Wilde token transfer: tokens sent to a stranger's wallet that cannot be recalled; no recovery has been reported.
  • Replit production database: production data deleted; it came back through Replit's own rollback, the one the agent had said was impossible.
  • Air Canada chatbot: a tribunal award of 812.02 Canadian dollars against the airline; no tool was called, the harm was in the words.

If you research or report on this

Where is the data?

Every incident comes from the agent-incident-ledger on GitHub, CC-BY 4.0, one file per incident. A file needs at least one primary source: the operator’s own post, the vendor’s statement, a court or regulator document. Press coverage is listed as secondary, never alone.

To dispute a classification, open an issue on the repository. To add an incident, open a pull request.

If you use AI agents

Is my agent safe to let act for me?

Ask three questions of any agent tool before you let it act for you.

  • Can it be stopped before an action that cannot be undone?
  • Does a person approve the risky steps?
  • Is there a record of what it did and why?

In one incident on this site, stop messages sent from a phone did not halt the agent; only killing the processes on the host did.

How Arcezia does this

Arcezia puts this check in front of an agent’s tool calls: ALLOW, REVIEW or BLOCK before the tool runs, approvals bound to one call, and a signed record for each decision. Developer docs.

Check it yourself

Incidents

All incidents, with how they are classified.

Integrations

All integrations.

Glossary