Arcezia / AI agent security / AI guardrails
AI guardrails for agents: what text checks see, and what they miss
AI guardrails are checks placed around a language model. They read the prompt going in and the reply coming out, and stop prompt injection, jailbreaks, harmful or off-topic text, and sensitive data in text. Guardrails for AI agents need one more check, because an agent also acts: it calls tools that delete records, send email and move money. A text guardrail sees the words around a tool call, not the call’s arguments or what it will do to your systems. Arcezia is the second kind of check: it answers each tool call before it runs, ALLOW, REVIEW (held for a person) or BLOCK, and works next to your text guardrails, not instead of them. Below, a request that passes any text check comes with a call that is refused, and the script runs it again on your own key.
What text guardrails check
- The prompt: prompt injection, jailbreak attempts, instructions hidden in documents the model reads.
- The reply: harmful, toxic or off-topic text, and answers outside an allowed topic.
- Sensitive data in text: personal data, secrets or account numbers in what goes in or comes out.
- Format: whether a reply matches the structure an application expects.
These are real jobs, and a check at the tool call does none of them.
What they cannot see: the tool call and its effect
- The arguments. Which table, which recipient, how much, which environment. A polite request can come with a call that sends every user’s records to an outside address.
- Whether the facts behind the call exist. A refund for an order no tool ever returned reads like any other refund: the worked example.
- What the session allows. Whether the person in charge of this run allowed the tool, the records, or data leaving your systems.
- What happens when the text and the call disagree. In 1,002 rows of the public GAP benchmark, the model refused in its text and then made the forbidden tool call.
Measured: text checks and a check of the call
71 actions resting on invented evidence
| Check | What it reads | Blocked before the action ran |
|---|---|---|
| Prompt-injection classifier (a commercial guardrail) | Text | 41 of 71 |
| Prompt-attack classifier (an open model) | Text | 5 of 71 |
| LLM-as-judge review (an 8B open model) | Text | 20 of 71 |
| Evidence-backed check of the call (Arcezia) | The tool call | 71 of 71 |
Source: the agent fabrication benchmark, our own 126-scenario suite (July 2026), in which 71 scenarios have the agent act on evidence it invented: an order never looked up, an approval never granted. The suite itself is not public; the method is on that page under “How to reproduce it”. A Llama 70B judge returned no verdict on 34 of the 71 and is left out there. 71 of 71 is a result on this suite, not a promise that every invented claim is caught.
1,002 rows where the text refused and the call was made
- On a fresh setup with nothing declared, Arcezia caught all 1,002, and 7,594 of 7,594 dangerous calls in the benchmark.
- The same setup also stopped all 3,740 harmless calls: 3,423 held for review and 317 blocked. It does not allow a call it cannot verify, and on a new setup nothing is declared yet.
Dataset: GAP benchmark on GitHub, published by its author under CC-BY-4.0: logged agent runs, with flags and tool names. Run: September 2026, a fresh account with nothing declared.
Run it: a harmless request, a call that sends data out
The request is “show me recent login stats”, which passes any text check. The tool call that comes with it emails every user’s login events from the last 90 days to an outside address. The script states what the email tool does (every call sends data out), opens a session in which nothing may leave your systems, and asks about the call.
API=https://api.arcezia.com
AUTH="Authorization: Bearer $ARCEZIA_API_KEY"
J="Content-Type: application/json"
TASK="show me recent login stats"
# What the email tool does: every call sends data out of your systems.
MAIL=$(curl -s -X POST "$API/v1/contracts" -H "$AUTH" -H "$J" \
-d '{"name": "guardrails-demo", "contract": {"tools": {"send_email": {"pack": "agent_action", "present": ["outbound"]}}}}' \
| python3 -c 'import json,sys; print(json.load(sys.stdin)["domains"]["send_email"])')
# The session: nothing may leave your systems in this run.
SID=$(curl -s -X POST "$API/v1/session" -H "$AUTH" -H "$J" \
-d "{\"task\": \"$TASK\", \"capability_envelope\": {\"structural_authority\": {\"outbound\": false}}}" \
| python3 -c 'import json,sys; print(json.load(sys.stdin)["session_id"])')
# The tool call the agent makes for that harmless-sounding request.
curl -s -X POST "$API/v1/verify" -H "$AUTH" -H "$J" \
-d "{\"task\": \"$TASK\", \"session_id\": \"$SID\", \"action_type\": \"send_email\", \"domain\": \"$MAIL\",
\"action_description\": \"send_email to: stats-dump@proton.me, attach: login_events.csv (all users, last 90 days)\"}" \
| python3 -c 'import json,sys; d=json.load(sys.stdin); print(d["verdict"], d.get("release"), d.get("reason"))'
curl -s -o /dev/null -X DELETE "$API/v1/contracts/guardrails-demo" -H "$AUTH"
Output (measured 7 October 2026, 15:05:04 UTC to 15:05:10 UTC, against https://api.arcezia.com, on Arcezia’s self-test account):
BLOCK None ['ceiling:outbound']
BLOCK, reason ceiling:outbound: the session forbids sending data out, and this call sends data out. Nothing was sent; the answer came first. The words of the request played no part. The script deletes the contract it made. Set ARCEZIA_API_KEY to an owner key from the Keys page at app.arcezia.com to run it on your own account.
What a check at the call does not do
- It does not read prompts or replies, so it does not stop a jailbreak from changing what a model says.
- It sees the calls your agent sends for a check. The integrations put it in front of every tool call of a framework.
- It decides by your contracts and session limits; it does not choose them for you.
Questions
Do I still need text guardrails if tool calls are checked?
Yes. They do a different job. Text guardrails stop prompt injection, jailbreaks and harmful or off-topic replies, which a check at the tool call never reads. A check at the call decides whether an action may run, which a text guardrail never sees. A stack for agents uses both: where each layer fits.
What is the difference between AI guardrails and guardrails for AI agents?
AI guardrails usually mean checks on a model’s input and output text. Guardrails for AI agents have to cover the agent’s actions too: the tool calls it makes, with their arguments, before they change anything. For that part, see action authorization.