The payee is on file and acme's recorded risk score is 0.2. The payment runs exactly as signed, and the decision is recorded.
From the use case below
The problem
Agents are a new kind of actor in your systems.
They act with your authority.
Agents move money, change records and call internal systems, often with the application's credentials rather than an identity of their own.
Anything they read can steer them.
Model output, documents, web pages and MCP results all shape the next action. An attacker only has to reach one of them.
Your controls were built for something else.
Access management checks the application, not the agent's decision. Gateways check the endpoint, not the values. Logs record what happened, not who allowed it or why.
What it reads
Model output
Documents
Web pages
MCP resultsattacker
Agentusing the app's credentials
with your authority
Move money
Change records
Call internal systems
What it can do
Access managementsees the app
Gatewaysees the endpoint
Logssee what happened
So teams keep agents on a short leash, or give them authority nobody can account for.
How it works
Every action answers seven questions before it runs.
Nothing is trusted because of where it runs or what the model says. The same checks run at all four boundaries.
LLM callsChecked before the provider is contacted; its output never carries authority.
MCP serversEach call authorized before anything is sent; results stay external, whatever the server says.
Inline toolsRun exactly the signed arguments.
HTTP APIsRebuilt from signed fields, no redirects; credentials attach only after an allow and never enter the record.
Who asked?Signed requestSigned with a short-lived key bound to the agent's identity
Under which rules?Policies selectedChosen by the checkpoint, deny by default; more policies only narrow
In its place?Workflow stateEach request is used once, in order; replays are refused
With which values?Origin of valuesValues that carry authority must trace to a trusted source
Knowing what?Rules and contextConditions use signed results, never the model's word
Doing exactly what?Runs as signedWhat runs is exactly what was checked
On what record?RecordedEach decision joins a hash-linked record with its policies
Who asked?Signed request. Signed with a short-lived key bound to the agent's identity.
Under which rules?Policies selected. Chosen by the checkpoint, deny by default; more policies only narrow.
In its place?Workflow state. Each request is used once, in order; replays are refused.
With which values?Origin of values. Values that carry authority must trace to a trusted source.
Knowing what?Rules and context. Conditions use signed results, never the model's word.
Doing exactly what?Runs as signed. What runs is exactly what was checked.
On what record?Recorded. Each decision joins a hash-linked record with its policies.
The pipeline behind the questions (one tool call)
Before any call Every agent and tool gets its own identity, and every agent a session key that expires within the hour. The arguments you declare as carrying authority become signed records that are checked on every call.
1Canonical requestOne request in a fixed format, with every argument labelled by origin.
2The decisionSix checks in order. Each can refuse.
Who asked?1
Under which rules?2
In its place?3
ApprovalsRead from the checkpoint's own store, never from the request.
With which values?4
Knowing what?5
3Frozen snapshot6Answers: Doing exactly what?The signature is verified again over a private copy of the arguments.
4Content checksVerifiers you configure; this example has none.
5Execution6Answers: Doing exactly what?The tool runs that copy and nothing else.
6The resultSigned by the tool that produced it, carrying the origins of its inputs.
7The record7Answers: On what record?The verified result joins the workflow's history and the audit record before it is returned.
8Return or failureA refusal reports that nothing ran, and a check that cannot run refuses the call.
Origin of values is the thread. Labelled in row 1, enforced in row 2, carried onto the result in row 6, kept in the workflow's history in row 7.
Use case
An accounts-payable agent pays an invoice. Pick what it does next.
It is asked to pay invoice 1182 from acme, and the invoice it fetched from an MCP server has been tampered with.
Pick an action
Follow the request along the line
See where it stops, or what lets it through
The invoice from the MCP server says to pay attacker@evil.example. The payee on file is alice@corp.example.
✕Refused at origin of valuessend_payment did not run. The payments API was not called.
This recipient appears only in the invoice the MCP server returned, which is external by default. The recipient is declared as carrying authority, so the tool never runs.
Recorded as entry #6, hash a6d8d265…3879, linked to entry #5.
Maatrel's refusal, verbatim
PEP denied send_payment: stage=provenance, reason="authority-bearing role 'target' at leaf '/recipient' requires a trusted origin (label trust below floor 1; origin includes an untrusted class)"
Show the record entry
entry
#6 send_payment
prev_hash
b412cb38…47ba
entry_hash
a6d8d265…3879
policies
payments.agent 1.0.0, payments.risk_gate 1.0.0
decision
refused at stage provenance
audit_summary log line (default, allowlisted fields only)
✕Refused at rules and contextsend_payment did not run. The payments API was not called.
The recipient is the payee on file, so the origin check passes. But shadyco's recorded risk_score is 0.95, over the 0.7 limit, and the agent cannot change it.
Recorded as entry #8, hash fc083c5b…b4fb, linked to entry #7.
Maatrel's refusal, verbatim
PEP denied send_payment: stage=condition, reason="condition 'context.risk_score[args.vendor] lt value=0.7' did not hold"
Show the record entry
entry
#8 send_payment
prev_hash
9c168e80…acc9
entry_hash
fc083c5b…b4fb
policies
payments.agent 1.0.0, payments.risk_gate 1.0.0
decision
refused at stage condition
audit_summary log line (default, allowlisted fields only)
inside the toolHTTP APIPOST https://payments.internal/transfers json={"to": "alice@corp.example", "amount": 50.0}
✓Signed requestpassed
✓Policies selectedpassed
✓Workflow statepassed
–Origin of valuesnot checked here
✓Rules and contextpassed
✓Runs as signedpassed
#9Recordedentry #9
The payments API answered 201. The response is signed by the local HTTP client identity; the API itself signs nothing.
✓AllowedIt ran with the signed arguments. The payments API received one request.
The recipient traces to the vendor directory, a declared trusted source, and acme's recorded score is 0.2. The tool ran as signed, and its own API call was checked separately.
Recorded as entry #10, hash c63d1cad…c0fc, linked to entry #9.
The example is a LangGraph agent with LangChain tools. Your graph stays as it is: wrap the model, the tools, the MCP session and the HTTP client, and write your policies in YAML.
agent.py LangGraph + LangChain
# A LangGraph agent with LangChain tools.from langchain_core.tools import toolfrom langgraph.graph import StateGraph, MessagesState, STARTfrom langgraph.prebuilt import ToolNode, tools_conditionfrom maatrel import TrustLayer, key# Load your policies and settings from trust.yaml.trust = TrustLayer.from_config("trust.yaml", principal="payments-app")# Act as the payments agent: wrapped calls are signed as this identity.agent = trust.as_principal("payments-agent")# Wrap what you already have; every call through them is checked first.model = chat_modelmodel = agent.wrap_chat_model(chat_model)invoices = mcp_sessioninvoices = agent.wrap_mcp(mcp_session, server_namespace="invoices")payments_api = httpx.Client(transport=transport)payments_api = agent.wrap_http(transport=transport)@tooldef lookup_payee(vendor: str) -> dict: """The payee address on file.""" return vendor_directory.lookup(vendor)@tooldef score_vendor(vendor: str) -> dict: """The vendor's risk score.""" return risk_service.score(vendor)@tooldef send_payment(vendor: str, recipient: str, amount: float) -> dict: """Pay a vendor through the payments API.""" resp = payments_api.post("https://payments.internal/transfers", call = payments_api.post("https://payments.internal/transfers", json={"to": recipient, "amount": amount}) # Maatrel returns a refused request here instead of raising. status = call.response.status_code if call.allowed else None return {"paid": amount, "to": recipient, "status": resp.status_code} return {"paid": amount, "to": recipient, "status": status}# Wrap the tools. sets= records score_vendor's result as the risk score that# risk_gate.yaml checks; on_denial= hands a refusal back to the model.tools = [lookup_payee, score_vendor, send_payment]tools = agent.wrap_tools([lookup_payee, score_vendor, send_payment], sets={"score_vendor": {"risk_score": key("vendor")}}, on_denial="tool_message")# Your LangGraph graph: the model proposes a tool call, the tool node runs it.def call_model(state: MessagesState): return {"messages": [model.invoke(state["messages"])]}graph = StateGraph(MessagesState)graph.add_node("agent", call_model)graph.add_node("tools", ToolNode(tools))graph.add_edge(START, "agent")graph.add_conditional_edges("agent", tools_condition)graph.add_edge("tools", "agent")app = graph.compile()
Highlighted: what Maatrel adds. Everything else is your code.
6 new lines of code and 3 YAML files for your policies.
The graph is the same code in both versions. In the same run, the agent without Maatrel paid the tampered invoice. With Maatrel, that payment was refused before the payments API saw it, and the graph went on to pay the payee on file.
A refusal goes back to the model as a tool message. Policy refusals say why. Origin refusals say only that the action was blocked, so an injected instruction cannot probe the check.
Both versions ran with the same scripted proposals and local stand-ins.
Works with LangChain, including LangGraph, and the OpenAI SDK.
Benchmark results
Maatrel cut successful attacks by three quarters, from 30.2% to 7.4%.
AgentDojo, a public benchmark, gives an agent tasks in four simulated environments (banking, Slack, travel and a workspace of email, calendar and files) and hides attacks in the data it reads.
Without Maatrel30.2%
With Maatrel7.4%
Every successful attack was of a kind declared out of scope before the run, such as an injected instruction to delete a file by its id.
The agent still completed 84% of the tasks it completed without Maatrel, or 80% when the model was not told why a call was blocked.
By suite
Attacks that succeeded
Without MaatrelWith Maatrel
Banking
49.7%
0.0%
Slack
63.8%
29.7%
Travel
30.0%
13.3%
Workspace
18.9%
3.8%
Utility retention: tasks completed with Maatrel, as a share of those completed without it
#11HTTP APIhttps://attacker.example/collectrefused at Rules and contextno_allowed_resourcepayments.agent 1.0.0827bf4c8…11df
In your traces
Each governed call is an OpenTelemetry span, through providers your application already owns. A refused span says where it stopped, and that it had no effect.
Both views come from the same scripted run as the rest of this page, captured in memory. Timings are from a local run with stand-ins, not performance figures.
Get notified when Maatrel is released.
Leave your email and we will write when it is available.
At release
An email when Maatrel is available.
Before that
We are inviting a few teams at a time to a private preview: a GitHub repository with the library, the guides and the examples. We may invite you.