AI agents that only do what you allow.
Write down what it's for and what it can touch. It asks before anything risky, and something other than the agent checks that it finished.
A simulation of the invoice agent below, played in your browser. The Acme receipt is over $300, so it waits for you. Allow it or deny it to see both outcomes.
Five promises capsule keeps, so the agent doesn't have to.
Most agent frameworks trust the AI model to behave. capsule doesn't. The model only suggests the next step. capsule decides whether it's allowed, logs it, and checks the result.
The examples use the invoice agent shown further down.
It can't do what you didn't allow.
Every action is checked against the permissions you wrote, before it happens. A tool you didn't allow isn't even offered. A receipt, an email or a model's reply is just text, and text can't give the agent new permissions.
Risky steps wait for your OK.
Mark a permission "ask me first" and every matching action stops until you approve or deny it. You don't have to be watching: the run waits, and you can answer later.
Every step is logged, and replays.
Each model call, tool call and answer is saved. Reopen a run and it ends up in exactly the same place, without calling the model or any tool again. A saved run doubles as a test you can replay any time, with no model and no API key.
"Done" is checked, not claimed.
A separate checker decides whether the job is finished, using evidence capsule gathers itself: rows present, files changed, tests re-run. It never takes the agent's word for it, and when it isn't sure, it asks you.
Its choices become training data.
A small, inexpensive model picks the agent's next step, but only from what its permissions allow. Every pick is logged with the options it had, so using the agent builds a labeled dataset as you go.
One small file in. A working agent out.
You describe the job in one form: what it's for, what it can touch, how much, what waits for you, and what "done" means. capsule mint turns it into a capsule, an agent whose permissions are exactly what you wrote.
; Post each receipt to the ledger. (agent invoice-ledger (purpose "post each receipt in receipts/ to the ledger, one row per receipt with its vendor and total") (capability receipts :files "receipts") (capability ledger :table "ledger.jsonl" :columns (vendor total) :insert "Insert one row. total is an integer number of cents.") (capability mail :spool "mail") (grant capability receipts/list :kind tool) (grant capability receipts/read :kind tool) (grant capability ledger/read :kind tool) ; Up to 300.00 is posted; up to 5000.00 waits for you; ; over that is refused; 1000.00 in all. (grant capability ledger/insert :kind spend :scope "..30000" :amount 1 :charges spend) (grant capability ledger/insert :kind spend :scope "..500000" :amount 1 :charges spend :ask) ; Every send waits for you. (grant capability mail/send :kind tool :ask) (budget :spend 100000) (done (rows ledger 2) (covers ledger vendor "Acme" "Globex") (empty mail)))
invoice-ledger
receipts/Turning the file into an agent is itself a logged run you can reopen.
How the limits hold.
Most agent "guardrails" are a line in a prompt, or a check someone remembered to write. In capsule, every action goes through the same check before it happens, and that check doesn't depend on the AI model behaving.
Permissions come from what you write.
You list what the agent may do, or state facts and rules that work it out, such as which team a task belongs to. Each permission keeps a note of where it came from, and capsule re-checks that note before using it.
Text can't change permissions.
An email, a web page or a model's reply is only ever data. Only sources you've approved can add facts, and withdrawing a fact removes every permission built on it, for every helper, at its next step.
For a real case, see the price monitor and the tampered page.
Every action traces back to a person.
Each logged step points to the steps it depended on, back to the permissions someone wrote. Anyone with the log can follow that chain and check nothing was altered, in their browser.
A small core does the checking.
The check rests on a small amount of simple, predictable code. The AI models and your agents sit outside it, and everything they produce is checked. Its rules are proved, not just tested.
Helpers only ever get less.
An agent can hand part of a job to a helper. The helper gets only what it asked for and what the agent itself can do. Anything beyond that is dropped, and the drop is logged. Budgets are split, never copied: two helpers with $500 each cost $1,000.
The permission rules are proved, not just tested.
The part of capsule that decides what an agent may do is modeled in Lean 4, a proof assistant. Each rule below is proved, and the proof is checked by computer. A test tries some cases. A proof covers every one.
- 177
- proofs, each checked by computer
- 0
- left unfinished or assumed
- 13
- gaps the proofs and a review found in the code, now fixed
A helper never gets more than the agent that started it.
Whatever it asks for, it can only do what every agent above it can do.
Borrowed code can’t bring its own permissions.
Code an agent is handed runs with that agent’s permissions, never those of whoever wrote it.
Taking a permission away reaches every helper.
Each one loses it at its next step, however deep it sits.
Combining permissions never turns “ask me first” into “go ahead.”
A step that waits for you at any level still waits.
Budgets are never created out of nothing.
Splitting a budget between helpers never makes more of it.
Each row’s limit is its own.
A price written for one product is held to that product’s limits, never another’s.
“Done” follows from the evidence.
What the checker accepts follows from the facts it was shown and the rules it was given, nothing else.
The proofs cover a precise description of how capsule checks permissions. The code itself is checked against that description with tests and review, not by proof.
One week, one model. The agent got cheaper anyway.
We pointed a coding agent built on capsule at its own codebase. It picked up tasks and opened pull requests, and after each run a second process read the log and improved the tools around it. The model stayed the same all week. Only the tools changed.
We call this compiling an agent: its own runs are the training data, and the goal is written down.
Model calls per run
Times it reached for a tool it didn't have
View as a table
| Run | Model calls | Reached for a missing tool |
|---|---|---|
| C2 (early) | 41 (816k tokens) | 11 |
| C5 | 7 | 0 |
| C11 | 15 | 1 |
| C12 | 9 | 1 |
| C10 | 9 | 1 |
pull requests merged with the agent's changes untouched.
Most of the savings came from tools giving clearer answers:
- numbered lines when reading files
- a cap on whole-file reads
- mistakes returned as messages the model can read, instead of crashes
- listing the tools on offer at every turn
The invoice agent's first live run posted one receipt out of two. The done-check caught it. After one fix to a tool description, both posted and every check passed.
The agent was wrong about itself.
At the end of each run, the agent wrote a short note about what slowed it down. Once, that note said certain tools had worked fine. The agent had never been given those tools.
It wasn't lying. It just didn't know. That's the trouble with letting an agent grade its own work: its account of what happened can be confidently wrong.
That's why the checker never reads the agent's account. It re-runs the checks itself and looks at what actually changed.
Said tools had worked that it was never offered.
The checker's first real verdict: done, at 0.92 confidence. It re-ran the checks itself and found no files changed outside the task.
On three tasks where checking meant reading code, it declined to decide and asked a person. That's what it's designed to do.
Five real jobs, built and on the record.
Our coding agent wrote four of them from a short task description, and we kept what it wrote unchanged. We wrote the fifth by hand, as the example it learns from. Each has a case study: what the agent may do, how it was built, and how it's tested.
Every job comes with automatic tests, 26 so far. Each one plays out a run with the agent's moves written in advance, so we can check that capsule allows, stops and refuses the right things, the same way every time.
The price-monitoring job ran live on three Claude models against pages with planted instructions. Read what happened, step by step, from its record.
The workbench · coming
Run capsule up to open a local app. Fill in a form, see the agent's source and exactly what it's allowed to do, then run it. You approve paused steps inline, watch the timeline, and press Improve to run the compile loop.
Live inspector · coming
Step through every job's runs, one logged step at a time, in your browser.
Apply as a design partner.
capsule is in private beta. Tell us about the job you'd hand an agent if you could trust it. If it's a fit, you'll get a 20-minute call with the engineer who built capsule.
Design partners use the workbench, so no Rust is needed. Logs stay on your side, and you can bring your own model provider or a local model.