capsule
Case studies

Lead enrichment

Scores each new sales lead from 0 to 100 from its website, writes a one-line summary, and tells the sales rep once.

  • Written by our coding agent
  • 4 automatic tests

The job

Read the new leads, open each company’s website, score how well it fits a buyer of developer tools, summarize what it does in one sentence, then send the sales rep one message.

The agent gets this description and the permissions on the right, nothing else. It never sees how “done” will be checked.

Can do on its own
Read the leads
Open the two companies’ websites
Add a row with a score from 0 to 100
Waits for your OK, every time
Send the sales rep a message
Refused outright
Open any other page
Write a score outside 0 to 100
Anything else
Counts as done when capsule confirms
Both leads have a row
Every score is from 0 to 100
One message was sent

How it was built

  1. Written from a task description

    One attempt, kept unchanged. The checker leaned yes but wasn’t sure enough, so a person decided.

  2. Review

    The first version limited the score as if it were money, which would have let a third lead slip through. The score became a plain number range.

Automatic tests

Each test plays out a run with the agent's moves written in advance, instead of asking an AI model, and checks that capsule allows, stops or refuses each one correctly. They run every time capsule's code changes.

Normal run
Stackforge scores 92 and Hearth Bakery 3, each with a summary. The message to the rep waits for you and is allowed.
Unlisted page
A pricing page that isn’t on the list is refused. The agent carries on and finishes.
Bad score
A score of 150 is refused outright. Nobody is asked.
Message denied
The message to the rep waits for you and is denied. Nothing is sent.

What’s proved

These hold for every run, not just the ones above. Each is proved, and the proof is checked by computer.

  • A number outside its allowed range is never accepted.
  • Combining permissions never turns “ask me first” into “go ahead.”

What this doesn’t show

  • This job hasn’t had a real run with an AI model yet. Its tests run on every code change.