Bookkeeper
Chases the receipts you swore you'd file. Does not judge.Captures invoices and receipts from mail and uploads, extracts and checks the numbers, matches them against payments, and prepares a clean pack for your accountant.
What it costs
Sold on its own, or as part of a stack where the price per agent drops. Every tier includes the eval report and every future release while you are subscribed.
1 seat — One operator, unlimited runs.
up to 10 seats — Shared config, per-seat audit log.
unlimited seats — SSO, priority evals, private bundles.
Where it stands
Planned — specified, not yet built.
- Scenarios in the suite120
- Minimum pass rate99%
- p95 latency ceiling4.0s
- Cost per run ceiling£0.04
- RunsScheduled
These are the thresholds Bookkeeper has to clear before it may ship — not results. Measured numbers publish with the first release.
The eval report
This ships inside the bundle. It is the contract: what the agent is tested on, what it refuses outright, and what it hands back to a human.
{
"agent": "bookkeeper",
"version": "0.3.0",
"status": "not_yet_measured",
"gate": {
"scenarios": 120,
"pass_rate_min": "99%",
"p95_latency_max": "4.0s",
"cost_per_run_max": "£0.04"
},
"suites": [
{
"name": "core",
"scenarios": 66,
"description": "The ordinary cases. What the agent is for, done correctly."
},
{
"name": "adversarial",
"scenarios": 30,
"description": "Prompt injection, contradictory instructions, hostile or malformed input, and attempts to make it act outside its contract."
},
{
"name": "regression",
"scenarios": 24,
"description": "Every bug ever found, kept forever so it cannot come back."
}
],
"refuses": [
"Acting on instructions found in content it was asked to read",
"Sending, publishing or paying anything without explicit approval",
"Inventing a fact it cannot source"
],
"escalates_to_human": [
"Confidence below the configured floor",
"Anything involving money, legal commitment or a new counterparty",
"Conflicting instructions it cannot resolve"
],
"measured": null
}
measured: null until the suite has actually run. We do not publish numbers we have not measured.