For Claude Code, Codex CLI, Gemini CLI and Cursor · local · open source

The reliability layer for AI coding agents.

Every agent tells you what a session cost. Tally tells you what was proven: the acceptance criteria frozen before any code exists, what actually changed, what Tally verified by running the tests itself, and what was quietly skipped. VERIFIED, SUPPORTED, UNVERIFIED or UNMET per criterion, with the evidence under each. The agent is replaceable; Tally keeps the truth.

Claude Code plugin
/plugin marketplace add kru3ish/tally
/plugin install tally@tally
Claude Code status line, always on
Tally · Rate limit the login endpoint · $1.33/$56.25 (2%) · ctx 34% · 2 coach flags (/tally:coach) · receipt: 67% · borderline
TALLY RECEIPTdemo session · 2026-09-10
Rate limit the login endpoint
github.com/acme/app/issues/42 · criteria frozen 14:01
JUDGED
BORDERLINE66.7% complete (1 unverifiable)quality 7/10ROI 84×
  • ✔c1 POST /api/login returns 429 after 5 failed attemptstier 2 · 1.00
  • ✔c2 A test covers the 429 pathtier 2 · 1.00
  • ✘c3 README documents the limittier 1 · 1.00
  • ?c4 Existing login behaviour unchanged below the limittier 1 · 1.00
Verificationnpm test → passed
Cost (API-equivalent)$1.33 of $56.25 budget
Per met criterion$0.660
Waste$0.629 (3 identical failing runs)
Tally's own spend$0.070 · 5.3%
Value credited$113 of $225
Two of the three verifiable criteria are met with an independent green test run, one unmet, one unverifiable. Spend is well inside budget, but the README item was skipped and stated as skipped, and $0.60 went into three identical failing test runs.
Next time
  1. When npm test fails twice with the same assertion, read the test before editing again.
  2. Read the ticket checklist before saying done.
  3. Stop calling the Jira MCP after the first 401.
Follow-up, 8 days later: PR #57 merged, then reverted in dba0f121 → NOT WORTH IT
01 · THE RECEIPT

Criteria in, verdict out.

Intake pulls the ticket from GitHub, Jira, Linear or a file and freezes its acceptance criteria before the first edit. On push, the Judge collects the diff and every test command the session ran, re-runs your tests itself, and rates each criterion met, partial, unmet or unverifiable with the evidence line. Completion, cost by phase, waste in dollars, ROI and the verdict are arithmetic, not opinion.

02 · THE COACH, ON ITS OWN

Runs after every turn.

No pane to open. After each of Claude's turns the hooks run fifteen deterministic rules: the same command failing three times, the same file read three times, context at 85%, an MCP server erroring, spend at 80% of budget, a usage limit at 80%, a credential read or a write outside the repo. Observations reach Claude on the next prompt as a Tally: line. Claude can ask Tally itself over MCP (tally_unmet) before it says done, and the Stop hook blocks once when the quick checks say a criterion is still unmet. Decisions only you can make show as flags in the status line.

03 · THE FOLLOW-UP, AND THE RECORD

Did it hold up? Audit it. Dispute it.

Seven days later Tally checks the PR: merged, reverted, reopened, review rounds, CI on the merge commit. Every criterion can be opened to its evidence (--explain), contested on the record (tally dispute), and a repo tally.json adds standing criteria and a budget with a real hard stop. When Tally cannot check enough, it says insufficient evidence instead of guessing. On library-shaped repos a maintainer-review tier reads the diff for layer, blast radius and untested surface, and can hold a verdict at borderline; it never raises one.

For one person

Your numbers on day one, no model calls.

Not only Claude Code. tally install --agent codex, --agent gemini or --agent cursor writes the same hooks into that agent, the Coach and Judge run unchanged, and the Judge itself can use any OpenAI-compatible model. Codex sessions carry cost from their rollouts; agents that do not expose token usage get receipts without a dollar figure rather than a guess.

The moment Tally is installed it reads the transcripts already on the machine and prints a one-page report: spend over 30 days, the three biggest sessions, where money bought nothing (retry loops, repeated reads, first-turn load, compactions), the one habit that would have saved the most, and what the same usage costs on each Claude plan at list price. Then tally replay <session> shows the timeline with the moment it went off track, tally prompts shows what your best opening prompts had in common, tally ask answers a question over your receipts citing sessions, and tally export --invoice turns receipts into invoice lines.

How it judges

Tiered, so most receipts cost almost nothing.

A model rates criteria. Arithmetic decides the verdict. A transcript saying "all tests pass" is a claim; Tally runs the tests itself, with your consent, in a scrubbed environment, and marks anything it cannot check as unverifiable rather than guessing.

TierWhat runsWhen
0Mechanical checks attached to each criterion at intake: tests pass, file exists, file changed, diff contains, command, PR. Cost, waste, ship events.Always. No model.
1A small model on the remaining judgment criteria with an evidence pack trimmed to ~6k tokens; returns a confidence per criterion.Whenever judgment criteria remain.
2A stronger model on only the criteria that need it: small again under $1 of session spend, sonnet up to $3, opus above.Session cost ≥ $3, --deep, low confidence, a partial/unmet call on a correctness criterion, or a criterion whose one-step change would flip the verdict.

Verdict rule: worth it needs ≥70% completion, ROI ≥2×, quality ≥6 and a green independent test run. Not worth it is under 40%, ROI under 1, or quality under 4. Everything between is borderline. Dollar figures are API-equivalent at list price; Tally's own calls are counted separately and printed on every receipt.

Accuracy

Early, and measured two ways.

Neither is a benchmark. The fixture numbers are reproduced in CI on every commit; the real-session numbers come from the author's own sessions, blind-graded before seeing Tally's answer.

19/ 19
criteria agreed, 5 fixture sessions
5/ 5
verdicts agreed, fixtures
4.1%
Tally's own spend, share of session spend
51%
criterion agreement on 11 real sessions, 51 criteria (early, n=11)

Real sessions, blind-graded (n=11, 2026-09-22): 51% exact criterion agreement (63% within one step), 3 abstentions; on the other 8 verdicts, 1 exact and 6 within one step, Tally stricter on 14 of 51. Coach: 54 of 83 replayed suggestions marked useful (65%). Grade your own with tally calibrate grade; that is the data this needs next.

Install

Two lines, either way.

As a Claude Code plugin

Hooks, autopilot Coach and the /tally:* commands. Needs Node 18+ and git, nothing else.

/plugin marketplace add kru3ish/tally
/plugin install tally@tally
/tally:statusline        # once, for the status bar

As an npm CLI

Hooks and status line go into ~/.claude/settings.json, backed up first and restored byte-identical on uninstall. Adds the one-key Coach pane, backfill and blind grading.

npm i -g @kru3ish/tally
tally install
tally doctor

Pick one. tally install refuses while the plugin is enabled, and tally doctor says so, so events are never recorded twice. tally demo replays a recorded session end to end in a throwaway directory before you install anything.

The rules

What the Coach watches for.

loop-detectthe same failing command or edit repeats 3×
rereadthe same file is read 3× without changing
context-pressurecontext passes 70% / 85%; writes HANDOFF.md
burn-ratespend with no file changes, or 80% / 100% of budget
mcp-errorsan MCP tool errors 3×
mcp-opportunity3+ shell or web fetches an MCP server would handle
dead-weighta skill or MCP server loaded but unused, with its first-turn cost in dollars
permission-frictionrepeated prompts for the same safe command
claude-mdno CLAUDE.md, or an instruction repeated across sessions
task-qualityspec quality under 5: clarify the ticket first, with the questions
task-confirmthe task was inferred from prompts and not yet confirmed
history-lessonlessons from past receipts and follow-ups in this repo
verification-consentfirst test re-run in a repo asks once

Noise control: at most one non-critical suggestion per three minutes and eight per session; critical ones always show. What reaches Claude is always an observation ("Tally observed npm test fail 3 times"), never an instruction.

Privacy

Nothing leaves your machine.

Hooks make no network calls and finish in under 150 ms. Model calls go through claude -p on your own login, isolated from your MCP servers, hooks and skills, or through any OpenAI-compatible local model (Ollama, vLLM, LM Studio) when models.provider says so: air-gapped, zero outbound traffic. Events are redacted (keys, tokens, bearer headers, passwords in URLs) and truncated before they are written. Write-back to an issue or PR is off by default and never includes prompts or code. Everything lives in ~/.tally; delete the folder and it is gone.

Next to the built-ins
/cost, ccusage/insightsTally
Unitsession, day30 daysone task
Tied to acceptance criterianonofrozen at intake
Runs your testsnonoyes
Post-merge outcomenonomerged / reverted
Live coachingnonoautopilot