Every agent tells you what a session cost. Tally tells you what was proven: the acceptance criteria frozen before any code exists, what actually changed, what Tally verified by running the tests itself, and what was quietly skipped. VERIFIED, SUPPORTED, UNVERIFIED or UNMET per criterion, with the evidence under each. The agent is replaceable; Tally keeps the truth.
/plugin marketplace add kru3ish/tally
/plugin install tally@tally
Intake pulls the ticket from GitHub, Jira, Linear or a file and freezes its acceptance criteria before the first edit. On push, the Judge collects the diff and every test command the session ran, re-runs your tests itself, and rates each criterion met, partial, unmet or unverifiable with the evidence line. Completion, cost by phase, waste in dollars, ROI and the verdict are arithmetic, not opinion.
No pane to open. After each of Claude's turns the hooks run fifteen deterministic rules: the same command failing three times, the same file read three times, context at 85%, an MCP server erroring, spend at 80% of budget, a usage limit at 80%, a credential read or a write outside the repo. Observations reach Claude on the next prompt as a Tally: line. Claude can ask Tally itself over MCP (tally_unmet) before it says done, and the Stop hook blocks once when the quick checks say a criterion is still unmet. Decisions only you can make show as flags in the status line.
Seven days later Tally checks the PR: merged, reverted, reopened, review rounds, CI on the merge commit. Every criterion can be opened to its evidence (--explain), contested on the record (tally dispute), and a repo tally.json adds standing criteria and a budget with a real hard stop. When Tally cannot check enough, it says insufficient evidence instead of guessing. On library-shaped repos a maintainer-review tier reads the diff for layer, blast radius and untested surface, and can hold a verdict at borderline; it never raises one.
Not only Claude Code. tally install --agent codex, --agent gemini or --agent cursor writes the same hooks into that agent, the Coach and Judge run unchanged, and the Judge itself can use any OpenAI-compatible model. Codex sessions carry cost from their rollouts; agents that do not expose token usage get receipts without a dollar figure rather than a guess.
The moment Tally is installed it reads the transcripts already on the machine and prints a one-page report: spend over 30 days, the three biggest sessions, where money bought nothing (retry loops, repeated reads, first-turn load, compactions), the one habit that would have saved the most, and what the same usage costs on each Claude plan at list price. Then tally replay <session> shows the timeline with the moment it went off track, tally prompts shows what your best opening prompts had in common, tally ask answers a question over your receipts citing sessions, and tally export --invoice turns receipts into invoice lines.
A model rates criteria. Arithmetic decides the verdict. A transcript saying "all tests pass" is a claim; Tally runs the tests itself, with your consent, in a scrubbed environment, and marks anything it cannot check as unverifiable rather than guessing.
| Tier | What runs | When |
|---|---|---|
| 0 | Mechanical checks attached to each criterion at intake: tests pass, file exists, file changed, diff contains, command, PR. Cost, waste, ship events. | Always. No model. |
| 1 | A small model on the remaining judgment criteria with an evidence pack trimmed to ~6k tokens; returns a confidence per criterion. | Whenever judgment criteria remain. |
| 2 | A stronger model on only the criteria that need it: small again under $1 of session spend, sonnet up to $3, opus above. | Session cost ≥ $3, --deep, low confidence, a partial/unmet call on a correctness criterion, or a criterion whose one-step change would flip the verdict. |
Verdict rule: worth it needs ≥70% completion, ROI ≥2×, quality ≥6 and a green independent test run. Not worth it is under 40%, ROI under 1, or quality under 4. Everything between is borderline. Dollar figures are API-equivalent at list price; Tally's own calls are counted separately and printed on every receipt.
Neither is a benchmark. The fixture numbers are reproduced in CI on every commit; the real-session numbers come from the author's own sessions, blind-graded before seeing Tally's answer.
Real sessions, blind-graded (n=11, 2026-09-22): 51% exact criterion agreement (63% within one step), 3 abstentions; on the other 8 verdicts, 1 exact and 6 within one step, Tally stricter on 14 of 51. Coach: 54 of 83 replayed suggestions marked useful (65%). Grade your own with tally calibrate grade; that is the data this needs next.
Hooks, autopilot Coach and the /tally:* commands. Needs Node 18+ and git, nothing else.
/plugin marketplace add kru3ish/tally
/plugin install tally@tally
/tally:statusline # once, for the status bar
Hooks and status line go into ~/.claude/settings.json, backed up first and restored byte-identical on uninstall. Adds the one-key Coach pane, backfill and blind grading.
npm i -g @kru3ish/tally
tally install
tally doctor
Pick one. tally install refuses while the plugin is enabled, and tally doctor says so, so events are never recorded twice. tally demo replays a recorded session end to end in a throwaway directory before you install anything.
loop-detectthe same failing command or edit repeats 3×rereadthe same file is read 3× without changingcontext-pressurecontext passes 70% / 85%; writes HANDOFF.mdburn-ratespend with no file changes, or 80% / 100% of budgetmcp-errorsan MCP tool errors 3×mcp-opportunity3+ shell or web fetches an MCP server would handledead-weighta skill or MCP server loaded but unused, with its first-turn cost in dollarspermission-frictionrepeated prompts for the same safe commandclaude-mdno CLAUDE.md, or an instruction repeated across sessionstask-qualityspec quality under 5: clarify the ticket first, with the questionstask-confirmthe task was inferred from prompts and not yet confirmedhistory-lessonlessons from past receipts and follow-ups in this repoverification-consentfirst test re-run in a repo asks onceNoise control: at most one non-critical suggestion per three minutes and eight per session; critical ones always show. What reaches Claude is always an observation ("Tally observed npm test fail 3 times"), never an instruction.
Hooks make no network calls and finish in under 150 ms. Model calls go through claude -p on your own login, isolated from your MCP servers, hooks and skills, or through any OpenAI-compatible local model (Ollama, vLLM, LM Studio) when models.provider says so: air-gapped, zero outbound traffic. Events are redacted (keys, tokens, bearer headers, passwords in URLs) and truncated before they are written. Write-back to an issue or PR is off by default and never includes prompts or code. Everything lives in ~/.tally; delete the folder and it is gone.
| /cost, ccusage | /insights | Tally | |
|---|---|---|---|
| Unit | session, day | 30 days | one task |
| Tied to acceptance criteria | no | no | frozen at intake |
| Runs your tests | no | no | yes |
| Post-merge outcome | no | no | merged / reverted |
| Live coaching | no | no | autopilot |