← Back to the token-goat README
CLI
Every command accepts a global --cwd <path>, which runs it as if invoked from that directory. It exists so a caller can name a project root without making that root its own working directory — a launcher should never resolve a binary name against a directory the workspace controls. It is applied before anything resolves the project root or loads config, so --cwd selects which .token-goat.toml applies. The one setting it does not select is whether index embeds a file, which follows that file’s own project (see index below).
Archive/document comparison workflow
These are agent-selected primitives, not a manual checklist. Give the agent the file and the question. The installed routing guide and read hook select the matching format flow; the commands below show the steps it can take without loading whole files:
token-goat sqlite-schema catalog.db
token-goat sqlite-query catalog.db "SELECT file_path, name FROM files WHERE name LIKE '%owner%' LIMIT 20" --json
token-goat xlsx-sheets link-map.xlsx
token-goat xlsx-query link-map.xlsx --sheet Links --columns publication,source,target --head 50
token-goat pdf-meta manual.pdf
token-goat pdf-outline manual.pdf
token-goat pdf-locate manual.pdf "torque spec" --ignore-case
token-goat pdf-extract manual.pdf --pages 12-15 --layout --head 120
token-goat intentionally does not render PDF pages or infer XML publication lineage: those operations produce binary/visual output or require schema-specific interpretation. Keep those steps in the document/PDF tooling, then pass only the bounded paths, rows, and page text needed for comparison.
| Command | What it does |
|———|————-|
| token-goat symbol [name] | Jump to a symbol definition. -p, --project [path] scopes the search to one project root instead of the default global (cross-project) index — pass no value to use the current directory’s project root, or a path to scope to a different one. Without --project, definitions from the current directory’s project are listed first, and a truncated page says how many of the matches are local. --json’s filePath renders root-relative when a project root resolves, absolute when none does — matching human output and the outline/skeleton/refs --json convention. --grep <pattern> searches project-wide by NAME PATTERN instead of an exact name (regex, falling back to a literal substring match when the pattern is not valid regex) — the positional name is omitted in that mode, and the two are mutually exclusive since regex-filtering an already-exact name can only match everything or nothing. This is the only project-wide symbol-name pattern search: skeleton/outline/exports --grep are per-file, types --grep covers only type-like kinds, and dead --grep only zero-reference symbols. The filter is applied before the --limit slice, so --limit N --grep P returns up to N matching rows; when it matches nothing among symbols that are in scope, the output names the active filter instead of reading as an empty project. --exclude-tests hides symbols DEFINED in a test file (opt-in — omitted, output is unchanged), the same definition-site sense dead --exclude-tests uses rather than the call-site sense of refs/callers. The high-value case is a common helper name that is mostly defined in tests: against this repo’s own index symbol run returns 18 rows of which 14 are test-file definitions, and symbol capture returns 9 of 9. Composable with --grep (a symbol must satisfy both) and applied before the --limit slice, so the flag selects from the whole match set rather than an already-capped page — without that ordering a --limit N window filled by test-file rows would report nothing for a symbol that is plainly indexed in src. --grep, --exclude-tests and --exclude-vendored therefore read every symbol in scope, but only its name and path, plus the body of each row that prints. That takes about two seconds on a 546,000-symbol project. When it hides every match there was, the output names how many were hidden and exits 0, instead of the exit-1 No matches a genuinely unindexed name returns. --stats adds a per-result reference count and doc-coverage flag, computed live from the index — the same flag read/skeleton/outline already carry, useful here for picking which of several same-named candidates is the real one. Note its known limitation: the count is keyed by symbol NAME project-wide, not by definition site, so under --grep several same-named symbols in different files all show the identical count. A name with no match exits 1 with No matches for '<name>' and then says why: an exact name that --kind/--file filtered out is named with its location, from one indexed lookup; otherwise the project’s closest indexed names follow as Did you mean:, or a Try: token-goat semantic pointer when none is close, plus a pointer when the name is a nested JSON/YAML key. The suggestions rank every distinct symbol name in the project, read with one names-only query and no symbol bodies, so a miss costs about a second on a 546,000-symbol project rather than the minutes a full-row walk used to take. A project with more than 500,000 distinct names skips the ranking and prints Near-name suggestions skipped: ... in its place, rather than ranking a subset and presenting it as the nearest. Under --json a miss prints the same {items, truncated, totalCount} envelope a hit does, with items: [] and totalCount: 0, and exits 0 like the command’s other empty results; none of the text-mode diagnosis runs for it. |
| token-goat read "file::symbol" | Pull one function or class, not the whole file. Supports qualified lookups (read "file.py::Class.method") and line ranges: read "file.py@10-40" for lines 10 to 40 inclusive, or read "file.py@42" for one line. Line ranges read straight from disk, so they work on any file, including paths outside an indexed project. A trailing @LINE on the symbol itself (read "file.py::run@42", or combined with a qualifier as read "file.py::Class.method@42") anchors an ambiguous spec to the one candidate starting on that exact line — for a top-level definition with no enclosing Class.method qualifier, this is the only way to pick it out when its bare name also matches something else in the same file; every ambiguity error’s retry suggestions already use this form where a plain qualifier wouldn’t be unique. Pass a comma-separated spec (file::a,b) to merge several symbols’ bodies from one file into a single call, each headed by its symbol name. Segments may also carry their own file (a.ts::x,b.ts::y) to merge symbols across several files in one call; a bare segment inherits the file to its left (a.ts::x,b.ts::y,z reads z from b.ts), and once more than one file is involved each block is headed by the full file::symbol so two files contributing the same symbol name stay distinct. --force-refresh reparses the file from disk and updates the index before querying — for files touched by git operations, external tools, or direct filesystem writes that bypass the normal post-edit indexing hook. --stats adds a per-symbol reference count and doc-coverage flag, computed live from the index. |
| token-goat replace <file> | Replace one string in a file using --old-from/--new-from or --old-b64/--new-b64; --all replaces every match. If the exact match fails but a unique match exists once CRLF/LF differences are ignored, it heals automatically, writing the replacement back in the file’s line-ending convention at that location. --normalize-newlines converts the old/new text’s CRLF/LF to match the target file’s dominant line ending before matching, for forcing normalization proactively. |
| token-goat insert-section <file> --after <heading> | Insert content immediately after a matched section (--content-from <source> or --content-b64 <payload>), resolved the same way section resolves headings (exact, normalized, or an unambiguous prefix) — avoids the stale byte-exact anchor replace would otherwise need for an append-to-a-running-log edit. |
| token-goat note-add <file> [--symbol NAME] | Attach a free-text architecture/rationale note (Markdown, --content-from <source> or --content-b64 <payload>) to a file, or to one specific indexed symbol within it. Captures a fingerprint of what the note describes (the symbol’s current body, or a digest of the file’s current top-level symbol manifest) so staleness can be detected later — re-running note-add for the same file/symbol overwrites rather than duplicates. |
| token-goat note-get <file> [--symbol NAME] | Read back the note attached to a file or one indexed symbol within it. Flags whether the note has gone stale (the underlying code changed since it was written) via a stale field under --json. |
| token-goat note-list [--stale-only] | List every recorded architecture note. --stale-only shows just the notes whose fingerprint no longer matches the current index — i.e. the file/symbol they describe changed since the note was written. Staleness is purely advisory: nothing here auto-rewrites or deletes a note. |
| token-goat write-file <dest> | Write exact bytes to a file, sidestepping shell-escaping trouble with backticks, quotes, $vars, and CRLF. --from <source> copies bytes from a source file; --b64 <payload> decodes a base64 payload; with neither, reads from stdin. |
| token-goat section "doc.md::Heading" | Pull one Markdown section by heading. Fully supports ATX (#/##) and Setext-style (Heading\n=== level 1 and Heading\n--- level 2) headings. A miss that is an unambiguous prefix of exactly one heading, or a distinctive suffix/word-subset of exactly one heading (e.g. Setup → “Installation and Setup”, Config Options → “Configuration Options”), auto-redirects with a (redirected from: …) marker (and a redirectedFrom field under --json); a query matching 2+ headings is never guessed and reports a miss instead. A genuine miss lists only headings similar to the query as “Did you mean” suggestions, not every heading in the file. Disambiguate duplicates with "doc.md::Heading#2". Comma-separated "doc.md::A,B" fetches several sections from one file in a single call, mirroring read’s file::a,b multi-symbol grammar, and automatically subsumes nested child sections ((already included in section 'Parent', lines X-Y)) to prevent duplicate tokens. Cross-file "a.md::Heading1,b.md::Heading2" fetches sections from several files in one call, mirroring read’s a.ts::x,b.ts::y cross-file grammar — a bare heading after a file::Heading segment inherits the previous file, and each section is keyed by its full file::Heading pair so two files sharing a heading name cannot overwrite each other. token-goat section doc.md --list lists every heading in the file instead of reading one; --grep <pattern> narrows that list to headings matching a regex (falls back to a literal substring match if the pattern doesn’t compile), same convention as outline/types/exports’s own --grep. Without --list, --grep filters the section body instead: only the matching lines print, each under the nearest sub-heading above it, followed by N of M lines matched. A pattern that matches nothing says so and exits 0. YAML front matter at the top of a file is skipped, so its closing --- is never read as a heading underline. |
| token-goat skill-section "<name>::<heading>" | Extract a named section from an installed skill without reading the full skill file. |
| token-goat skeleton "file" | Show all signatures in a file without bodies — typically 70–90% fewer tokens than a full read. --force-refresh reparses from disk first, bypassing a stale index. --stats adds a per-symbol reference count and doc-coverage flag, computed live from the index. --grep <pattern> narrows to symbols whose name matches a regex (a literal substring when the pattern is not valid regex), which is how you skim one area of a large file without dumping its whole symbol list; --min-lines <n> drops symbols shorter than N lines. Both compose, and if a filter removes everything the output says so and names the filter, rather than looking like a file with no symbols. Accepts a comma-separated file list ("a,b,c") to cover several files in one call, one clearly-headed block per file; extra space-separated file arguments are reported in a note naming that comma form instead of being silently dropped. With --json, a comma-separated list returns one merged document (rows carry their own filePath), not one document per file. |
| token-goat outline "file" | List top-level symbols with line ranges and docstring hints — one-glance file map. Doc hints are clipped to about a sentence with a visible ellipsis; the full doc comment is one read "file::symbol" away (--json carries it whole). --force-refresh reparses from disk first, bypassing a stale index. --stats adds a per-symbol reference count and doc-coverage flag, computed live from the index. --grep <pattern> narrows to symbols whose name matches a regex (a literal substring when the pattern is not valid regex), which is how you skim one area of a large file without dumping its whole symbol list; --min-lines <n> drops symbols shorter than N lines. Both compose, and if a filter removes everything the output says so and names the filter, rather than looking like a file with no symbols. Accepts a comma-separated file list ("a,b,c") to cover several files in one call, one clearly-headed block per file; extra space-separated file arguments are reported in a note naming that comma form instead of being silently dropped. With --json, a comma-separated list returns one merged document (rows carry their own filePath), not one document per file. |
| token-goat yaml-outline <file> | Structural summary of a YAML document (array shape / object key types) instead of a raw Read. Multi-document streams (----separated) outline as an array of documents. --filter TEXT keeps only the top-level keys containing the text and reports how many matched. |
| token-goat yaml-query <file> <path> | Extract one value or a projected/filtered subset from a YAML document by dot-path instead of a raw Read (same grammar as json-query: [n] index, [*] wildcard, [field=value] filter — e.g. items[status=active].name). --head <n> caps a projected/filtered result. |
| token-goat xml-outline <file> | Structural summary of an XML document (element tag hierarchy, attribute keys, child counts) instead of a raw Read. Supports --depth <n> (alias for --max-depth) and --json. |
| token-goat xml-query <file> [path] | Extract one value, element text/XML, or a projected/filtered subset from an XML document by dot-path or --xpath <expr> instead of a raw Read (dot-path grammar matches json-query: element tags, @attr, [n], [*], [attr=value]). Supports --xpath <expression> with namespace and attribute predicate support, --with-lines to output exact source line ranges ((lines start-end)), --decode-embedded-xml to pretty-print and bound entity-encoded nested AML/XML payloads, --head <n>, and --json. |
| token-goat html-outline <file> | Structural outline of an HTML document (DOM hierarchy, tags, IDs, classes, element counts, depth, landmarks, tables, forms) instead of reading thousands of lines of markup. Supports --json. |
| token-goat html-query <file> <selector> | Extract matching HTML elements or text using standard CSS selectors (tags, #id, .class, attribute operators [attr], [attr=val], [attr*=val], [attr^=val], [attr$=val], child > and descendant combinators) without a headless browser or heavy DOM dependencies. Supports --text to strip tags, --head <n>, and --json. |
| token-goat html-lint <file> | Fast, zero-dependency HTML5 structural linter checking for unclosed tags, void element violations, duplicate IDs, missing viewport/charset, missing alt attributes, and inline script bloat. Supports --json. |
| token-goat json-outline <file> [--filter TEXT] | Structural summary of a JSON document (array shape / object key types) instead of a raw Read. --filter TEXT keeps only the top-level keys containing the text, for a registry with hundreds of them: json-outline registry.json --filter 11862 prints ten keys and (10 of 93 keys contain "11862"). |
| token-goat json-query <file> <path> | Extract one value or a projected/filtered subset from a JSON document by dot-path instead of a raw Read: dot-separated keys with optional bracket segments — [n] index, [*] wildcard (projects every element/value), [field=value] filter. A quoted bracket segment is one literal key, the only way to name a key holding a dot or a space: "['a.b'].c" reads c under the key a.b, where a.b.c walks three levels. Examples: data.items[3].name, items[*].id, items[status=active]. A trailing projection picks several fields at once: members[*].[id, name] returns one list per element, members[*].{id, city: address.city} one object per element with optional aliases, and a missing field comes back as null. [] iterates like [*], | chains a projection onto a path (members[] \| {id, name}), and a projection on a plain object applies once ({orgName: name, tier}). yaml-query takes the same projections. Pass - as the file to read the document from stdin (gh api repos/o/r \| token-goat json-query - owner.login); with nothing piped in, the command says so instead of waiting. yaml-query and xml-query accept - the same way. |
| token-goat brief "file::symbol" | Bundle a symbol’s body, resolved callers (grouped by enclosing function), and its containing doc section into one round-trip instead of three separate read/callers/section calls. --limit <n> caps the callers shown per symbol (default 20; the true caller count is reported even when truncated). Comma-separated "file::a,b" fetches several symbols’ bundles from one file in a single call, mirroring read’s file::a,b multi-symbol grammar. Cross-file "a.ts::x,b.ts::y" bundles symbols from several files in one call, mirroring read’s cross-file grammar — a bare segment inherits the file to its left, and once more than one file is involved each bundle is keyed by the full file::symbol so two files contributing the same symbol name stay distinct. Also accepts read’s symbol@LINE anchor to pick out an otherwise-ambiguous candidate. -C, --context <n> adds N lines of real call-site source around each entry of the caller block. --json’s symbol.filePath and callers[].file render root-relative when a project root resolves, absolute when none does — matching the plain-text block above. --exclude-tests hides callers whose call site is in a test file, matching refs/callers; the caller count and the elided tail both count the filtered set, so they never disagree with the rows shown, and when the filter empties the block it says so instead of reporting a bare zero that would read as “nothing calls this”. --json adds hiddenByExcludeTests only when the filter actually hid something. --grep <pattern> narrows the caller block to callers whose enclosing symbol name matches this regex (literal substring if it is not valid regex), the same filter refs --grep/call-chain --grep apply to their own results — useful for a high-fanout symbol whose default 20-caller window is otherwise mostly noise; composes with --exclude-tests, and reports hiddenByGrep under --json only when it hid something. |
| token-goat scope "file:line" | Show symbols in scope at a given line — avoids reading the whole file to understand locals. |
| token-goat exports "file" | List public (exported) symbols with types, docstring hints, and line ranges ((lineStart-lineEnd) in text mode, lineStart/lineEnd fields under --json). Names caught only by the source-text scan (no corresponding index row — e.g. certain re-export forms) report no location: omitted from text mode, null under --json. Accepts a comma-separated file list ("a,b,c") to cover several files in one call, one clearly-headed block per file; extra space-separated file arguments are reported in a note naming that comma form instead of being silently dropped. --grep <pattern> only shows exported symbols whose NAME matches this regex (literal substring if it is not valid regex), applied before output is built; when it matches nothing among real exports, the output names the active filter instead of reading like the file has no exports at all. |
| token-goat refs "<name>" | Show all files and line numbers where a symbol is referenced. Pass a comma-separated spec (a,b,c or file::a,b) to merge several symbols’ references into one call, each group headed by its symbol name. Segments may also carry their own file (a.ts::x,b.ts::y) to merge references across several files in one call, mirroring read’s cross-file grammar — a bare segment inherits the file to its left, and once more than one file is involved each block is headed by the full file::symbol so two files contributing the same symbol name stay distinct. --top <n> groups references by file (count only) and shows just the top N by reference count with an elision note, instead of a per-line dump — for high-fanout symbols referenced in hundreds of places. -C, --context <n> shows N lines of real call-site source either side of each hit, rendered exactly like grep -C; omit it (or pass 0) and output is unchanged. Under --json each item gains a contextLines array alongside the existing context field (which names the enclosing symbol, not source text). --exclude-tests hides references whose call site is a test file (opt-in — omitted, output is unchanged); the summary line reports the filtered count plus how many were hidden. --grep <pattern> only shows references whose call-site FILE PATH matches this regex (literal substring if it is not valid regex) — rows render as file:line: symbol, so this is the field each row is keyed on. The pattern is tested against the path exactly as the row renders it, so an anchored --grep "^src/" matches what you see, identically in the single, multi-symbol and cross-file forms. Every form renders a call-site path the same way – root-relative when a project root resolves, absolute when none does, never cwd-dependent – and --json carries that same spelling in filePath (and in --top’s fileCounts[].file), so a payload is reproducible rather than tied to one machine’s drive-letter casing. The high-value case is narrowing a wide-fanout symbol to drop test/vendored hits. Applied before --top’s grouping and before any --limit slice, so it selects from the whole reference set, not an already-capped page; when it matches nothing among references that do exist, the output names the active filter instead of reading like the symbol is unreferenced. A bare name that isn’t indexed at all reports Symbol not found: <name> (with a Did you mean: suggestion when a near-name candidate is indexed) instead of the misleading “no references found”, which is reserved for a real, indexed symbol that genuinely has zero references. |
| token-goat callers <symbol> | Show which functions call a given symbol, grouped by caller with file, caller name, and every invoking line. Complements refs, which shows raw reference sites without grouping by enclosing function. Accepts file::symbol to disambiguate WHICH same-named definition is meant when several files define a symbol with that name — the file only narrows which definition, callers can still be found in any file. -C, --context <n> shows N lines of real call-site source either side of each hit, rendered exactly like grep -C; omit it (or pass 0) and output is unchanged. Under --json each item gains a contextLines array alongside the existing context field (which names the enclosing symbol, not source text). --exclude-tests hides callers whose call site is a test file (opt-in — omitted, output is unchanged); prints a note naming how many were hidden. --grep <pattern> only shows callers whose enclosing symbol NAME matches this regex (literal substring if it is not valid regex) — rows render as symbol<TAB>file:line, so this is the field each row is keyed on. Applied before the --limit slice, so it selects from the whole caller set, not an already-capped page; when it matches nothing among callers that do exist, the output names the active filter instead of reading like the symbol has no callers. A bare name that isn’t indexed at all reports Symbol not found: <name> (with a Did you mean: suggestion when a near-name candidate is indexed) instead of the misleading “no references found”, which is reserved for a real, indexed symbol that genuinely has zero callers. --json emits each item’s path under both file and filePath with the identical value; file is kept for this release only and will be removed in a future one, so filePath is the spelling to migrate to (matching symbol/types --json). --json emits the shared {items, truncated, totalCount} envelope — the same shape symbol/refs/skeleton/outline --json return, present whether or not truncation occurred, so a script never has to branch on shape. |
| token-goat call-chain <symbol> | Trace every caller layer from a symbol back to the entry points — one step deeper than callers. Use when you need to know what reaches a function across the whole call graph, not only who invokes it directly. Pairs with impact for the downstream direction. Accepts file::symbol to disambiguate WHICH same-named definition the chain starts from — the file only narrows which definition, callers can still be found in any file. --exclude-tests prunes callers whose call site is a test file BEFORE they’re admitted to the traversal, so nothing walks through a test node either (opt-in — omitted, output is unchanged); when every caller was a test, the “no callers” line names how many were hidden instead of reading as genuinely unreferenced. Symbol not found: <symbol> for an unindexed name (bare or file::symbol) now carries a Did you mean: suggestion when a near-name candidate is indexed. --grep <pattern> keeps only completed chains containing a symbol name matching this regex (literal substring if it is not valid regex) — the BFS still walks the full graph, this only narrows which finished chains are reported, so a chain passing through a matching symbol on its way to an unrelated root still surfaces; when it matches none of the chains that do exist, the output names how many were filtered out rather than reading as genuinely caller-less. A chain the walk abandoned because it ran out of --depth ends in a (depth-limit) marker, so a truncated chain is never mistaken for one that reached a real entry point. |
| token-goat impact <symbol> | Walk the call-reference graph forward (breadth-first) and list every function that depends on a symbol, with hop depth; module-scope callers are surfaced as (module scope) <file> entries. Run before a refactor to size up the blast radius without starting a build. Accepts file::symbol to disambiguate WHICH same-named definition the walk starts from — the file only narrows which definition, callers can still be found in any file. --exclude-tests prunes callers (including module-scope entries) whose call site is a test file BEFORE they’re enqueued for further traversal, so nothing walks through a test node either (opt-in — omitted, output is unchanged); when every caller was a test, the “no callers found” error names how many were hidden instead of reading as genuinely unreferenced. A bare name that isn’t indexed at all reports Symbol not found: <name> (with a Did you mean: suggestion when a near-name candidate is indexed) instead of the misleading “no callers found”, which is reserved for a real, indexed symbol that genuinely has zero impact. --grep <pattern> only shows impacted entries whose symbol name (or (module scope) <file> key) matches this regex (literal substring if it is not valid regex), the same filter call-chain --grep/dead --grep apply to their own results — applied BEFORE the --top slice, so it selects from the whole impacted set rather than an already-capped page; when it matches none of the impacted entries that do exist, the output names how many were filtered out instead of reading as genuinely impact-free. |
| token-goat context-for <task> | Takes a natural-language task description, runs semantic search across the indexed codebase, and emits a prioritized list of token-goat read commands trimmed to a token budget. Fetches only the relevant slices instead of loading entire files. --budget N sets the token ceiling; --top N limits the file count; --json for structured output. Every emitted command carries the file::symbol@LINE anchor, so a suggestion still runs when the same symbol name has more than one definition in its file; --json entries carry the matching line field. |
| token-goat answer "<question>" | Deterministic question router: classifies a plain-English question to one of eight intents (where a symbol is, what it does, who calls it, what tests cover it, what a file exports, what a file imports, what imports a file, what breaks if a symbol changes), resolves the subject against the index, and delegates in-process to the command that already answers it. No model call, no inference, no fallback to a fuzzy match. Because the subject is looked up before anything runs, the router picks the right ARGUMENT FORM too: answer "what tests cover foldPath" resolves foldPath to src/path_containment.ts and runs test-for against the file, where passing a symbol to test-for directly fails. A bare basename resolves when exactly one file in the project carries it, and lists the candidates when several do. Every answer is preceded by via: token-goat <command> <argument> so the delegate can be verified or re-run with its own flags. Anything else refuses: cannot answer deterministically: <why>; try: <command>, exit 1. what does X do and what is X for run brief when X has exactly one definition in the project; several definitions refuse with the count and where they are. Questions of judgement, intent or runtime behaviour (why does X..., how does X work, is X correct, should I...) refuse even when they name a real symbol. what imports <file>, which files import it and importers of <file> run deps <file> --importers; asked about a symbol, answer says importers are per file and names deps for the file that defines it. token-goat stats counts every call as answer:<command> (the delegate named on the via: line, with answered or its exit code) or answer:refused with the reason; the question text is not stored. |
| token-goat ask "<question>" (experimental) | Retrieves relevant slices via full-text (BM25) search over the symbol index — not semantic/embedding search — and lists them as pointer-citations plus token-goat read commands. Set TOKEN_GOAT_ASK_BACKEND=claude or TOKEN_GOAT_ASK_BACKEND=codex to synthesize a short answer via that CLI (whatever model it defaults to; token-goat does not force Haiku or any particular tier); with the env var unset, or the named CLI missing from PATH, ask degrades to printing the retrieved pointers with no network call. --top N caps the number of FTS hits (default 8); --json for structured output. Answers are not cached — each call re-retrieves and re-synthesizes from scratch. Every emitted command carries the file::symbol@LINE anchor, so a suggestion still runs when the same symbol name has more than one definition in its file; --json entries carry the matching line field. |
| token-goat changed [<ref>] | List files (or --symbol for symbols) changed since a git ref, without reading the full diff. <ref> and --since <ref> are equivalent (default HEAD~5); --since wins if both are given. --json for structured output. --grep <pattern> only lists changed files whose path matches this regex (literal substring if it is not valid regex) — applied to the file list even in --symbol mode, before any downstream slicing; when it matches none of the files that did change, the output names the active filter instead of reading like nothing changed. --exclude-tests hides changed files that live in a test file (opt-in — omitted, output is unchanged), completing the flag family already on refs/callers/dead/call-chain/impact/semantic/symbol. --grep can only ever select a path, so there was no reliable way to ask for the non-test half of a diff: the negative-lookahead regex that expresses “not a test” silently degrades to a literal substring match whenever the regex-compile fallback fires. Test files are a large share of a typical diff — measured against this repo, 35–54% of changed files across the last 5, 10 and 20 commits. Like --grep it filters the file path, so it applies in --symbol mode too, and it prunes before the per-file index lookup rather than after, so a test file is never queried at all. Composable with --grep (a file must satisfy both; when both are active and --grep is what emptied the list, the --grep notice takes priority, and when --grep left only test files so that --exclude-tests emptied it, the message names both filters rather than claiming no non-test file changed). When it hides every changed file there was, the output names how many were hidden and exits 0, rather than a bare “No files changed.” that would read as a clean diff. Every zero-row path emits the shared {items, truncated, totalCount} envelope under --json. |
| token-goat diff "file::symbol" [range] | Show only the git diff hunk(s) that fall within one symbol’s line range, e.g. token-goat diff "file.ts::myFn" HEAD~3..HEAD, instead of the whole file’s diff. Also accepts read’s symbol@LINE anchor to pick out an otherwise-ambiguous candidate. |
| token-goat blame "file::symbol" | Git blame narrowed to a specific symbol’s lines — no whole-file blame needed. Also accepts read’s symbol@LINE anchor to pick out an otherwise-ambiguous candidate. |
| token-goat log "file::symbol" [ref] | Git commit history scoped to one symbol’s line range via git’s own -L line-range history, instead of a raw git log -- file dump of every commit that touched the whole file. --max-count <n> caps commits shown (default 20); --json for structured output. Also accepts read’s symbol@LINE anchor to pick out an otherwise-ambiguous candidate. |
| token-goat types ["file"] | List type definitions (TypedDict, Protocol, dataclass, Pydantic models) in a file or across the project. --grep <pattern> only shows type declarations whose NAME matches this regex (literal substring if it is not valid regex), applied before output is built; when it matches nothing among declarations that do exist, the output names the active filter instead of reading like there are none. --exclude-tests hides type declarations DEFINED in a test file (opt-in — omitted, output is unchanged), the same definition-site sense dead --exclude-tests uses. Applied before the per-kind --limit slice, so the flag selects from the whole matching set rather than an already-capped page; when it hides every declaration there was, the output names how many were hidden and exits 0, instead of the exit-1 No type declarations found a genuinely empty scope returns. --json’s filePath renders root-relative when a project root resolves, absolute when none does, matching plain-text output. --json emits the shared {items, truncated, totalCount} envelope — the same shape symbol/refs/skeleton/outline --json return, present whether or not truncation occurred, so a script never has to branch on shape. |
| token-goat openapi-outline <spec> | Per-operation listing (method, path, operationId, summary, tags) of an OpenAPI 3.x / Swagger 2.0 spec (JSON or YAML) instead of a raw Read. |
| token-goat openapi-op <spec> <operation> | Full detail (parameters, request body schema, response schemas, description) for exactly one OpenAPI operation instead of a raw Read. operation may be an operationId (exact match) or a "METHOD path" spec, e.g. "GET /users/{id}". |
| token-goat sqlite-tables <db> [--json] | Ultra-compact overview of tables, views, row counts, and column counts in a SQLite database instead of a raw Read. |
| token-goat sqlite-schema <db> | Tables/views, columns, indexes, foreign keys, and row counts of a SQLite database instead of a raw Read. |
| token-goat sqlite-query <db> "<SELECT ...>" | Run a read-only SELECT against a SQLite database instead of a raw Read or shelling out to sqlite3 — rejects any non-SELECT statement. |
| token-goat session-schema [table] [--json] | Authoritative schema catalog for session_store_sql (DuckDB/SQLite) and session SQLite tables (todos, todo_deps), complete with column types, caveats, and common column misnomer mappings (title → summary, role → agent_name), eliminating trial-and-error SELECT * schema exploration (85–95% smaller). |
| token-goat describe <target> [table] [--json] | Unified database schema inspection helper: inspects any known session store table or SQLite .db file on disk, displaying table structure and column definitions instead of exploratory queries. |
| token-goat imports "file" | Show the import graph for a file one level deep. Accepts a comma-separated file list ("a,b,c") to cover several files in one call, one clearly-headed block per file; extra space-separated file arguments are reported in a note naming that comma form instead of being silently dropped. --grep <pattern> only shows imports whose MODULE SPECIFIER matches this regex (literal substring if it is not valid regex), applied before --json’s truncation; when it matches nothing among real imports, the output names the active filter instead of reading like the file has no imports at all. |
| token-goat dep-docs <package> | Extract one installed npm package’s README, package.json metadata, and (if resolvable) a compact .d.ts signature outline, instead of grepping node_modules. |
| token-goat find "<query>" | Find the FILES defining a symbol whose name matches a pattern: a case-insensitive substring scan over indexed symbol names, emitting the distinct file paths. When no name contains the pattern, falls back to an edit-distance match so a mistyped name still lands (getUserr → getUser) — the same ranking Did you mean: uses. The fallback runs only when the substring pass found nothing, so an exact match is never reordered or displaced, and a query near nothing still reports a clean miss instead of unrelated names. A recovered match names what it actually matched on stderr rather than silently answering for a name you didn’t type; --json marks it with fuzzy: true and matchedNames, both absent on an exact hit. --limit <n> caps the file count. It reads every symbol name and path in the project, without bodies, which takes about two seconds on a 546,000-symbol project and twice that when the fallback runs. Matches on NAMES only — for meaning-based search over file content use token-goat semantic. |
| token-goat locate <spec> | Locate a symbol or landmark with exact line spans across the project or in a specific file. Supports name or file::name, with fuzzy edit-distance fallback when no exact match is found. --file <path> restricts search to a specific file; --limit <n> caps results; --json emits the shared envelope with exact line spans. Like find, it reads every symbol name, path and line span in scope without bodies: about two seconds across a 546,000-symbol project. |
| token-goat similar "file::symbol" | Find the top-k symbols most similar to a given symbol, via full-text search over symbol names and bodies. Also accepts read’s symbol@LINE anchor to pick out an otherwise-ambiguous candidate. |
| token-goat test-for "file" | Find test file(s) for an implementation file and list their test functions. --json’s testFile renders root-relative when a project root resolves, absolute when none does, matching plain-text output. --json emits the shared {items, truncated, totalCount} envelope — the same shape symbol/refs/skeleton/outline --json return, present whether or not truncation occurred, so a script never has to branch on shape. |
| token-goat dead | Surface functions, methods, and classes with no recorded callers in the project index. Private names and common entry points (main, app, etc.) are excluded by default. --include-private lifts the underscore filter; --kind narrows to specific symbol types, comma-separated for a union (--kind function,method) — an unrecognized kind errors instead of silently reading as a clean codebase; --top N caps output; --json for structured output. --exclude-tests hides dead symbols DEFINED in a test file (opt-in — omitted, output is unchanged); prints a note naming how many were hidden. --grep <pattern> only shows dead symbols whose NAME matches this regex (literal substring if it is not valid regex), applied before --top’s slice; when it matches nothing among dead symbols that do exist, the output names the active filter instead of reading like a genuinely clean codebase. Results are a heuristic lead — dynamic dispatch and external callers are invisible to static indexing. --json emits both file and filePath with the identical value, root-relative when a project root resolves, absolute when none does, matching plain-text output; file is retained for this release only and will be removed in a future release, filePath is the spelling to migrate to (matching symbol/types --json). --json emits the shared {items, truncated, totalCount} envelope — the same shape symbol/refs/skeleton/outline --json return, present whether or not truncation occurred, so a script never has to branch on shape. |
| token-goat coverage-gaps | Find callables in non-test source files that never appear in a test file’s reference records. Useful for spotting untested surface area before a refactor or release. --top N caps output; --json for structured output. |
| token-goat recent [N] | Show the N most recently edited/accessed files with their symbols. |
| token-goat grep "<pattern>" [paths...] | Built-in fallback regex search over files (no rg shell-out, no caching) — session-aware dedup for raw rg/grep Bash calls is a separate hook, not this command. Accepts zero or more paths: omit to walk cwd, or pass several to search them together with hits merged in argument order under one --max-lines cap. -C, --context <n> shows n lines before and after each match. --symbol annotates each hit with its enclosing indexed symbol — ` [name (kind)] appended in text mode, a symbol: {name, kind, lineStart, lineEnd} | null field per item under –json — null/no tag when the hit falls outside any indexed symbol (e.g. module-level code). |
| token-goat search “
Missed lookups recover surgically: read and section print a “Did you mean…?” list on a miss, and section auto-redirects on an unambiguous heading-prefix match — a typo costs at most one extra glance, not a re-read.
Skill efficiency — the <!-- COMPACT_END --> marker
When Claude Code invokes a skill, it re-injects the full skill body on every subsequent turn. A large skill file (e.g. a 10k-token /improve or /ralph) can cost 40–65k tokens per session across 6 active skills. The <!-- COMPACT_END --> marker solves this: place it in any skill file to split it into a compact form (above the marker, ~400 tokens) and a reference section (below). Token-goat detects the marker the first time the skill fires, caches only the compact slice, and injects that from then on — labeled --- compact form (N tokens) --- so the model knows to request the full body only when it needs the detail.
To add the marker to a skill, open the file and insert <!-- COMPACT_END --> on its own line where the “quick reference ends and the detail begins” — typically after the quick-start table and before step-by-step instructions. The full reference section is still reachable via token-goat skill-section "<name>::<heading>" or token-goat skill-body <name> when needed.
Re-load and direct-read protection. Even without the marker, token-goat protects against the two other ways large skills burn context in a long session:
- If the model tries to
Reada skill file directly (~/.claude/skills/improve/SKILL.md), the pre-read hook intercepts it and emits atoken-goat skill-body improvehint instead — the full 10k–65k tokens never enter context. - If the same skill is invoked a second time in the session (e.g.
/improvecalled again after a/compact), re-load detection fires: instead of re-caching the full body, token-goat emits the cached token count andskill-body/skill-sectionrecall hints. The model can retrieve any section it actually needs rather than absorbing the whole skill again.
To check overhead for your current skills: token-goat skill-size. To inspect compact freshness, run token-goat skill-list — the compact_stale column shows [stale] when a skill’s compact was generated from an older version of the file. Run token-goat skill-compact --all to refresh every stale compact in the current session in one pass.
token-goat install now pre-generates compacts for all installed skills as its final step, so compacts are ready from the first session. If you install new skills after the initial install, run token-goat skill-compact --all manually — or check token-goat doctor --context which reports how many skills were added since the last pre-gen pass and shows the exact command to run.
Memory analysis and cleanup
token-goat memory audits the CLAUDE.md files Claude Code loads for a project for wasted tokens: exact-duplicate lines within one file, duplicate headings, content that overlaps verbatim across files, and near-duplicate sibling auto-memory files (~/.claude/projects/<slug>/memory/*.md). Default mode is --analyze (read-only):
$ token-goat memory
# token-goat memory
Project: C:\Projects\example
## CLAUDE.md files (1)
C:\Projects\example\CLAUDE.md (842 tok)
exact-duplicate lines: 1
line 40 duplicates line 12: "Always run the full test suite before committing."
duplicate headings: none
cross-file overlaps: none
## Duplicate-content clusters (sibling auto-memory files)
none
--fix builds on --analyze. The only change it can apply automatically is removing exact-duplicate lines (keeping the first occurrence) — a pure structural dedup with no judgment call. Duplicate headings and cross-file overlaps are printed as advisory findings only; they often mean content should move into a path-scoped .claude/rules/ file or a subdirectory CLAUDE.md, but token-goat never picks where for you, so no diff is proposed for those.
Every proposed exact-duplicate-line fix is shown as a diff before anything is written, gated by the same confirm-before-write flow: pass --yes to apply non-interactively (scripts, CI), or run it from a terminal without --yes to be prompted per file. Running --fix without --yes from a non-interactive shell (no TTY) prints the diffs as a dry run and writes nothing.
Session waste ledger
token-goat waste parses the current project’s Claude Code session transcript — the JSONL file Claude Code writes under ~/.claude/projects/<slug>/*.jsonl — and attributes token cost to every tool call in it, then flags a few concrete waste signals: files that were Read once and never referenced again, and Bash commands run repeatedly without ever hitting token-goat’s own bash-output cache. By default it auto-discovers the most-recently-modified Claude Code or Copilot CLI transcript for the current project, preferring a Copilot CLI session that is still running; pass --transcript <path> to point at a specific one instead (useful when several sessions are open, or for CI/testing):
$ token-goat waste
# token-goat waste
Transcript: C:\Users\you\.claude\projects\C--Projects-example\a1b2c3d4-....jsonl
Total tokens: 18420
## Tokens by tool
Read: 9120 tok
Bash: 6210 tok
Grep: 2140 tok
Edit: 950 tok
## Top expensive tool calls
[3400 tok] Read: src/big_module.ts
[1800 tok] Bash: npm test
## Read once, never touched again
src/unrelated_helper.ts: 640 tok, never referenced again
## Repeated Bash commands not hitting the token-goat cache
"git status": ran 4 times, 210 tok each, 840 tok total, uncompressed
## Assistant output (re-send CEILING, not real spend)
42 turns, 21300 tok generated
Re-send upper bound: 187400 tok if every turn were resent at full price on every later request
Real cost is substantially lower: prompt caching bills resent conversation history at cache-read rates, not full input price.
--top <n> controls how many entries appear under “Top expensive tool calls” (default 10). --json prints the same report as machine-readable JSON instead.
--copilot reads a GitHub Copilot CLI session instead, from <copilot-home>/session-state/<id>/events.jsonl. It is a different report rather than the same one with different inputs, because Copilot writes down its own token accounting at shutdown and token-goat reports those numbers rather than estimating them:
$ token-goat waste --copilot
## Per-request fixed overhead (Copilot's own token counts)
System prompt: 7,897 tok
Tool definitions: 10,993 tok
Conversation: 577 tok
## Tool definitions by MCP server (estimated)
github-mcp-server: 6 tools, 6 KB, ~2,135 tok, 0 calls this session
~2,135 tok estimated across 1 server, re-sent every request.
Copilot counted 10,993 tok of tool definitions in total, so this is roughly 19.4% of it.
Dropping a server saves its line above on every request for the rest of the session.
Never called this session: `copilot mcp disable github-mcp-server` drops it from future sessions.
'copilot mcp enable <name>' restores one, and '--disable-mcp-server <name>' drops one for a single run instead.
Calls per server come from the server name Copilot records on every MCP tool call. A server with none gets the command that turns it off, which Copilot keeps in its settings until copilot mcp enable undoes it. A log that records no tool calls at all says so instead, since it cannot tell a server nobody needed from a session that never ran a tool. token-goat audit carries the same verdict into its recommended fix.
The fixed overhead is the largest number in a Copilot session and no hook can reach it: Copilot assembles the system prompt and the tool definitions natively, with nothing between assembly and send. Only configuration moves it. The per-server breakdown exists to make that configuration decision possible, since one aggregate says the tool definitions are expensive without saying which tools. It is read from Copilot’s own MCP tool cache, counts only the fields a model is actually sent, and is labeled an estimate throughout: it comes from byte length rather than Copilot’s tokeniser, and it deliberately does not add up to Copilot’s total, because Copilot’s own built-in tools are not cached there.
The “Assistant output” section is separate from the tool-call ledger above it: generatedTokens is what was actually paid, once, to produce the assistant’s own text turns. resendCeilingTokens is a cache-unaware upper bound on how much re-sending those turns as conversation history on every later request could cost — not real spend, since Claude Code’s prompt caching bills a repeated conversation prefix at cache-read rates, a fraction of full input price. Treat it as a ceiling on how bad unbounded verbosity could get, not as a dollar figure.
Cross-cache recall
token-goat recall "<query>" searches every cached bash-output, web-output, and mcp-output entry at once, so you don’t need to remember which cache type holds the result you want — a single full-text query ranks hits across all three:
$ token-goat recall "eslint warnings"
[bash] a1b2c3d4e5f6a7b8 (token-goat bash-output a1b2c3d4e5f6a7b8)
npx eslint src tests
npx eslint src tests\n[token-goat: delta] 2 of 5 prior issues resolved; remaining: 3
[mcp ] mcp_9f8e7d6c5b4a3210 (token-goat mcp-output mcp_9f8e7d6c5b4a3210)
mcp:mcp__plugin_github_github__get_check_runs {"owner":"..."}
... eslint warnings found in 2 files during CI ...
Results are ranked by relevance (BM25 via SQLite FTS5, falling back to a plain substring scan if FTS5 is unavailable), newest indexed entries win ties. --type bash|web|mcp narrows to one cache type; --limit <n> caps the result count (default 10); --json emits { id, cacheType, label, snippet, storedAt }[] instead. The index is built incrementally as entries are cached — there is no separate rebuild step.
Run token-goat recall with no query to browse instead of search: every cached entry across all three types, newest first, in the same format and honoring the same --type/--limit/--json flags. This is the case where the index matters most — the ids have scrolled out of context and you have no term to search for, so the alternative is running bash-history, web-history, and mcp-history in turn.
Hint efficacy tracking
Every hint hook (the re-read/dedup/surgical-read nudges in the Bash, Read, Edit, Grep and Glob hooks, and the call-streak hints) is
worth its keep only if it’s actually followed. token-goat hint-stats reports, per hint
category: how many times it fired, how many times a later Bash command in the same session
actually invoked the specific token-goat command (or referenced the specific cached-output id)
the hint pointed at, the resulting efficacy percentage, whether the category is currently
auto-suppressed, and the spent-bytes column (bytes of hint text actually injected into context for
that category — the real cost of emitting it, not just how often it fired):
$ token-goat hint-stats
category emitted undisplayed acted-on efficacy suppressed manual+ manual- spent-bytes
bash_redirect 42 +3 - 9 21.4% no 0 0 3150
bash_recall 18 - 15 83.3% no 0 0 1080
read_reread_dedup 11 - 2 18.2% * no 0 0 660
read_structural_nav 7 3 1 14.3% * yes 0 1 420
edit_reread_suggest 3 - 0 0% * no 0 0 180
read_batch 6 - 4 66.7% no 0 0 900
search_brake 2 - 1 50.0% no 0 0 330
grep_dedup_hint 0 ~4 - 0 n/a no 0 0 240
glob_dedup_hint 0 ~1 - 0 n/a no 0 0 60
TOTAL saved-bytes=48200 (all-time, every hint kind) spent-bytes=7020 (hint_emissions ledger only)
spent-bytes (and the TOTAL line’s spent-bytes) render n/a instead of a fake 0 whenever a
category — or, for the total, the whole store — has no tracked spend figure at all: either
nothing has fired yet, or every emission predates this feature and was recorded before spend
tracking existed. A partially-tracked category shows the real sum plus how many legacy rows it
excludes, e.g. 120 (2 legacy), rather than silently blending unknown-cost rows into the total
as if they cost nothing. n/a means nothing in that category reached the agent: a category whose
only emissions are still pending or were never scorable shows their real spend.
emitted counts only the emissions that have been scored. Two kinds sit beside it instead:
+Nis how many are still waiting on a verdict. Each hint gets the next few tool calls in its session to be followed; until those calls arrive, it has been scored neither way. A session that ends before they arrive leaves its hints pending for good, and they are never scored.~Nis how many carried nothing a later call could match (seegrep_dedup_hintbelow), so no verdict was ever possible.
Neither is counted in emitted, acted-on, the efficacy figure or the suppression decision.
Both are counted in spent-bytes, because the hint text reached the agent either way. Counting
pending hints as failures once muted bash_redirect on a real ledger at 1 in 11 (9.1%) when its
scored hints stood at 1 in 6 (16.7%): five of the eleven were left pending by a session that ended
mid-window.
A category is auto-suppressed for its harness only once the numbers show it is confidently below
the bar, not merely below it on a small sample. token-goat takes the 95% Wilson interval on the
category’s efficacy and mutes it when the top of that interval sits under
hint_stats.suppress_threshold_pct (default 15%). With the defaults, a hint nobody ever follows is
muted after 22 scored emissions; 1 followed in 7 is not muted, though its raw rate is 14%. For a
starred category (below) the test runs on the defiance rate instead: it is muted once the bottom of
that interval sits over hint_stats.defiance_threshold_pct, which defaults to 100 minus the
suppress threshold. hint_stats.min_sample_size (default 5) is a floor under both tests. At the
default threshold the interval never allows muting sooner, so the floor only matters once the
threshold is raised. Pending hints are not part of the sample.
A suppressed category is not silenced for good. On the suppressed occasions listed in
hints.backoff_thresholds (default the 1st, 3rd, 10th and 30th, then every 30th) the hint is shown
anyway as a probe, and a probe the agent follows lifts the upper bound back over the bar. token-goat
hint-stats --reset clears the tracked data outright. Configure the knobs with token-goat config set
hint_stats.suppress_threshold_pct <pct>, hint_stats.defiance_threshold_pct <pct> and
hint_stats.min_sample_size <n>. To see how the rule behaves across uptake rates, and how each
category on a ledger scores under it, run npm run eval:hints from a checkout (add
-- --db <path to global.db> to read a ledger; it is opened read-only).
A * on an efficacy figure means that category is scored on an absence. Those hints ask the
agent not to do something, so a window that expires with no re-read counts as compliance, while
an unmarked category only scores when the agent actually runs the command the hint named. The two
percentages sit on different scales and comparing them straight across gives the wrong answer: a
high starred figure means the warned-against read was not seen, not that the hint persuaded
anyone.
“Acted on” is a real, session-scoped signal (the exact file path or cached-output id the hint’s
own text pointed at is checked against the next few tool calls in that session) — not a guess —
but it is a proxy for correlation, not proof of causation: a match means the agent ran the
suggested command shortly after the hint, not that the hint necessarily caused it. A hint whose
text has no extractable path/id (a small minority of branches) cannot be scored at all: it is
left out of emitted and the efficacy figure, shown as the ~N beside emitted, and counted
only in spent-bytes. --mark-effective <category> / --mark-ineffective <category>
record a separate manual vote as a human override/supplement for exactly that gap — manual votes
are shown alongside the automatic percentage but never blended into it. --json emits
{ category, emitted, actedOn, efficacyPct, pending, unobservable, detected, suppressed, suppressionPermanent, manualEffective, manualIneffective, bytesEmitted, legacyEmissions }[].
With --session-id it emits an object instead, carrying the same rows under rows beside a
scope that names which fields are limited to that session.
Note that what this feature calls “harness” (Claude Code, Codex, Gemini, …) is not the same as
“which LLM model” — no bridge in this codebase exposes an LLM model identifier to hooks, so
harness is the closest real signal available.
Two categories are scored on a pattern of calls rather than a named command. read_batch (three or more reads or searches in a row, each sent a full turn after the previous result) counts as acted on when the first later turn that makes read-only calls makes at least two of them together. search_brake (three searches in a row found nothing) counts as acted on when the next search from a later turn is token-goat answer or token-goat semantic. Until that later call arrives the emission stays pending (+N) and is counted neither way. grep_dedup_hint and glob_dedup_hint (an identical Grep or Glob already ran this session) are never scored: the note rides on the re-run it describes, so no later call can show whether it was heeded. They appear in spent-bytes and as the ~N beside emitted, never in emitted itself or the efficacy figure.
Scoring the compressors — token-goat bench
token-goat bench replays a fixed corpus of captured command output through the same function
that decides what a real shell command hands back to the model, and reports how much smaller the
result is. It exists so a change to a compressor can be measured instead of guessed at:
$ token-goat bench
case filter in out saved fidelity
---------------------------------------------------------------
git-log-stat git-log 57227 3027 94.7% 2/2
npm-ls-all dep-list 34546 1060 96.9% 2/2
vitest-run vitest 22684 322 98.6% 2/2
---------------------------------------------------------------
TOTAL 114457 4409 96.1% 6/6
ratio 96.1% saved (PRIMARY -- must improve; measured floor 0.0%, headroom 3.9%)
fidelity 6/6 kept (GUARD -- must not regress; any miss exits 1)
coverage 3/157 filters exercised, 3/3 cases compressed
There are two numbers on purpose. Ratio is the thing to push up. Fidelity counts the lines each case declares it must never lose, and it is what stops the ratio from being gamed: deleting output raises the ratio, so a benchmark reporting only the ratio would reward destroying the very content the compressors exist to preserve. A dropped must-keep line exits 1, so a loop can revert on the exit code alone.
The corpus lives in tests/fixtures/bench: one <id>.json naming the command, its exit code,
and its must-keep lines, beside an <id>.txt holding that run’s real output. Point `–corpus