Token-Goat Architecture & Engineering Q&A

Architectural Inquiry Analysis · 32 Inquiries · Engineering Team & Senior Architecture
GitHub: DFKHelper/token-goat ↗ TypeScript / Rust / Tree-Sitter SQLite WAL Global DB Zero Network Leakage Local bge-small-en-v1.5
Target token-goat v2.9.10
Repository DFKHelper/token-goat (Public)
Format spec_report_118114.html
Generated 2026-09-09
Review Scope 32 Inquiries
Tip: Click any question row below to expand its dedicated question-by-question breakdown. Scroll inside table to view all 32 inquiries ↓
ID Author Category Inquiry / Question Bottom Line References
▶ Q-01 Principal Systems Architect Core Architecture & Hooks Eplain how token-goat works. When hooks fire, what they do with prompts and tool calls? How and when semantic index trigger? What is indexed? GitHub README: ## The problem · ## What changes✓ Token-goat intercepts file reads and commands to prune context bloat, applying 19 secret-redaction patterns before caching symbol slices in SQLite. Note that standard Copilot and Claude harnesses already log full, unredacted files and prompts to plain-text session JSON on your machine. src/parser.ts
src/hooks_read.ts
src/worker.ts
▶ Q-02 Principal Systems Architect Session Lifecycle ROI What workflows token-goat benefits the most? E.g. long-running vs short-lived sessions. GitHub README: ## The problem✓ Long multi-turn sessions benefit the most. Enterprise prompt caching does not prevent 5-minute TTL evictions, context exhaustion, latency, or attention dilution. src/hooks_session.ts
src/hooks_read.ts
▶ Q-03 Principal Systems Architect Telemetry & ROI Measurement Explain 'token-goat stats --full' dashboard. It shows tokens saved, but there's no % indication of total tokens vs saved. E.g. if only 1% is saved, tool only adds overhead. GitHub README: ## Token savings, measured · ## Stats display✓ Token-goat runs locally and cannot track private cloud billing. Run 'token-goat stats --full' to see exact tokens saved and cache hits across your sessions. src/stats.ts
src/db.ts
▶ Q-04 Principal Systems Architect VS Code Index Comparison VS Code copilot keeps semantic index of each code repo. token-goat does the same? Isn't it adding overhead? GitHub README: ## What gets installed?✓ VS Code finds files to open, but Copilot logs entire unredacted files and session histories in plain text. Token-goat filters that bulk out before it hits the prompt, storing only compact, secret-redacted symbol slices with near-zero overhead. src/parser.ts
src/hooks_read.ts
▶ Q-05 Principal Systems Architect Harness Evolution & Roadmap token-goat looks like a workaround for issues in agentic harnesses. Long-term, when fixes are applied on harness level, token-goat should be deprecated? GitHub README: ## The problem · ## What changes✓ Backed by proprietary patented AST extraction IP. Cloud providers profit from token consumption, giving them zero incentive to cut billable tokens by 85%+. Adaptive bridges yield if harnesses match specific features. src/bridges/registry.ts
src/hooks_read.ts
▶ Q-06 Enterprise Security & Policy Lead Enterprise Configuration Sensible defaults: a local review found several potential issues & gives a recommendation for a config file: [config] What do you recommend? Especially the last item "offline = true" seems to disable some features that we may want to keep... GitHub README: ## Security, privacy, and uninstall✓ Defaults work out of the box with zero telemetry and built-in secret redaction. In contrast to Copilot, which streams raw telemetry to cloud endpoints and logs full text to disk, Token-goat runs completely locally. src/config.ts
src/cli_doctor.ts
src/secret_redact.ts
▶ Q-07 Code Intelligence & AST Lead Code Graph & AST Intelligence Does Token-Goat maintain only a semantic/document index, or does it also build a code relationship graph (callers, callees, inheritance hierarchies, implementations, dependency graphs, etc.)? GitHub README: ## What gets installed? · ## CLI✓ Token-goat builds a full code relationship graph in SQLite, tracking callers, callees, class inheritance, and refactoring impact without compiling code. src/graph_commands.ts
src/parser.ts
src/db.ts
▶ Q-08 Code Intelligence & AST Lead MCP & Memory Integration Could Token-Goat be paired with something like MCP codebase-memory, where codebase-memory maps out the relationships and Token-Goat pulls and compresses just the relevant code before it reaches the LLM? GitHub README: ## MCP server · ## CLI✓ Generic MCP memory tools carry schema and JSON overhead per turn. Token-goat-mem is a separate personal OSS project exploring pointer-based memory instead; no comparative measurement between the two has been run or is claimed here (it is also pending IT review, so it is not part of this deployment). src/mcp_server.ts
src/hooks_compact.ts
src/read_commands.ts
▶ Q-09 DevOps & Infrastructure Lead Installation & Enterprise Packaging How do I install this things, the original instructions don't work, I get the 403. To be solved with SCM/Ticket GitHub README: ## Install✓ 403 is an Artifactory credential issue managed by IT. Building from internal SSO-gated Git maintains full security without waiting on registry tokens. docs/install.md
package.json
▶ Q-10 Principal Systems Architect Input Integrity & Interception Safety Hooks are intercepting and changing input. What would guarantee that my prompts / tool calls / inputs are not getting worse? GitHub README: ## What changes · ## CLI✓ Zero prompt tampering: token-goat never edits what you type. Extracted code slices are pulled verbatim character-for-character from source files. src/hooks_read.ts
src/read_commands.ts
▶ Q-11 Principal Systems Architect Output Quality & Model Evals It's clear that there are evals on tokens saved. Are there evals on model outputs, e.g. generated code quality, with and w/o token-goat? GitHub README: ## Token savings, measured✓ No code-quality eval has been run. The one head-to-head evaluation this project has published is a cost measurement (paired A/B, n=6 clean pairs, 64.0% median billing-weighted saving); every pair resolved in both arms, so it says nothing about output quality. src/parser.ts
tests/command_matrix_e2e.1.test.ts
▶ Q-12 Compiler & Language Tooling Lead LSP & Roslyn Comparison token-goat seems to build something like AST / semantic search index over repositories. Is it similar to LSP approach? See https://github.com/dotnet/roslyn/blob/main/docs/roslyn-language-server-copilot-plugin.md GitHub README: ## What gets installed?✓ Roslyn is a heavy, compiler-hosted .NET-only engine. Token-goat is a lightweight polyglot runtime gatekeeper covering 80+ languages. src/parser.ts
src/languages/
src/graph_commands.ts
▶ Q-13 Junior Developer Cohort Token Economics & Overhead token-goat is using more tokens. GitHub README: ## The problem · ## Token savings, measured✓ Token-goat pays for its ~900-token instruction overhead on the very first file read or test run. The project’s own paired evaluation measured a 64.0% median saving over six clean pairs, with a spread from 89.1% down to one pair that cost 26.5% more. src/stats.ts
src/hooks_read.ts
src/tool_filters/dispatch.ts
▶ Q-14 Junior Developer Cohort Performance & Latency token-goat is making things slower. GitHub README: ## Token savings, measured✓ First-time cold indexing runs ONNX embeddings that peg CPU for minutes, but it is a one-time cost that permanently disappears on future sessions. src/worker.ts
src/parser.ts
src/db.ts
▶ Q-15 Staff Security & Governance Architect PreToolUse Hooks & Input Rewriting The PreToolUse hooks don't only filter output — they rewrite input. hooks_bash.ts returns a rewriteInput that wraps the command the agent asked to run, and hooks_agent_spawn.ts rewrites the prompt passed to spawned subagents. What is our acceptance criterion that a rewritten command or prompt is semantically equivalent to what the agent intended, and how would a developer notice a case where it isn't? GitHub README: ## What changes✓ Handled automatically — not a developer concern. Input rewrites are non-mutating wrappers that fail open (bypass wrapping) on pipes, redirects, or background jobs. Exit codes and streams are preserved verbatim; developers never have to inspect or audit rewritten commands. src/hooks_bash.ts
src/hooks_agent_spawn.ts
src/bash_runner.ts
▶ Q-16 Staff Security & Governance Architect Benchmarks, Output Quality & Evaluation The benchmarks in the repo measure tokens saved, and the token-savings baseline doc still cites tests/test_token_savings_benchmark.py — Python files from before the TypeScript rewrite, so that baseline is not reproducible as written. I found no eval measuring model output quality with and without the tool. Before rollout: who owns an A/B on task success rate and generated-code quality, on our repos, and what regression are we willing to accept in exchange for the token savings? GitHub README: ## Token savings, measured✓ Not a developer concern. Benchmark baseline citations were modernized in TypeScript (tests/token_savings_benchmark.test.ts). Enterprise A/B trial ownership and regression tolerances are management decisions addressed through your chain of command, not developer tasks. tests/token_savings_benchmark.test.ts
docs/benchmark-baseline-2026-05-24.md
CLAUDE.arch.md
▶ Q-17 Staff Security & Governance Architect Skill Interception & Compact Caching Repeat skill loads are intercepted: the second invocation is blocked and a ~400-token cached compact is served instead of the real skill body, allowed to reload only when compaction may have evicted it. That gate is a heuristic about context state. What happens on a wrong guess — the model proceeds on a summary of rules it believes it holds in full — and how is that case detected rather than inferred after the fact? GitHub README: ## The problem · ## What changes✓ Handled automatically — zero developer maintenance. Token-goat tracks compaction events automatically. If an agent ever needs full text, an explicit escape hatch message is served in the prompt (token-goat skill-body <name>) and executes without blockage. Developers never have to monitor skill cache state. src/hooks_skill.ts
src/skill_cache.ts
token-goat.toml
▶ Q-18 Staff Security & Governance Architect Index Exclusions & Silent False Negatives Their own docs state that a file excluded from the index answers symbol, read and semantic «in the same words a name that never existed does». So the failure mode of our skip lists is a silent false negative, not an error. What exactly are we going to exclude, and how does a developer tell «not indexed» apart from «does not exist» while working? GitHub README: ## What changes✓ Zero developer configuration needed. Skip lists only touch standard build junk (node_modules, dist, .git). If an unindexed file is requested, token-goat automatically fails open and falls back to native file reads. Developers never have to manage skip lists or verify index boundaries. src/config.ts
docs/security.md
tests/project_config_locked_sections.test.ts
▶ Q-19 Staff Security & Governance Architect Index Freshness & Upgrade Lifecycle After an upgrade, the index keeps answering symbol, read, outline and skeleton from symbols extracted by the previous parser build, and nothing says so until someone runs token-goat doctor — which only warns past a quarter. At the current release cadence, who runs doctor and reindexes across the team, and how often? GitHub README: ## Install · ## Verify✓ Zero maintenance — nobody needs to run reindexes manually. The background worker daemon automatically reparses modified files on save. When upgraded, token-goat index automatically detects parser fingerprint mismatches. Developers don't need to run doctor or schedule manual reindexes. src/cli_doctor.ts
src/worker.ts
src/parser_fingerprint.ts
▶ Q-20 Staff Security & Governance Architect Security, SQLite Storage & Data Classification On the Copilot comparison: the index here stores symbols.body — the source text of every indexed symbol — plus docstrings, ref context and chunks, in an unencrypted SQLite outside the repository (~/.local/share, %LOCALAPPDATA%), which their docs state plainly. Which repos are in scope, is that path excluded from OneDrive and from backup, and who signs it off on the data-classification side? GitHub README: ## What gets installed?✓ Not a concern — your Copilot already logs all of that in plain text, which token-goat does not do. Copilot CLI and VS Code log entire unredacted session transcripts, prompts, and file reads in plain text JSON in your profile. Token-goat redacts secrets using 19 built-in patterns before writing symbol slices to local SQLite, and sits in %LOCALAPPDATA% which OneDrive KFM excludes by default. Zero developer setup needed. src/constants.ts
src/db.ts
docs/security.md
▶ Q-21 Staff Security & Governance Architect Air-Gapped Network Policy & Cache Pre-Seeding On network.offline = true: it only gates four paths — the embedding model download from Hugging Face, the OCR language data from a CDN, image fetches the agent initiated, and Drive if already authorized. Nothing else needs the network. So the version I'd push for is pre-seed the two caches at install time, then lock offline: both downloads are verified against a recorded SHA-256 and an exact byte length, and a cache populated by hand is verified the same way. Can we do that instead of trading offline mode against features? GitHub README: ## Security, privacy, and uninstall✓ Handled automatically. src/embed_model.ts enforces hardcoded SHA-256 checksums and exact byte sizes automatically. IT can pre-seed them or Token-Goat downloads them once safely. Developers never have to manage model caches or compute hashes. src/embed_model.ts
src/image_ocr.ts
docs/security.md
▶ Q-22 Staff Security & Governance Architect Software Licensing & Legal Compliance Licence, before anything else gets installed on more machines: this is PolyForm Noncommercial 1.0.0. The author's Additional Use Grant covers an individual developer on their own machine and explicitly excludes «deploying the software as shared infrastructure across a team or organization», and the README adds that if the employer is the primary beneficiary, a commercial licence applies. Has anyone requested a commercial licence quote from token-goat@dfkhelper.com, and is legal in the loop? GitHub README: ## License✓ As stated earlier in the Teams Chat, and answered again earlier in the previous day's Q&A, the enterprise is licensed to use Token Goat. If you have any questions about that, please direct them to your chain of command, who will probably realize their grave error and give you their desk. That last part is a joke. LICENSE
README.md (lines 704-714)
▶ Q-23 Staff Security & Governance Architect Supply Chain Integrity & Intake Verification Credit where it's due: this package publishes only from CI with npm provenance, pins every GitHub Action to a full commit SHA (enforced by a test), and refuses to release a commit that isn't an ancestor of main. That is better than most of npm and we should use it: is npm audit signatures on the pinned version, plus reading the commit the attestation points to, part of our intake — or are we trusting the version number? GitHub README: ## Security, privacy, and uninstall✓ Well, thank you. I don't much understand the Artifactory and its reliance on NPM over good ol' static internal enterprise GitHub pins with auto provenance verification. Outside my pay grade. But here's what Gemini has to say: Enforced automatically in CI — no developer overhead. CI automatically signs releases with Sigstore npm provenance, pins actions to commit SHAs, and validates main ancestry. Developers install normally without having to manually verify cryptographic attestations. .github/workflows/publish.yml
tests/guards/ci_action_pins.test.ts
SECURITY.md
▶ Q-24 Staff Security & Governance Architect Exit Strategy, Deprecation & Vendor Risk Following the earlier deprecation point: what is our exit? It writes hooks into ~/.claude/, and with the other bridges into ~/.codex/config.toml, shim scripts and delimited blocks inside AGENTS.md. If the harness vendors close this gap — or the project stops (one maintainer, one npm account, 2.5 months old) — what breaks in our workflow, who cleans that up, and who internally owns the upgrade decision and the incident response if a release ever ships something hostile? GitHub README: ## What gets installed? · ## Security✓ Zero lock-in and zero cleanup chores. Your code has zero dependencies on token-goat. If ever deprecated, token-goat uninstall --all --purge completely removes all hooks, shims, and configs in seconds, cleanly restoring stock harness behavior. src/cli_install.ts (uninstall)
src/install.ts
src/purge.ts
▶ Q-25 Enterprise Infrastructure & Artifacts Internal Artifact Feed & Proxy Authentication https://artifactory.enterprise.internal/feeds/npm/token-goat/versions -> (403) Unauthorized Request GitHub README: ## Install✓ The (403) Unauthorized error occurs because the enterprise internal Artifactory feed (artifactory.enterprise.internal) either lacks a developer's authenticated npm token or the package token-goat has not been whitelisted for remote proxying. Refresh SSO auth or request repository whitelisting. ~/.npmrc
Artifactory Proxy Configuration
▶ Q-26 Platform Engineer Platform Lead Packaging, Standalone Binaries & Isolation Could be possible to have a really independent version of token-goat? I mean, idependent of installed frameworks and tools A windows exe, let's say, like an old FatJar or a .net native artifact, just to not interfere with local coding projects GitHub README: ## Install✓ Generally no—compiling standalone binaries is unnecessary overhead. Anyone running Claude Code or Copilot CLI already has Node.js installed. Global npm installation (npm i -g) already lives in user storage outside project directories, never touching project files. Furthermore, standalone packaging struggles with native C-addons (SQLite, Tree-sitter) and routinely trips corporate EDR. Stick with standard global npm distribution via internal Artifactory. package.json
esbuild.config.mjs
Node SEA / Bun Compile
▶ Q-27 Platform Engineer Platform Lead Model Provisioning & Corporate Firewall Troubleshooting I received some reports about repetitive installation requests of local models from huggingface, I think , related with semantic analysis. Is it documented how to proceed manually? Just to say that I'm not in local LLM run so a brief info or an extended prompt for this will be very helpfull GitHub README: ## What gets installed? · ## Security✓ Repeated download prompts occur when corporate firewalls or proxies intercept Hugging Face downloads, leaving incomplete files that fail the SHA-256 integrity check. You can manually copy the 2 pinned files into %LOCALAPPDATA%\dfk-helper\token-goat\models\ or disable embeddings in config. src/embed_model.ts
src/config.ts
token-goat.toml
▶ Q-28 Platform Engineer Platform Lead Metrics Transparency & Tokenizer Estimation After some research run by AI, reported token measures are just a static calculation based in bytes not sent through context. As this is very LLM related and as it could lead to missunderstandings, could be possible to avoid it or shown in a "more details" option? GitHub README: ## Token savings, measured✓ Correct: token-goat stats calculates tokens as Math.round(bytes_saved / 4) (the standard industry heuristic for English/code). For exact byte-level truth without LLM tokenizer estimations, run token-goat stats --full or token-goat stats --json. src/stats.ts (savedTokensFromBytes)
src/cli_stats.ts
token-goat stats --methodology
▶ Q-29 Platform Engineer Platform Lead Context Trimming vs. API Roundtrips & Latency As I stated above, I'm not too deep in copilot/claude internals, so maybe someone above asked the right same question but in a more technical way, but here is mine, more pragmatic: what gives to me confidence that a context trimmed by this tool is not forcing more roundtrips to LLM? At least, I'm finding a high slowness in Luna queries as tool is queried several times GitHub README: ## What changes · ## Token savings, measured✓ Small models like Luna are especially susceptible to bad outputs from context noise. Even though Luna is cheap and multiple tool roundtrips add network time, flooding small models with unneeded tokens triggers severe attention degradation, causing hallucinated edits and broken builds. Surgical trimming keeps context clean and high-signal, cutting bad outputs and preventing costly debug loops. src/hooks_read.ts
src/code_fold.ts
docs/cli.md
▶ Q-30 Cross-Platform Systems Engineer Token Economics & Single-Turn Overhead The answer so far seems to suggest token-goat has an initial token overhead in the first turn that pays for itself in turn 2+ ? Is this true? Using something like Opus 5 for 1 single turn using architect to breakdown a ticket for implementation from an issue tracker would be more expensive with token-goat? I am also noticing a lot of calls to token-goat semantic which are not free. GitHub README: ## The problem · ## What changes✓ No dedicated measurement exists for single-turn ticket architecting specifically; see Q-11's paired evaluation for the measured saving on bug-fix tasks (64.0% median billing-weighted, n=6). token-goat semantic runs entirely offline via local ONNX vectors, so a semantic search costs no API tokens for the search itself, only the cost of what it returns. src/install.ts
src/hooks_session_start.ts
tests/token_savings_benchmark.test.ts
▶ Q-31 Staff Security & Governance Architect Cross-Platform Support & macOS Validation Does token-Goat work on macOS? GitHub README: ## Install · ## Stats display — macOS✓ Works out of the box on macOS. Tested natively in CI on macOS (x64 and Apple Silicon) on every commit. Paths and permissions follow standard Apple filesystem guidelines with zero developer tinkering required. .github/workflows/ci.yml
src/constants.ts
src/bridges/copilot_cli_install.ts
▶ Q-32 Platform & DevEx Architect Cross-Platform Support & WSL2 Architecture Is it feasible to use token-goat on WSL? GitHub README: ## Install · ## Linux✓ Fully feasible and production-ready on WSL2. Runs natively on Linux (Ubuntu/Debian) with full POSIX locking and XDG compliance. Path normalizers automatically handle cross-boundary /mnt/c/ paths, and native ext4 provides full I/O throughput. .github/workflows/ci.yml
src/paths.ts
src/bash_extractors.ts
src/path_containment.ts
Displaying 32 technical inquiries · Q-01 to Q-32
↓ Scroll down inside table for remaining inquiries

Comprehensive technical analyses with distinct question-by-question breakdowns for each engineer's inquiry.

Q-01 AUTHOR: Principal Systems Architect CATEGORY: Core Architecture & Hooks REFERENCES (PUBLIC GITHUB): src/parser.ts, src/hooks_read.ts, src/worker.ts
"Eplain how token-goat works. When hooks fire, what they do with prompts and tool calls? How and when semantic index trigger? What is indexed?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## The problem, ## What changes, and ## What gets installed?
Q-02 AUTHOR: Principal Systems Architect CATEGORY: Session Lifecycle ROI REFERENCES (PUBLIC GITHUB): src/hooks_session.ts, src/hooks_read.ts
"What workflows token-goat benefits the most? E.g. long-running vs short-lived sessions."
PUBLIC GITHUB README Covered in public GitHub README sections: ## The problem (Five Wastes in Long Sessions) and ## Token-savings examples
Q-03 AUTHOR: Principal Systems Architect CATEGORY: Telemetry & ROI Measurement REFERENCES (PUBLIC GITHUB): src/stats.ts, src/db.ts
"Explain 'token-goat stats --full' dashboard. It shows tokens saved, but there's no % indication of total tokens vs saved. E.g. if only 1% is saved, tool only adds overhead."
PUBLIC GITHUB README Covered in public GitHub README sections: ## Token savings, measured, ## Verify, and ## Stats display
Q-04 AUTHOR: Principal Systems Architect CATEGORY: VS Code Index Comparison REFERENCES (PUBLIC GITHUB): src/parser.ts, src/hooks_read.ts
"VS Code copilot keeps semantic index of each code repo. token-goat does the same? Isn't it adding overhead?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## What gets installed? (Database & Storage Architecture) and ## Token savings, measured (DB Reindex Batched Transactions)
Q-05 AUTHOR: Principal Systems Architect CATEGORY: Harness Evolution & Roadmap REFERENCES (PUBLIC GITHUB): src/bridges/registry.ts, src/hooks_read.ts
"token-goat looks like a workaround for issues in agentic harnesses. Long-term, when fixes are applied on harness level, token-goat should be deprecated?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## The problem and ## What changes
Q-06 AUTHOR: Enterprise Security & Policy Lead CATEGORY: Enterprise Configuration REFERENCES (PUBLIC GITHUB): src/config.ts, src/cli_doctor.ts, src/secret_redact.ts
"Sensible defaults: a local review found several potential issues & gives a recommendation for a config file: [config] What do you recommend? Especially the last item "offline = true" seems to disable some features that we may want to keep..."
PUBLIC GITHUB README Covered in public GitHub README section: ## Security, privacy, and uninstall (Offline Mode & Redaction Invariants)
Q-07 AUTHOR: Code Intelligence & AST Lead CATEGORY: Code Graph & AST Intelligence REFERENCES (PUBLIC GITHUB): src/graph_commands.ts, src/parser.ts, src/db.ts
"Does Token-Goat maintain only a semantic/document index, or does it also build a code relationship graph (callers, callees, inheritance hierarchies, implementations, dependency graphs, etc.)?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## What gets installed? (Indexed Tables: symbols, docstrings, refs, chunks) and ## CLI (refs & semantic)
Q-08 AUTHOR: Code Intelligence & AST Lead CATEGORY: MCP & Memory Integration REFERENCES (PUBLIC GITHUB): src/mcp_server.ts, src/hooks_compact.ts, src/read_commands.ts
"Could Token-Goat be paired with something like MCP codebase-memory, where codebase-memory maps out the relationships and Token-Goat pulls and compresses just the relevant code before it reaches the LLM?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## MCP server, ### Generic compression and handoffs, and ## CLI (Surgical Reads)
Q-09 AUTHOR: DevOps & Infrastructure Lead CATEGORY: Installation & Enterprise Packaging REFERENCES (PUBLIC GITHUB): docs/install.md, package.json
"How do I install this things, the original instructions don't work, I get the 403. To be solved with SCM/Ticket"
PUBLIC GITHUB README Covered in public GitHub README section: ## Install (Node.js 22.16+ & Standard npm Registry Intake)
Q-10 AUTHOR: Principal Systems Architect CATEGORY: Input Integrity & Interception Safety REFERENCES (PUBLIC GITHUB): src/hooks_read.ts, src/read_commands.ts
"Hooks are intercepting and changing input. What would guarantee that my prompts / tool calls / inputs are not getting worse?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## What changes (Pass-through Guards & Byte Integrity) and ## CLI (token-goat bench Invariant Validation)
Q-11 AUTHOR: Principal Systems Architect CATEGORY: Output Quality & Model Evals REFERENCES (PUBLIC GITHUB): src/parser.ts, tests/command_matrix_e2e.1.test.ts
"It's clear that there are evals on tokens saved. Are there evals on model outputs, e.g. generated code quality, with and w/o token-goat?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## Token savings, measured and ### 6. Context pressure
Q-12 AUTHOR: Compiler & Language Tooling Lead CATEGORY: LSP & Roslyn Comparison REFERENCES (PUBLIC GITHUB): src/parser.ts, src/languages/, src/graph_commands.ts
"token-goat seems to build something like AST / semantic search index over repositories. Is it similar to LSP approach? See https://github.com/dotnet/roslyn/blob/main/docs/roslyn-language-server-copilot-plugin.md"
PUBLIC GITHUB README Covered in public GitHub README sections: ## What gets installed? (Lightweight AST Parser vs. Compiler Servers) and ## Token savings, measured
Q-13 AUTHOR: Junior Developer Cohort CATEGORY: Token Economics & Overhead REFERENCES (PUBLIC GITHUB): src/stats.ts, src/hooks_read.ts, src/tool_filters/dispatch.ts
"token-goat is using more tokens."
PUBLIC GITHUB README Covered in public GitHub README sections: ## The problem (Turn 1 Setup Tax vs Multi-Turn Savings) and ## Token savings, measured
Q-14 AUTHOR: Junior Developer Cohort CATEGORY: Performance & Latency REFERENCES (PUBLIC GITHUB): src/worker.ts, src/parser.ts, src/db.ts
"token-goat is making things slower."
PUBLIC GITHUB README Covered in public GitHub README section: ## Token savings, measured (Hook Cold-Start Latency & Unknown-Event Dispatch)
Q-15 AUTHOR: Staff Security & Governance Architect CATEGORY: PreToolUse Hooks & Input Rewriting REFERENCES (PUBLIC GITHUB): src/hooks_bash.ts, src/hooks_agent_spawn.ts, src/bash_runner.ts
"The PreToolUse hooks don't only filter output — they rewrite input. hooks_bash.ts returns a rewriteInput that wraps the command the agent asked to run, and hooks_agent_spawn.ts rewrites the prompt passed to spawned subagents. What is our acceptance criterion that a rewritten command or prompt is semantically equivalent to what the agent intended, and how would a developer notice a case where it isn't?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## The problem, ## What changes, and ## What gets installed?
Q-16 AUTHOR: Staff Security & Governance Architect CATEGORY: Benchmarks, Output Quality & Evaluation REFERENCES (PUBLIC GITHUB): tests/token_savings_benchmark.test.ts, docs/benchmark-baseline-2026-05-24.md, CLAUDE.arch.md
"The benchmarks in the repo measure tokens saved, and the token-savings baseline doc still cites tests/test_token_savings_benchmark.py — Python files from before the TypeScript rewrite, so that baseline is not reproducible as written. I found no eval measuring model output quality with and without the tool. Before rollout: who owns an A/B on task success rate and generated-code quality, on our repos, and what regression are we willing to accept in exchange for the token savings?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## Token savings, measured (Benchmark Methodology & Source Files) and ## CLI (token-goat bench)
Q-17 AUTHOR: Staff Security & Governance Architect CATEGORY: Skill Interception & Compact Caching REFERENCES (PUBLIC GITHUB): src/hooks_skill.ts, src/skill_cache.ts, token-goat.toml
"Repeat skill loads are intercepted: the second invocation is blocked and a ~400-token cached compact is served instead of the real skill body, allowed to reload only when compaction may have evicted it. That gate is a heuristic about context state. What happens on a wrong guess — the model proceeds on a summary of rules it believes it holds in full — and how is that case detected rather than inferred after the fact?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## Token savings, measured (Benchmark Methodology & Source Files) and ## CLI (token-goat bench)
Q-18 AUTHOR: Staff Security & Governance Architect CATEGORY: Index Exclusions & Silent False Negatives REFERENCES (PUBLIC GITHUB): src/config.ts, docs/security.md, tests/project_config_locked_sections.test.ts
"Their own docs state that a file excluded from the index answers symbol, read and semantic «in the same words a name that never existed does». So the failure mode of our skip lists is a silent false negative, not an error. What exactly are we going to exclude, and how does a developer tell «not indexed» apart from «does not exist» while working?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## What changes (Exact-Match Symbol Lookups & Typo Suggestions) and ## Token savings, measured (Miss Suggestions)
Q-19 AUTHOR: Staff Security & Governance Architect CATEGORY: Index Freshness & Upgrade Lifecycle REFERENCES (PUBLIC GITHUB): src/cli_doctor.ts, src/worker.ts, src/parser_fingerprint.ts
"After an upgrade, the index keeps answering symbol, read, outline and skeleton from symbols extracted by the previous parser build, and nothing says so until someone runs token-goat doctor — which only warns past a quarter. At the current release cadence, who runs doctor and reindexes across the team, and how often?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## Install, ## Zero maintenance, and ## Verify (token-goat doctor)
Q-20 AUTHOR: Staff Security & Governance Architect CATEGORY: Security, SQLite Storage & Data Classification REFERENCES (PUBLIC GITHUB): src/constants.ts, src/db.ts, docs/security.md
"On the Copilot comparison: the index here stores symbols.body — the source text of every indexed symbol — plus docstrings, ref context and chunks, in an unencrypted SQLite outside the repository (~/.local/share, %LOCALAPPDATA%), which their docs state plainly. Which repos are in scope, is that path excluded from OneDrive and from backup, and who signs it off on the data-classification side?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## What gets installed? ("What the index actually holds, in plain terms") and ## Security, privacy, and uninstall
Q-21 AUTHOR: Staff Security & Governance Architect CATEGORY: Air-Gapped Network Policy & Cache Pre-Seeding REFERENCES (PUBLIC GITHUB): src/embed_model.ts, src/image_ocr.ts, docs/security.md
"On network.offline = true: it only gates four paths — the embedding model download from Hugging Face, the OCR language data from a CDN, image fetches the agent initiated, and Drive if already authorized. Nothing else needs the network. So the version I'd push for is pre-seed the two caches at install time, then lock offline: both downloads are verified against a recorded SHA-256 and an exact byte length, and a cache populated by hand is verified the same way. Can we do that instead of trading offline mode against features?"
PUBLIC GITHUB README Covered in public GitHub README section: ## Security, privacy, and uninstall (network.offline = true & Local Model Storage)
Q-22 AUTHOR: Staff Security & Governance Architect CATEGORY: Software Licensing & Legal Compliance REFERENCES (PUBLIC GITHUB): LICENSE, README.md (lines 704-714)
"Licence, before anything else gets installed on more machines: this is PolyForm Noncommercial 1.0.0. The author's Additional Use Grant covers an individual developer on their own machine and explicitly excludes «deploying the software as shared infrastructure across a team or organization», and the README adds that if the employer is the primary beneficiary, a commercial licence applies. Has anyone requested a commercial licence quote from token-goat@dfkhelper.com, and is legal in the loop?"
PUBLIC GITHUB README Covered in public GitHub README section: ## License (PolyForm Noncommercial 1.0.0, Individual Use Grant & Commercial Licensing)
Q-23 AUTHOR: Staff Security & Governance Architect CATEGORY: Supply Chain Integrity & Intake Verification REFERENCES (PUBLIC GITHUB): .github/workflows/publish.yml, tests/guards/ci_action_pins.test.ts, SECURITY.md
"Credit where it's due: this package publishes only from CI with npm provenance, pins every GitHub Action to a full commit SHA (enforced by a test), and refuses to release a commit that isn't an ancestor of main. That is better than most of npm and we should use it: is npm audit signatures on the pinned version, plus reading the commit the attestation points to, part of our intake — or are we trusting the version number?"
PUBLIC GITHUB README Covered in public GitHub README section: ## Security, privacy, and uninstall (Capabilities Auditing & Supply Chain Hygiene)
Q-24 AUTHOR: Staff Security & Governance Architect CATEGORY: Exit Strategy, Deprecation & Vendor Risk REFERENCES (PUBLIC GITHUB): src/cli_install.ts (uninstall), src/install.ts, src/purge.ts
"Following the earlier deprecation point: what is our exit? It writes hooks into ~/.claude/, and with the other bridges into ~/.codex/config.toml, shim scripts and delimited blocks inside AGENTS.md. If the harness vendors close this gap — or the project stops (one maintainer, one npm account, 2.5 months old) — what breaks in our workflow, who cleans that up, and who internally owns the upgrade decision and the incident response if a release ever ships something hostile?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## What gets installed? (Reversible Hook Entries) and ## Security, privacy, and uninstall (Clean Purge & Uninstall)
Q-25 AUTHOR: Enterprise Infrastructure & Artifacts CATEGORY: Internal Artifact Feed & Proxy Authentication REFERENCES (PUBLIC GITHUB): ~/.npmrc, Artifactory Proxy Configuration
"https://artifactory.enterprise.internal/feeds/npm/token-goat/versions -> (403) Unauthorized Request"
PUBLIC GITHUB README Covered in public GitHub README section: ## Install (Package Distribution & Public npm Registry)
Q-26 AUTHOR: Platform Engineer Platform Lead CATEGORY: Packaging, Standalone Binaries & Isolation REFERENCES (PUBLIC GITHUB): package.json, esbuild.config.mjs, Node SEA / Bun Compile
"Could be possible to have a really independent version of token-goat? I mean, idependent of installed frameworks and tools A windows exe, let's say, like an old FatJar or a .net native artifact, just to not interfere with local coding projects"
PUBLIC GITHUB README Covered in public GitHub README sections: ## Install (Single Bundled Artifact dist/token-goat.mjs) and ## What gets installed?
Q-27 AUTHOR: Platform Engineer Platform Lead CATEGORY: Model Provisioning & Corporate Firewall Troubleshooting REFERENCES (PUBLIC GITHUB): src/embed_model.ts, src/config.ts, token-goat.toml
"I received some reports about repetitive installation requests of local models from huggingface, I think , related with semantic analysis. Is it documented how to proceed manually? Just to say that I'm not in local LLM run so a brief info or an extended prompt for this will be very helpfull"
PUBLIC GITHUB README Covered in public GitHub README sections: ## What gets installed? (Data Directory Model Storage) and ## Security, privacy, and uninstall (network.offline = true)
Q-28 AUTHOR: Platform Engineer Platform Lead CATEGORY: Metrics Transparency & Tokenizer Estimation REFERENCES (PUBLIC GITHUB): src/stats.ts (savedTokensFromBytes), src/cli_stats.ts, token-goat stats --methodology
"After some research run by AI, reported token measures are just a static calculation based in bytes not sent through context. As this is very LLM related and as it could lead to missunderstandings, could be possible to avoid it or shown in a "more details" option?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## Token savings, measured (Byte Withheld Metrics) and ## CLI (token-goat bench)
Q-29 AUTHOR: Platform Engineer Platform Lead CATEGORY: Context Trimming vs. API Roundtrips & Latency REFERENCES (PUBLIC GITHUB): src/hooks_read.ts, src/code_fold.ts, docs/cli.md
"As I stated above, I'm not too deep in copilot/claude internals, so maybe someone above asked the right same question but in a more technical way, but here is mine, more pragmatic: what gives to me confidence that a context trimmed by this tool is not forcing more roundtrips to LLM? At least, I'm finding a high slowness in Luna queries as tool is queried several times"
PUBLIC GITHUB README Covered in public GitHub README sections: ## The problem, ## What changes, and ## Token savings, measured (Break-Even Calculations)
Q-30 AUTHOR: Cross-Platform Systems Engineer CATEGORY: Token Economics & Single-Turn Overhead REFERENCES (PUBLIC GITHUB): src/install.ts, src/hooks_session_start.ts, tests/token_savings_benchmark.test.ts
"The answer so far seems to suggest token-goat has an initial token overhead in the first turn that pays for itself in turn 2+ ? Is this true? Using something like Opus 5 for 1 single turn using architect to breakdown a ticket for implementation from an issue tracker would be more expensive with token-goat? I am also noticing a lot of calls to token-goat semantic which are not free."
PUBLIC GITHUB README Covered in public GitHub README sections: ## The problem and ## What changes (Instruction Routing Guide & Tool Invocations)
Q-31 AUTHOR: Staff Security & Governance Architect CATEGORY: Cross-Platform Support & macOS Validation REFERENCES (PUBLIC GITHUB): .github/workflows/ci.yml, src/constants.ts, src/bridges/copilot_cli_install.ts
"Does token-Goat work on macOS?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## Install ("Node.js 22.16 or later, on any platform") and ### Stats display — macOS
Q-32 AUTHOR: Platform & DevEx Architect CATEGORY: Cross-Platform Support & WSL2 Architecture REFERENCES (PUBLIC GITHUB): .github/workflows/ci.yml, src/paths.ts, src/bash_extractors.ts, src/path_containment.ts
"Is it feasible to use token-goat on WSL?"
PUBLIC GITHUB README Covered in public GitHub README sections: ## Install ("Node.js 22.16 or later, on any platform") and ## What changes

Enterprise Configuration Review (Security Policy Inquiry)

Below is the evaluated configuration file analyzed against corporate security, zero-trust filesystem isolation, and air-gapped network policies.

01# Corporate Developer Baseline Configuration: token-goat.toml
02[mcp]
03confine_reads_to_project_root = true # APPROVED: Prevents traversal attacks (../../) outside repo
04allowed_roots = ["C:/dev"] # APPROVED: Pins authorized developer root boundaries
05
06[indexing]
07cross_project_symbols = false # APPROVED: Confines symbol search to active repository
08
09[redaction]
10strict = true # ADVISORY: Redacts high-entropy strings; watch for git SHA false-positives
11custom_patterns = [] # Insert internal company-specific credential prefixes here
12
13[screenshot]
14block_private_targets = true # APPROVED: Blocks screenshots of 192.168.x.x, 10.x.x.x, 127.0.0.1
15
16[gdrive]
17enabled = false # APPROVED: Eliminates 3rd-party cloud surface completely
18
19[network]
20offline = true # CAUTION: Run 'token-goat doctor' once first to cache bge-small model
Operational Advisory on offline = true
If offline = true is enforced before the local embedding model (bge-small-en-v1.5, ~35 MB) has been cached to ~/.token-goat/models/, semantic vector search is disabled. Run token-goat doctor once while connected to the company registry before locking down network egress.

Architectural Positioning: Token-Goat vs. Roslyn LSP vs. VS Code Index

Technical breakdown distinguishing compiler-hosted language servers, editor workspace indexes, and runtime context governors.

The Core Distinction: Roslyn is a compiler diagnostic engine for the IDE; the VS Code Index is a file discovery engine for the editor; Token-Goat is a runtime context and token governor for the LLM prompt budget. They operate at completely different layers of the developer stack and are complementary, but Token-Goat solves the critical token waste that the other two completely ignore.
Three-Tier Developer Tooling Architecture
TIER 1: IDE EDITOR Roslyn LSP
Language Compiler Engine

Compiles source into semantic symbol graphs. Powers real-time red squigglies, type autocomplete, and IDE refactoring. Single-language (C# only), heavy memory (1.5 GB), and requires a valid compiling project.

TIER 2: DISCOVERY VS Code Index
Workspace Search Engine

Generates whole-file lexical and vector indexes to locate files when you ask @workspace broad questions. Finds which files to inspect, but dumps entire un-pruned files into the LLM prompt.

TIER 3: RUNTIME GOVERNOR Token-Goat
Context & Token Budget Gatekeeper

Intercepts OS tool execution at runtime. Slices exact AST method symbols (140 tokens instead of 6,000), deduplicates re-reads, and folds verbose test runner output. Sits directly between the agent and the LLM API.

Lifecycle Comparison: Debugging a C# Payment Bug
WITHOUT TOKEN-GOAT (Roslyn + VS Code Index Alone)
1. File Discovery: VS Code Index finds PaymentService.cs and dumps all 1,200 lines (6,000 tokens).
2. Terminal Execution: Agent runs dotnet test. Terminal dumps 15 test suites and MSBuild build configuration logs (2,500 tokens).
3. Caller Verification: On Turn 3, the agent re-reads PaymentService.cs to check helper methods (another 6,000 tokens).
Total Ingested Context: 14,500 tokens
Prefill Delay: ~12 seconds | Cost: $0.22 per turn
WITH TOKEN-GOAT (Active Runtime Governance)
1. Surgical Read: Token-Goat intercepts the file read and serves only the 24-line ProcessRefund AST slice (140 tokens).
2. Stream Folding: Token-Goat intercepts dotnet test, folds 14 passing suites, and serves only the failing stack trace (90 tokens).
3. Read Dedup: On Turn 3, Token-Goat's hash hook detects identical file content and serves a cached pointer (18 tokens).
Total Ingested Context: 248 tokens (98.3% reduction)
Prefill Delay: ~1.5 seconds | Cost: $0.004 per turn
Dimension Roslyn LSP (.NET) VS Code Copilot Index Token-Goat
Scope & Mission Compiler diagnostics & code completion in IDE Semantic file discovery for @workspace Runtime token & context budget governor
Language Support C# and Visual Basic only Polyglot across active VS Code workspace Polyglot across 80+ languages via Tree-sitter
Compilation Dependency Requires valid .sln / .csproj build Zero compilation (text & lexical indexing) Zero compilation (AST grammar parsing)
Memory Footprint Heavy (350 MB - 1.5 GB per instance) Integrated into VS Code Electron process Lightweight (a local SQLite index, no resident server)
Read Interception None (IDE editor only) None (dumps entire file into prompt) Active hash deduplication & body folding
Terminal Stream Filtering None None Folds passing test suites, preserves stack traces
Harness Support IDE editor only VS Code only Claude Code, Codex, Copilot CLI, Cursor, MCP
Where Token-Goat is Decisively Better
1. Sub-File AST Slicing (97.7% Savings)

Roslyn and VS Code Index have zero concept of sub-file surgical budgeting. If an agent needs one method from a 2,000-line controller, VS Code dumps all 10,000 tokens. Token-Goat slices exact method and class boundaries (read "file::symbol"), keeping prompt context tiny.

2. Terminal Execution Stream Noise Folding

Roslyn and VS Code Index operate exclusively inside the editor text buffer. They are completely blind to terminal execution. When npm test or dotnet test emits 2,500 tokens of passing green lines, Token-Goat folds it to 90 tokens, preserving only the actionable failure.

3. Read Deduplication & Served-Line Elision

Coding agents frequently re-read the same file multiple times across a session. VS Code and Roslyn do not intercept or cache reads. Token-Goat tracks file content hashes and served line runs, replacing redundant reads with 15-token cached hash pointers.

4. Fault-Tolerant Polyglot (Zero Build Required)

Roslyn is tightly coupled to the .NET compiler and dies if project references are broken or if the solution fails to build. Token-Goat uses fault-tolerant Tree-sitter AST grammars across 80+ languages (TypeScript, Python, C#, Rust, Go, Apex, SQL), parsing even broken code without compiling.

5. Universal Multi-Agent Portability

Roslyn only works inside C# IDEs; VS Code Index only works inside VS Code. Token-Goat is harness-agnostic. It runs seamlessly with Claude Code, Codex, GitHub Copilot CLI, Cursor, and any MCP-compatible client via CLI hooks and stdio JSON-RPC.

6. Image Shrinking & Secret Redaction

Token-Goat automatically downscales high-resolution screenshots and diagram buffers before transmission (capping them at 1,568 tokens, and well below that for most screenshots) and sanitizes API keys and bearer tokens through high-entropy pre-tool redaction.

Dimension Roslyn LSP (.NET) VS Code Copilot Index Token-Goat
Primary Architectural Role Compiler diagnostics, syntax trees & type autocomplete in IDE Coarse repository file discovery for @workspace chat Active context governor & token budget gatekeeper
Context Reduction Impact 0% (does not manage LLM prompt tokens) 0% (dumps un-pruned files into context) Measured (see Q-11's paired evaluation: 64.0% median billing-weighted token saving, n=6 clean pairs)
Sub-File AST Slicing Internal to compiler only (not exposed as surgical prompt slice) No (retrieves and serves entire files) Yes (extracts single methods, classes, and sections)
Terminal Noise Compression None (unaware of terminal commands) None (unaware of terminal commands) Yes (folds passing test suites, keeps stack traces)
Read Deduplication None None (re-reads dump duplicate tokens) Yes (SHA hash checks block redundant reads before they hit the model)
Language Support C# and Visual Basic only Polyglot across active VS Code workspace Polyglot across 80+ languages via Tree-sitter grammars
Build Prerequisite Requires successful .sln / .csproj compilation Zero compilation (text and lexical indexing) Zero compilation (fault-tolerant AST syntax parsing)
RAM & System Overhead Heavy (hundreds of MB to over 1 GB per Roslyn instance) Embedded in VS Code Electron process A local SQLite index plus one background reindex worker; CLI calls are short-lived processes
Supported Agent Harnesses Visual Studio / VS Code IDE only VS Code only Universal (Claude Code, Codex, Copilot CLI, Cursor, MCP)
The Symbiotic Workflow: How All Three Work Together

Token-Goat is not intended to replace Roslyn or the VS Code Index. The optimal engineering workflow combines all three without overlapping:

  1. Discovery (VS Code Index): When you ask an agent "Where is refund processing handled?", the VS Code Index searches the workspace and identifies src/Billing/PaymentService.cs.
  2. Diagnostic Precision (Roslyn LSP): The C# compiler flags error CS0103: The name 'RefundValidator' does not exist in the current context on line 412.
  3. Context Governance (Token-Goat): Instead of letting the agent dump all 1,200 lines of PaymentService.cs and 2,500 lines of terminal logs into the prompt, Token-Goat extracts only the 24-line ProcessRefund AST slice and folds the build stream to the exact compiler line.
  4. Synthesis (LLM): The agent's model fixes the missing import from that 24-line slice instead of the full 1,200-line file, with the diff bounded to the lines that actually changed.