token max/ evidenceInstallation & guide ↗

IMPLEMENTATION REVIEW · RESEARCH · REPRODUCIBLE MEASUREMENTS

Smaller context.
Show the evidence.

The collector can make an audit cheaper to read. Whole-task savings and preserved model quality are still unmeasured.

Token Max has no runtime hooks. Its plugin manifest packages an on-demand Codex skill and local Python collector. The command writes reports only; implementing any proposal is separate work. It never intercepts normal tool calls, removes conversation history, changes model routing, or compresses prompts on their way to a model. This review covers this repository’s mechanisms; it does not audit every plugin or hook installed on a user’s computer.

The demonstrated result is narrower: compact metadata needs fewer text tokens than reading a large, verbose synthetic log in full. Small inputs can get larger. Metadata correctness is tested; downstream coding success has not been evaluated in paired model runs. No percentage below is a promised reduction in API bills, subscription usage or future sessions.

Decision: use the report to select evidence worth reading. Keep the original log available. A size ranking cannot replace a transcript or establish that a task passed.

Hooks & mechanisms: what actually runs

MechanismImplementation and possible benefitCost or limit
Skill installation.codex-plugin/plugin.json packages skills/token-audit. The standalone installer links that same skill into ~/.agents/skills; do not install duplicate copies.No hook registration or automatic optimization. Discovery metadata and invoked instructions consume context.
Local collectionaudit.py parses logs and inventories files. Python counts, hashes and sorts instead of asking a model to read every body.The collector makes no model calls. Agent invocation and interpretation still consume tokens.
Compact evidencesession.py and report.py export usage, sizes, fingerprints and line references.Lossy metadata extraction. Bodies stay local; diagnostic references survive size ranking. Fingerprint-checked retrieval recovers relevant evidence.
Progressive readingSKILL.md routes usage questions to --usage-only, then compact summaries and bounded retrieval for deeper reviews.Procedural guidance, not an enforced retrieval budget. Loading all references every time adds overhead.
Instruction/code reviewproject.py locates large files, Python function spans, repeated instruction lines and routing candidates.It does not rewrite files or prove waste. Presence on disk does not establish active context.
Plugin and routing adviceThe rubric compares overhead with displaced work on the same task.No plugin installation or executable model router. A cheaper model may use more tokens or require retries.
GitHub ActionsThe Pages workflow tests and publishes site/.A deployment workflow, not an agent token hook. Private evidence directories are outside its artifact.

Source entry points: installer, collector, and rubric.

Measured improvements, with overhead included

Version 0.2 implements the earlier review's findings: shorter default reports, a usage-only path, diagnostic references independent of size, and verified bounded retrieval. The current results and preserved pre-polish baseline use identical synthetic log fixtures and tiktoken 0.14.0 / o200k_base. Fixture hashes match. This named encoding is not asserted to be the active Codex model's tokenizer.

FixtureRaw JSONLPrevious summaryCompact summaryCompact + query + error retrieval resultsUsage-only CLI result
tiny26262720640677
verbose18,25362920840877
repeated29,55779926446477
crowded19,62181626746777

Units: text tokens. Tiny has one routine output line; verbose has 2,000 lines; repeated has eight identical 400-line outputs; crowded has twenty 100-line outputs. Every fixture also has one short critical error, duplicate usage snapshots and a two-file project. These intentionally repetitive inputs are controlled examples, not a representative production sample.

The compact summaries are about 67% smaller than their previous summaries on the same fixtures. The tiny summary now fits below the raw log's text size, while adding the error-retrieval result makes it larger again. A simpler usage-only response costs 77 text tokens here and skips project scanning/report writes. A bare projection of just the known counters is 40 tokens but excludes provenance and warnings, so it answers an even narrower question.

FixtureSkill + compact summarySkill + both references + compact summary
tiny7573,500
verbose7593,502
repeated8153,558
crowded8183,561

Loading the skill still outweighs a tiny raw log. Do not load both references for a simple counter question. Composite columns tokenize concatenated text; reference loading is conditional. Bodies omitted from the summary are not magically free: retrieving them requires more context and a tool round trip.

The query-and-retrieval column includes the bounded diagnostic selector output, returned excerpt and metadata. It excludes command arguments, tool schemas and model reasoning. No column includes message framing, history replay, follow-up exploration, retries or validation. Raw-log size excludes original project source, while summaries include some project metadata. These are reading-strategy comparisons, not lossless compression or proof of whole-task savings. Source and fixture hashes identify the measured inputs. The current result can be reproduced from this checkout. The prior result is a preserved measurement artifact; its older source snapshot is not bundled here.

Quality: retained facts and lost information

Eight assertions pass in each of four fixtures: provider counters survive; cached input is subtracted correctly; duplicate snapshots are not added; output counts survive; the largest output’s byte count is correct; repeated calls have the expected count; the short critical error retains a diagnostic reference; and bounded retrieval recovers its exact text. These are 32 deterministic evidence checks, not solved coding tasks.

The negative control is a short ERROR: migration checksum mismatch observation. Its body stays out of automatic summaries for privacy and size. In the crowded fixture it falls outside the top-15 size ranking, but its independent diagnostic reference survives and its original text is recovered exactly. The previous implementation lost that reference. Recovery of this known fact is now tested; general model task-quality preservation remains unmeasured.

AvailableOmitted or unknownFollow-up
Cumulative counters, sizes, selected line referencesPer-file causal cost, exact model-visible framing and replayCompare provider usage deltas on matched tasks.
Largest outputs, repeated calls, all detected diagnostic references and recognized exit codesBodies and semantic failure/success judgments; unrecognized failuresInspect original excerpts and acceptance results before deleting work.
Current inventory and instruction candidatesHistorical revision, instruction activation, non-Python function spansVerify revision, scope and code dependencies.

JSON retains the largest 15 outputs, 10 repeated-call groups and 10 messages, plus every detected diagnostic reference regardless of size. Default Markdown displays three outputs, three repeated groups and three diagnostic candidates, with a count and pointer when more diagnostics exist. Expanded tables are opt-in with --full-report. These row limits are not a hard token budget. Project JSON has the scanned inventory, not a full transcript. Unsupported multimodal content is not measured as text.

Preserving quality requires retrievable source evidence, checking failures independently of size, and retaining necessary tests, constraints and project knowledge. Tests cover short failures below the size cutoff, nonzero exits after success banners, structured outputs, tail errors outside a returned excerpt, stale fingerprints, output bounds and unchanged source/configuration files. Error-word detection can flag benign examples or miss unusual failures. Absence of a diagnostic never proves success. The report-only agent policy is explicit guidance, not an OS sandbox.

Source review findings and correction

An independent read-only review inspected the installer, parser, accounting, tests and claims. It found a reproducible defect: --out PROJECT/reports allowed later runs to inventory their own earlier reports. A one-file fixture grew to three inventoried files; generated evidence outranked actual source. This could misdirect recommendations and contaminate comparisons.

Fixed: resolve the output location before scanning and exclude its own evidence.md, evidence.json and AUDIT.md. Unrelated files in that directory remain visible, including when --out is the project root. A regression test repeats the audit three times at both locations and verifies stable source inventories. Other, previously chosen output locations are not inferred; keep those ignored or outside comparison scope.

Accounting tests confirm snapshots are not summed and cached input/reasoning are not counted twice. Missing counters stay unknown. Decreases, malformed lines and fork overlap receive warnings. The last valid snapshot means last in file order; the parser does not reconcile out-of-order events into a billing ledger. Logs do not establish invoice totals or subscription-quota usage.

Validation: 39 automated tests passed with the optional tokenizer installed, including benchmark/error-recovery controls, report-only boundaries and the output-location regression. The plugin manifest and skill validators also pass. A separate agent completed an isolated synthetic report-only audit, recognized a small late failure, avoided unsupported token-saving claims and left source/instruction files unchanged. That exercise also exposed a broad JSON read; the new bounded query helper addresses that overhead. Its row limits and field projection are tested, but one agent exercise is not a paired quality study. Standard-library runs skip tokenizer-dependent checks. These tests verify contracts, not general absence of defects or end-to-end savings.

Privacy: bodies are processed locally and omitted from automatic evidence reports. Explicit retrieval returns a chosen excerpt to the local terminal; keep that excerpt private. Paths, model/tool names, hashes and counts remain metadata and can still be sensitive. Filename exclusions are not a universal secret detector. Do not publish real reports or send private logs to web services. The downloadable benchmark is synthetic.

Papers and engineering blogs

These original papers and first-party write-ups support mechanisms and experimental design. None evaluates Token Max. Each entry distinguishes implementation from recommendations and unimplemented techniques.

Local filtering and tool discovery

Anthropic’s Code execution with MCP describes processing large results in code and loading relevant tools on demand. Its illustrated 150,000-to-2,000-token scenario demonstrates a mechanism, not a transferable savings guarantee.

Applied: Python extracts counts and candidate metadata before model inspection. Not implemented: an MCP execution sandbox, dynamic tool discovery or a tool-output hook. Plugin recommendations must compare schemas, instructions, results and extra calls with the work displaced.

Observation masking versus model summarization

The Complexity Trap (Lindenbauer et al., version 3, 2025) compares context strategies in SWE-agent on SWE-bench Verified across five model configurations, with additional OpenHands evidence. It reports observation masking halving cost relative to the raw agent while matching or sometimes exceeding model summarization’s solve rate in its experiments.

Applied as evaluation guidance: test a simple deterministic baseline before adding model calls, and measure cost together with solve rate. Not implemented: live history masking. Selecting metadata and retrieving evidence for an audit does not reproduce the paper’s agent-history intervention or establish its cost ratio here.

Relevance, position and retrieval

Lost in the Middle (Liu et al., TACL 2024) evaluates multi-document question answering and key-value retrieval. Performance varies with information position, including degradation when relevant information lies in the middle. It does not show every short prompt is better or test current Codex coding tasks.

Anthropic’s Effective context engineering for AI agents recommends just-in-time retrieval and retained identifiers for lookup, and warns that aggressive compaction can lose subtle context. This is engineering guidance, not a controlled Token Max experiment.

Implemented: compact default reports, a usage-only path, independent diagnostic references, retained local logs and bounded field selection and fingerprint-checked retrieval. The short-error recovery test directly checks a failure of size-only ranking. The default skill loads review/source references only when needed.

Learned prompt compression

LLMLingua-2 (Pan et al., ACL Findings 2024) learns extractive compression through token classification. Its in-domain and out-of-domain evaluations include MeetingBank, LongBench, ZeroScrolls, GSM8K and BBH, addressing downstream behavior as well as compression.

Applied as a requirement: test task quality before adopting a compressor. Not installed or implemented: LLMLingua or any learned compression model. Deleting arbitrary code tokens, paths or constraints is not equivalent to that algorithm.

Instruction files can cause extra work

Evaluating AGENTS.md (Gloaguen et al., version 1, 2026) studies repository context, task success and agent behavior. It finds extra exploration and increased inference costs in its tested settings; success effects vary by condition and instruction source. Human-written instructions are not uniformly harmful.

Applied: review requirements that cause unnecessary work, not just file length. Keep relevant constraints and test commands. Move infrequent procedures behind references only when that actually reduces loading frequency. Removing required checks to lower token counts fails the quality requirement.

Symbols, file boundaries and caching

Aider’s repository map selects relevant symbols within a token budget. Applied as advice: try targeted search or symbol navigation before splitting a large file. Token Max measures Python function spans; it does not implement Aider’s graph ranking. A split helps only when equivalent tasks retrieve less without missing dependencies.

OpenAI’s prompt-caching documentation describes prefix reuse. Applied: separate cached/uncached input and record cache conditions. Cache hits can affect processing cost or latency while tokens remain in context. They do not establish a shorter prompt or subscription savings.

The source ledger also links pinned Codex protocol sources, instruction discovery and the Tokenoscope reference revision. Verify current product behavior before changing client configuration.

Establishing whole-task savings and quality

This next evaluation has not been performed. Choose representative tasks with clear acceptance checks before examining optimization results. Include small tasks, verbose exploration, failures/recovery and tasks dependent on repository-specific instructions.

  1. Freeze the comparison. Use the same initial revision, task, model/version, reasoning setting, tools, permissions and acceptance criteria. Start independent sessions, change one mechanism at a time, and randomize or alternate run order. Record cache conditions rather than assuming equal cache hits.
  2. Record complete outcomes. Capture start/end counters, input, cached input, output, reasoning, latency, retries, clarifications and validation. Do not subtract across resets or add overlapping forks. Never replay side-effecting logged commands as a benchmark.
  3. Charge overhead. Include audit calls, follow-up retrieval, schemas/instructions, compressor model calls and required validation. Report setup separately and amortize only across explicitly observed tasks. Tokens, money and latency are different outcomes.
  4. Judge quality independently. Use unchanged acceptance tests, behavior checks and human review for requirements tests miss. Include failures and incomplete tasks. Excluding failed attempts can make a cheap-but-unreliable strategy appear better.
  5. Repeat paired trials. Report per-task differences, distributions and uncertainty, not only a best run. Predefine the acceptable quality difference. A small pilot is feasibility evidence; it cannot establish broad non-inferiority.
task, revision, model, settings, arm, run
input_tokens, cached_input_tokens, output_tokens, reasoning_output_tokens
elapsed_seconds, retries, acceptance_pass, review_result, missing_evidence

Use total = input + output when consistent with provider fields; cached input and reasoning are subsets. Also report success rate and tokens across all attempts divided by accepted completions; the latter is undefined with zero completions. Do not infer dollars without applicable pricing/cache rules.

An instruction rewrite immediately supports only tokens(before) − tokens(after) per inclusion under the same encoding. Future inclusion frequency and behavior require observation. A smaller file or removed routing table cannot establish those quantities.

Before adding an automatic hook

No hook was installed during this review. A future output-filtering hook should pass the paired evaluation above. Preserve command identity, exit status, failure diagnostics, truncation indicators and a local reference to omitted output. Keep result schemas valid and allow full retrieval.

Begin with a tool-specific deterministic filter, not a universal first-N-lines rule. Test errors at the beginning, middle and end; misleading success banners followed by failures; Unicode; structured/binary results; and filter failures. Return unfiltered results for unsupported formats or processing errors, and measure that fallback’s cost.

Disable or revise a hook that hides a required fact, changes meaning, lowers the agreed acceptance rate or increases total tokens after retries. A published compression percentage is insufficient reason to enable it.

Reproduce the review

Ordinary collection/tests need Python 3.10+ and Git (or ripgrep for inventory fallback). Exact text-token benchmarking also needs the optional tokenizer. The recorded run used an isolated environment with tiktoken==0.14.0. Installation and initial encoding retrieval may require network access; fixture processing makes no model calls.

python3 -m venv .venv
.venv/bin/python -m pip install tiktoken==0.14.0
.venv/bin/python scripts/benchmark.py --output .token-audit-benchmark.json
.venv/bin/python -m unittest discover -s tests -q
python3 scripts/check_site.py

Results include fixture hashes, tokenizer version, source hash, sizes, fidelity checks and information-loss controls. Compare with the recorded synthetic result. Regenerate it and this table when collector text or instructions change. A standard-library test checks the published source fingerprint to catch stale results; tokenizer-enabled tests verify the measurements and recovery assertions. No real transcript or credentials are in the fixture.

For real projects, use the local collector and keep reports private. Use comparable before/after task runs for stronger claims. Running repeated full audits is itself work with a token cost.