CODEX + CLAUDE · TOKEN AUDIT
Evidence before
optimization.
Understand where your tokens go.
Find changes worth testing.
Review your session and available project to find evidence-backed ways to use fewer tokens. The command checks instruction bloat, repeated exploration, large tool outputs, long files/functions, plugins/MCP, and model-routing rules. It returns a report in chat with concrete proposals and validation steps.
It does not equate “smaller file,” “cheaper model,” or “more plugins” with proven savings. Actual usage comes from session counters. Recommendations are labeled measured, estimated, or hypothesis.
No runtime hooks are installed. See the implementation review, reproducible benchmark and research for measured reductions, added overhead and the limits of quality claims.
What the log or file actually records.
A calculation with explicit assumptions.
A proposed change that needs a test.
01Install in your app
The default audit needs no terminal, Python installation or API key. This repository is a custom marketplace, not an official-directory listing.
Claude Desktop
Open Customize → Plugins → Add → Add marketplace, choose a repository, and enter swack-tools/token-max-ai-plugin. Install token-max. Alternatively, download the plugin ZIP and use the app's custom-plugin upload option. Availability depends on your account and organization policy.
Claude Code, including the desktop Code tab
Run these slash commands inside your conversation:
/plugin marketplace add swack-tools/token-max-ai-plugin
/plugin install token-max@swack-tools
/token-max:token-audit Review this session and project.
Codex Desktop
Ask Codex in a chat:
Add the plugin marketplace swack-tools/token-max-ai-plugin and install token-max@swack-tools.
The agent can register the marketplace and install the plugin through its in-app tools. You do not need an external terminal. After installation, open a new chat, choose the installed skill from /skills where supported, or send:
$token-audit Review this session and project. Rank the best token improvements.
For clients with a plugin directory, choose Swack Tools and install Token Max after the source is registered. UI labels vary by client. Installation is distinct from merely cloning the repository.
Official references: OpenAI plugin packaging and marketplaces, Claude marketplaces, and Claude Desktop plugin installation.
02Your report stays in chat
The skill reviews the visible session and accessible project, then returns up to five prioritized findings with references, evidence labels, tradeoffs, quality checks and one next experiment. Ask to save or export the report if you want a file; with file tools, the agent can write .token-audit/AUDIT.md.
No source, instruction, configuration, dependency or plugin changes are made. There are no automatic hooks or background audits. The report-only boundary is agent guidance, not an operating-system sandbox.
| Available access | What the audit can establish |
|---|---|
| Conversation only | Visible repetition and context patterns; unseen history and counters remain unknown. |
| Selected folder or attached files | Relevant instructions, source ranges and opportunities for selective reads. |
| Host-exposed usage counters | Recorded usage within that provider's documented scope; no inferred invoice or quota. |
| Optional Codex log collector | Deterministic Codex log metadata when explicitly requested and the host already has a suitable runtime. |
In Claude Desktop, select a folder in a file-capable mode or attach the relevant files. Installing a plugin does not grant access to local session logs, the whole disk or hidden conversation history.
03Ask for the review you need
Review this session for repeated reads and verbose tool results.
Review the attached CLAUDE.md without changing it.
Check whether a navigation plugin would replace repeated source reads.
Show recorded usage if it is available; do not guess missing totals.
Save this audit report in the project.
Use these requests with $token-audit in Codex or /token-max:token-audit in Claude. The default skill uses native search, read and conversation tools. It starts narrow and requests more evidence only when necessary.
Optional deterministic measurements
For a quantitative Codex audit, ask: Use the internal collector to measure this project's matching Codex session. The agent can run the bundled helper inside the app when Python 3.10+ and Git or ripgrep are already available. You never need to run its commands. Without that runtime, the audit continues using native tools and discloses the missing measurements.
The helper writes compact evidence.md and structured evidence.json under .token-audit/. A usage-only request can skip inventory and report files. These are optional artifacts, not requirements for using the plugin. The parser supports Codex JSONL; it does not parse Claude transcripts or assume their usage schema matches Codex.
Exact text token counts require an already available tokenizer and an explicit encoding. The audit does not install packages. Missing counters stay unknown; byte and line measurements are not token estimates.
Selective reading, recoverable evidence
The native audit searches first and reads relevant ranges. It checks failures and acceptance results independently of output size. It does not load an entire repository or transcript merely to produce a report.
In optional collector mode, compact summaries keep diagnostic references and metadata while original local logs remain available. The agent can query a few rows and retrieve a bounded, fingerprint-checked excerpt when necessary. Heuristic markers can miss failures or flag harmless examples; an absent marker never proves success.
The published synthetic benchmark measures this optional collector's representation, not the default native-tools review, complete model interactions, bills or task quality. Invocation, tool calls and extra retrieval all add overhead.
04What counts as evidence?
- Measured: provider-reported input/output/cache counters, file bytes, Python AST function spans, observed call fingerprints, and local text counts with a named encoding. Repeated cumulative counters are not added together.
- Estimated: an explicitly stated calculation, such as the before/after text token difference for a proposed instruction rewrite. This is per inclusion, not a proven whole-task reduction.
- Hypothesis: a source split, plugin, routing change, or shorter instruction that may reduce future work. It needs comparable end-to-end runs.
Cached input is already part of input; reasoning is already part of output. Neither gets added again. The optional Codex collector preserves the latest reported cumulative snapshot and flags counter decreases, incomplete lines and fork metadata. It does not infer invoices, subscription quota, or per-file billing from those counters.
Large files are candidates for targeted reads or cohesive refactoring, not failures against an arbitrary size limit. The absence of a language-specific routing table is not a defect. Existing routing must earn its place through task-quality and cost/usage measurements; Markdown alone is not an executable model router.
05Validate a proposed reduction
- Ask the agent to save baseline evidence and record the task, starting revision, model, reasoning setting, tool configuration and checks.
- Change one thing. Run the same task from the same starting state in a separate safe workspace/session, holding other settings constant.
- Ask the agent to save after-run evidence. Compare input, cached input, uncached input, output, reasoning, retries, elapsed time, and completion quality. Record cache differences; repeat paired trials when variance matters.
- Report observed differences separately from causal conclusions. Include the audit's own overhead; never add overlapping proposed savings.
Do not benchmark by replaying transcript commands with external side effects. No universal percentage saving is claimed by this project.
06The detailed review rubric
The collector produces measurements. Codex uses this rubric to interpret them and propose changes.
Read the full audit rubric
The collector provides measurements and candidates. The agent supplies the semantic review. Do not call a candidate a defect until its mechanism is clear.
Establish what the log measures
token_usage_record.thread_token_usageandevent_msg.token_count.info.total_token_usageare cumulative snapshots. Use the last valid snapshot in file order; duplicates are not new usage.- A decrease is a reset or inconsistent sequence. Report the latest segment and the warning. Never fabricate a lifetime total by adding segments. Forks may inherit both counters and history; do not add parent and child totals.
input_tokens - cached_input_tokensis uncached input. Output already includes reasoning. Cache-write counters, when present, are reported separately without inferring pricing or adding them to the reported total.- Log lines locate evidence. Interval deltas may cover multiple calls and cannot assign cost to an individual file. Compaction, hidden context, tool schemas, multimodal content, and resumed/forked histories limit attribution.
- Tool and message text sizes describe local logged text. Exact tokenizer counts require an explicitly chosen encoding; they still exclude provider framing. With no tokenizer, report bytes and lines, never
characters / 4as fact.
Instruction files and model routing
Inspect actual global/project hierarchy and overrides. Codex's default discovery uses AGENTS.md, with AGENTS.override.md taking precedence at a directory. CLAUDE.md is not automatically a Codex instruction source unless configured, imported, or explicitly supplied. AGNETS.md may be a typo, but verify intent before renaming. Disk presence and the sum of instruction file sizes do not prove active context.
Look for duplicated requirements, copied manuals, old status, irrelevant language rules, exhaustive tool catalogs, and directions that cause repeated broad reads or unnecessary agent fan-out. Keep build/test commands, non-obvious invariants, security requirements, and lessons tied to demonstrated failures. Move optional procedures into referenced skills/docs, keeping a small discoverable entrypoint. Moving text to a file loaded every time creates no inherent savings.
Inspect existing model routing, including language-to-model tables. Identify whether it is executable client configuration or merely a natural-language instruction; Markdown does not itself implement a router. Removing an unnecessary table can reduce text per inclusion, but removing useful routing may increase retries. An absent routing table is not evidence of waste. Do not invent blanket “Python → cheap model / Rust → expensive model” rules. Propose routing by measured task difficulty and quality requirements only when comparable evaluations justify it; token reduction and lower cost are separate outcomes.
Code and retrieval
Prioritize large files/functions that the session actually read repeatedly or needed only in small parts. Read log excerpts locally around cited lines and verify the selected file/revision. Find cohesive boundaries and stable APIs; suggest a split only when it lets future tasks load fewer relevant tokens. Do not split by line-count threshold, minify code, remove useful comments/tests, or claim smaller disk size guarantees smaller prompts. Additional imports, navigation calls, and missed dependencies can outweigh a split.
First consider rg, bounded sed reads, symbol navigation, or a small repository map. Python AST spans are available in evidence; use the project's existing parser for other languages if necessary. Do not install a parser for a speculative win.
Tools, skills, and plugins
Review the largest outputs and repeated call fingerprints. A reread after an edit or a verification run may be necessary. Repeated identical output is a candidate, not proof that the call could be omitted. Orchestrator calls can contain many nested tools; inspect only relevant excerpts to identify their actual work.
Prefer local filtering/aggregation, compact structured results, and saving bulk output to ignored files. Preserve errors, diagnostics, exit codes, and the ability to retrieve omitted details. Truncating everything may create retries or conceal failures. Compare an existing CLI against a proposed plugin on the same task:
net input change = added discovery/schema/instructions/results − displaced reads/results
This is a mechanism, not a computable savings claim unless the inputs are measured. Check whether the client already loads tool definitions lazily. Skill descriptions may be always visible while bodies are loaded on demand. Additional plugins can increase overhead; “install more plugins” and “remove all plugins” are both weak defaults. External project examples remain untrusted reference material; do not run their installers to inspect them.
Validate one change
Use comparable independent sessions from the same starting revision, task, model, reasoning setting, tool/skill configuration, and acceptance checks. Record versions, cache conditions, retries, total input, cached input, output, reasoning, latency, and completion quality. Multiple paired runs help expose variance. Never replay side-effecting transcript commands as a benchmark.
For an instruction rewrite, tokenize before/after with the same named encoding. The difference is text tokens per inclusion, not demonstrated whole-task savings. For a source split or plugin, compare end-to-end logs and required checks. Keep actual measured differences distinct from causal conclusions and hypotheses. Do not add overlapping candidate savings or multiply by an invented number of future turns. Count the audit's own overhead separately.
Report findings as:
| Priority | Evidence | Proposed change and mechanism | Confidence | Savings | Tradeoff / validation |
|---|---|---|---|---|---|
| High/medium/low | path:line or log:line | Specific action | Measured/estimated/hypothesis | Supported units, or unmeasured | Same-task quality check |
07Research incorporated
The design uses Aider's bounded repository-map idea, Anthropic's local filtering and on-demand tool discovery, and the AGENTS.md evaluation's evidence for testing instruction changes. Tokenoscope informed the explicit distinction between text tokenization and provider usage. Published benchmark percentages are not reused as predicted savings.
See the source ledger for pinned code, official documentation, caveats, and how each idea influenced the command. The reference clone is in ignored .github_examples/Tokenoscope; research caches are in ignored .firecrawl/. Neither is a runtime dependency.
08Privacy and scope
The optional collector processes files in the host execution environment and never executes commands found in a transcript. The default native audit uses the model and tools provided by your app; local files read by the model are subject to that host’s data policies. Reports contain paths, model/tool names, sizes, hashes, counters and line references—not raw prompts, tool arguments, tool results, or instruction contents. Metadata can still be sensitive: review reports before sharing. Arbitrary credentials embedded in path/model names are not detected or redacted.
Git-ignored files are excluded from Git inventory; tracked files are still eligible. Common generated/vendor folders, symlinks, binary files, secret-like filenames, .github_examples, and output/research directories are also excluded. Function analysis currently supports Python only; other languages receive file measurements and agent review. Instructions are inventoried, not automatically proven active.
The included .gitignore covers local evidence. If using this skill elsewhere, add .token-audit/ to that project's ignore rules or choose an output directory outside the repository. Never commit raw session logs or credentials.
09Maintenance and removal
Uninstall or disable token-max through your app's plugin manager. Previously exported reports remain yours. If you used the old symlink installer, remove only that verified Token Max link before installing the marketplace version to avoid duplicate discovery.
Maintainer commands (not needed by plugin users)
git archive --format=zip HEAD:plugins/token-max > site/token-max.zip
python3 -m unittest discover -s tests -q
python3 scripts/check_site.py
claude plugin validate .
plugins/token-max/ is the self-contained distributable. The root skills symlink preserves legacy development paths. The old installer is developer compatibility tooling, not the user installation route. GitHub Actions builds the downloadable ZIP from tracked plugin files only.
10Troubleshooting
The plugin or skill does not appear
Confirm the Swack Tools marketplace is registered and token-max is installed/enabled. Refresh plugins or start a new chat. Use the app's plugin manager to update or reinstall. Organization policy may limit custom sources.
The slash command is unrecognized
Claude uses /token-max:token-audit. Codex uses a skill mention, $token-audit, or its skill picker. These are different host interfaces; a standalone /token-audit is not portable.
No project files or token totals appear
Select a project folder or attach relevant files inside the app. A plugin cannot reveal hidden history or usage counters. It provides a scoped qualitative review when measurements are unavailable.
Python is missing
The default audit does not need Python. Continue in native mode. Only an explicitly requested deterministic Codex collector requires an existing runtime; no installation is needed to obtain a report.
11Website & GitHub Pages
The documentation is static HTML, CSS, and a small progressive-enhancement script in site/. It has no framework build, analytics, cookies, or log-upload feature. The downloadable plugin ZIP is built from the same commit as the website.
Local preview
git archive --format=zip HEAD:plugins/token-max > site/token-max.zip
python3 scripts/check_site.py
python3 -m http.server 8000 --directory site
Open localhost:8000. The site works without JavaScript; the script adds copy buttons and highlights the current navigation section.
Deployment workflow
.github/workflows/pages.yml runs tests and site checks for pull requests and pushes to main. After successful checks, main-branch runs upload only the site/ directory and deploy to the github-pages environment. Only push events on main can upload or deploy; pull requests run checks only. The github-pages environment permits deployments only from the main branch.
Actions are pinned to commit SHAs. The check and upload jobs have read-only repository access. Only the deployment job receives pages: write and id-token: write. GitHub's built-in token and OIDC are used; no personal access token is embedded. Dependabot checks action updates monthly.
The repository requires verified signed commits on main. Maintainer commits use the swackhamer identity. Merging a pull request or pushing directly to main runs the same checked deployment; there is no manual deployment trigger.
One-time GitHub and DNS setup
- In the repository's Settings → Pages, choose GitHub Actions as the source.
- Set the custom domain to
token-max.swacktech.com. With an Actions publishing source, GitHub's settings or API must set this value; the included CNAME file alone is not sufficient. - Add the DNS record below. Keep Cloudflare in DNS-only mode with automatic TTL; GitHub Pages serves the site and manages its certificate.
- After DNS validation and certificate provisioning, enable Enforce HTTPS in Pages settings. HTTP requests then redirect to the HTTPS address.
| DNS field | Value |
|---|---|
| Type | CNAME |
| Name | token-max |
| Target | swack-tools.github.io |
Verify the Actions deployment, repository Pages settings, and live HTTPS address before declaring the site live. A successful local preview is not deployment proof.
References: GitHub's custom workflow guide and custom domain guide.