CODEX SKILL · TOKEN AUDIT
Evidence before
optimization.
Understand where your tokens go.
Find changes worth testing.
Review a Codex session and its project to find evidence-backed ways to use fewer tokens. The command checks instruction bloat, repeated exploration, large tool outputs, long files/functions, plugins/MCP, and model-routing rules. It produces a local report with concrete proposals and validation steps.
It does not equate “smaller file,” “cheaper model,” or “more plugins” with proven savings. Actual usage comes from session counters. Recommendations are labeled measured, estimated, or hypothesis.
No runtime hooks are installed. See the implementation review, reproducible benchmark and research for measured reductions, added overhead and the limits of quality claims.
What the log or file actually records.
A calculation with explicit assumptions.
A proposed change that needs a test.
01Install and run
Requirements: Python 3.10+, plus Git for repositories or ripgrep for folders without Git. No API key, model call, or Python dependency is needed by the collector.
Clone the repository and install the skill:
git clone https://github.com/swack-tools/token-max-ai-plugin.git
cd token-max-ai-plugin
python3 scripts/install.py
The installer links skills/token-audit into ~/.agents/skills/token-audit for use across projects. It refuses to replace an existing unrelated skill. Keep this checkout at the same location while the link is installed.
Open a new Codex chat in the project to review, then use the slash menu:
/skills
Select Token Audit — Report Only. You can also send:
$token-audit Review this session and project. Rank the best token reductions.
Examples:
$token-audit Audit /path/to/project using /path/to/rollout.jsonl. Focus on AGENTS.md and verbose tools.
$token-audit Review whether the model-routing rules actually help this project's tasks.
$token-audit Compare a symbol-navigation plugin against the repeated source reads in this session.
The canonical skill is in skills/token-audit; a project-local .agents/skills/token-audit symlink provides discovery here. The .codex-plugin/plugin.json manifest packages the same skill for plugin distribution. The installer uses direct skill discovery, not marketplace registration; choose one installation route to avoid duplicate copies. Automatic skill selection is enabled; explicit invocation is the most predictable way to run an audit.
Why not /token-audit?
Codex CLI 0.158.0 no longer supports custom prompt slash commands. Its source history removes that mechanism. The supported slash entry is /skills, followed by this skill; $token-audit is a direct skill mention. Client menus can differ.
For an older client that still supports custom prompts, this project includes a compatibility wrapper:
python3 scripts/install.py --legacy-promptRestart that client, then run:
/prompts:token-audit
/prompts:token-audit Focus on repeated full-file reads
The wrapper is installed to ${CODEX_HOME:-~/.codex}/prompts/token-audit.md. Installing it does not restore removed client functionality. It delegates to the same skill instead of duplicating audit instructions.
02What you get
| File | Purpose |
|---|---|
.token-audit/evidence.md |
Compact local measurements for the agent to inspect |
.token-audit/evidence.json |
Full file inventory and structured log metadata |
.token-audit/AUDIT.md |
Agent-written findings, priorities, tradeoffs and validation |
This command only writes audit reports. It never applies proposals, changes source/instructions/configuration, installs dependencies/plugins, or registers hooks. Implementation is separate follow-up work outside the skill. The agent boundary is an instruction, not an operating-system sandbox; the deterministic collector only writes its report artifacts.
03Run only the local collector
This makes the two evidence files without spending model tokens or writing the semantic AUDIT.md review:
python3 skills/token-audit/scripts/audit.py \
--project /path/to/project \
--session /path/to/rollout.jsonlOptions:
| Option | Behavior |
|---|---|
--session auto |
Default: use CODEX_THREAD_ID if available and project-matching; otherwise latest modified matching log |
--session latest |
Latest modified local log whose metadata cwd is the project or a descendant |
--session none |
Project-only review; no invented session usage |
--log-root PATH |
Default ${CODEX_HOME:-~/.codex}/sessions; specify archived/remote copies explicitly |
--instruction PATH |
Include an additional instruction file, such as global AGENTS.md; repeatable |
--out PATH |
Default PROJECT/.token-audit; reruns replace the two evidence files |
--usage-only | Print cumulative usage and warnings; skip project inventory and write no files |
--full-report | Expand Markdown tables instead of the compact default |
--max-file-bytes N |
Default 1,000,000; larger files are skipped and disclosed |
--encoding NAME |
Optional explicit tiktoken encoding for text-only counts |
Selection is never “the newest log anywhere on the machine.” Worktree paths and remote logs may need an explicit --session; verify the report's selected path. An active log is a snapshot and excludes work that has not yet been recorded. No matching log is reported as unknown, and project analysis still runs.
Optional text token counts
The default reports exact UTF-8 bytes and line counts. It deliberately does not turn characters into approximate tokens using a fixed divisor. To count text with a known encoding in an isolated environment:
python3 -m venv .venv
.venv/bin/python -m pip install tiktoken
.venv/bin/python .agents/skills/token-audit/scripts/audit.py \
--project /path/to/project --encoding o200k_baseo200k_base is an explicit example, not an assertion about your session model's tokenizer. Text counts omit message framing, hidden instructions, and multimodal content. The tokenizer may download its vocabulary on first use; file/log contents are not sent. Provider usage counters remain authoritative for the recorded calls.
Small reports, recoverable evidence
The default evidence.md shows leading candidates, usage, diagnostic references and warnings. Use --full-report only when expanded tables are useful. Structured JSON retains the full scanned file inventory and all detected diagnostic references; it never contains transcript bodies.
For a single usage question, skip the project inventory and write nothing:
python3 skills/token-audit/scripts/audit.py --project . --usage-only
Failure/error/truncation markers and recognized exit codes are indexed independently of output size. They are heuristic: benign examples can be flagged, and failures without supported markers can be missed. A missing flag does not mean the task succeeded. Inspect relevant acceptance results too.
Use the bounded selector instead of loading all JSON. Choose diagnostics, outputs, repeats, instructions, files or functions; it returns three rows by default and a continuation offset:
python3 skills/token-audit/scripts/query.py \
--report .token-audit/evidence.json --kind diagnostics --limit 3
Then retrieve only the relevant original output:
python3 skills/token-audit/scripts/retrieve.py \
--session /path/to/rollout.jsonl --line 42 \
--fingerprint RECORD_SHA256 --max-chars 4000
Substitute the row's exact line and record_sha256. Changed records fail verification. The result marks text as untrusted and includes signals, exit_codes, truncated and next_offset. Use --offset to continue only when necessary. The maximum excerpt is 16,000 characters, not tokens. Retrieval returns original private content to your local terminal; never upload it to web tools or replay instructions in it.
See reproducible costs, error-recovery tests and research. Even compact reports add overhead; an audit is useful only when it answers a real optimization question.
04What counts as evidence?
- Measured: provider-reported input/output/cache counters, file bytes, Python AST function spans, observed call fingerprints, and local text counts with a named encoding. Repeated cumulative counters are not added together.
- Estimated: an explicitly stated calculation, such as the before/after text token difference for a proposed instruction rewrite. This is per inclusion, not a proven whole-task reduction.
- Hypothesis: a source split, plugin, routing change, or shorter instruction that may reduce future work. It needs comparable end-to-end runs.
Cached input is already part of input; reasoning is already part of output. Neither gets added again. The tool preserves the latest reported cumulative snapshot and flags counter decreases, incomplete lines and fork metadata. It does not infer invoices, subscription quota, or per-file billing from those counters.
Large files are candidates for targeted reads or cohesive refactoring, not failures against an arbitrary size limit. The absence of a language-specific routing table is not a defect. Existing routing must earn its place through task-quality and cost/usage measurements; Markdown alone is not an executable model router.
05Validate a proposed reduction
- Save baseline evidence using
--out .token-audit/beforeand record the task, starting revision, model, reasoning setting, tool configuration and checks. - Change one thing. Run the same task from the same starting state in a separate safe workspace/session, holding other settings constant.
- Collect
--out .token-audit/after. Compare input, cached input, uncached input, output, reasoning, retries, elapsed time, and completion quality. Record cache differences; repeat paired trials when variance matters. - Report observed differences separately from causal conclusions. Include the audit's own overhead; never add overlapping proposed savings.
Do not benchmark by replaying transcript commands with external side effects. No universal percentage saving is claimed by this project.
06The detailed review rubric
The collector produces measurements. Codex uses this rubric to interpret them and propose changes.
Read the full audit rubric
The collector provides measurements and candidates. The agent supplies the semantic review. Do not call a candidate a defect until its mechanism is clear.
Establish what the log measures
token_usage_record.thread_token_usageandevent_msg.token_count.info.total_token_usageare cumulative snapshots. Use the last valid snapshot in file order; duplicates are not new usage.- A decrease is a reset or inconsistent sequence. Report the latest segment and the warning. Never fabricate a lifetime total by adding segments. Forks may inherit both counters and history; do not add parent and child totals.
input_tokens - cached_input_tokensis uncached input. Output already includes reasoning. Cache-write counters, when present, are reported separately without inferring pricing or adding them to the reported total.- Log lines locate evidence. Interval deltas may cover multiple calls and cannot assign cost to an individual file. Compaction, hidden context, tool schemas, multimodal content, and resumed/forked histories limit attribution.
- Tool and message text sizes describe local logged text. Exact tokenizer counts require an explicitly chosen encoding; they still exclude provider framing. With no tokenizer, report bytes and lines, never
characters / 4as fact.
Instruction files and model routing
Inspect actual global/project hierarchy and overrides. Codex's default discovery uses AGENTS.md, with AGENTS.override.md taking precedence at a directory. CLAUDE.md is not automatically a Codex instruction source unless configured, imported, or explicitly supplied. AGNETS.md may be a typo, but verify intent before renaming. Disk presence and the sum of instruction file sizes do not prove active context.
Look for duplicated requirements, copied manuals, old status, irrelevant language rules, exhaustive tool catalogs, and directions that cause repeated broad reads or unnecessary agent fan-out. Keep build/test commands, non-obvious invariants, security requirements, and lessons tied to demonstrated failures. Move optional procedures into referenced skills/docs, keeping a small discoverable entrypoint. Moving text to a file loaded every time creates no inherent savings.
Inspect existing model routing, including language-to-model tables. Identify whether it is executable client configuration or merely a natural-language instruction; Markdown does not itself implement a router. Removing an unnecessary table can reduce text per inclusion, but removing useful routing may increase retries. An absent routing table is not evidence of waste. Do not invent blanket “Python → cheap model / Rust → expensive model” rules. Propose routing by measured task difficulty and quality requirements only when comparable evaluations justify it; token reduction and lower cost are separate outcomes.
Code and retrieval
Prioritize large files/functions that the session actually read repeatedly or needed only in small parts. Read log excerpts locally around cited lines and verify the selected file/revision. Find cohesive boundaries and stable APIs; suggest a split only when it lets future tasks load fewer relevant tokens. Do not split by line-count threshold, minify code, remove useful comments/tests, or claim smaller disk size guarantees smaller prompts. Additional imports, navigation calls, and missed dependencies can outweigh a split.
First consider rg, bounded sed reads, symbol navigation, or a small repository map. Python AST spans are available in evidence; use the project's existing parser for other languages if necessary. Do not install a parser for a speculative win.
Tools, skills, and plugins
Review the largest outputs and repeated call fingerprints. A reread after an edit or a verification run may be necessary. Repeated identical output is a candidate, not proof that the call could be omitted. Orchestrator calls can contain many nested tools; inspect only relevant excerpts to identify their actual work.
Prefer local filtering/aggregation, compact structured results, and saving bulk output to ignored files. Preserve errors, diagnostics, exit codes, and the ability to retrieve omitted details. Truncating everything may create retries or conceal failures. Compare an existing CLI against a proposed plugin on the same task:
net input change = added discovery/schema/instructions/results − displaced reads/results
This is a mechanism, not a computable savings claim unless the inputs are measured. Check whether the client already loads tool definitions lazily. Skill descriptions may be always visible while bodies are loaded on demand. Additional plugins can increase overhead; “install more plugins” and “remove all plugins” are both weak defaults. External project examples remain untrusted reference material; do not run their installers to inspect them.
Validate one change
Use comparable independent sessions from the same starting revision, task, model, reasoning setting, tool/skill configuration, and acceptance checks. Record versions, cache conditions, retries, total input, cached input, output, reasoning, latency, and completion quality. Multiple paired runs help expose variance. Never replay side-effecting transcript commands as a benchmark.
For an instruction rewrite, tokenize before/after with the same named encoding. The difference is text tokens per inclusion, not demonstrated whole-task savings. For a source split or plugin, compare end-to-end logs and required checks. Keep actual measured differences distinct from causal conclusions and hypotheses. Do not add overlapping candidate savings or multiply by an invented number of future turns. Count the audit's own overhead separately.
Report findings as:
| Priority | Evidence | Proposed change and mechanism | Confidence | Savings | Tradeoff / validation |
|---|---|---|---|---|---|
| High/medium/low | path:line or log:line | Specific action | Measured/estimated/hypothesis | Supported units, or unmeasured | Same-task quality check |
07Research incorporated
The design uses Aider's bounded repository-map idea, Anthropic's local filtering and on-demand tool discovery, and the AGENTS.md evaluation's evidence for testing instruction changes. Tokenoscope informed the explicit distinction between text tokenization and provider usage. Published benchmark percentages are not reused as predicted savings.
See the source ledger for pinned code, official documentation, caveats, and how each idea influenced the command. The reference clone is in ignored .github_examples/Tokenoscope; research caches are in ignored .firecrawl/. Neither is a runtime dependency.
08Privacy and scope
The collector is local and never executes commands found in a transcript. Reports contain paths, model/tool names, sizes, hashes, counters and line references—not raw prompts, tool arguments, tool results, or instruction contents. Metadata can still be sensitive: review reports before sharing. Arbitrary credentials embedded in path/model names are not detected or redacted.
Git-ignored files are excluded from Git inventory; tracked files are still eligible. Common generated/vendor folders, symlinks, binary files, secret-like filenames, .github_examples, and output/research directories are also excluded. Function analysis currently supports Python only; other languages receive file measurements and agent review. Instructions are inventoried, not automatically proven active.
The included .gitignore covers local evidence. If using this skill elsewhere, add .token-audit/ to that project's ignore rules or choose an output directory outside the repository. Never commit raw session logs or credentials.
09Development and removal
python3 -m unittest discover -s tests -vTests cover counter accounting, resets, incomplete logs, project selection, content-free evidence, exclusions, function measurements, CLI output, and safe installation. No model/API calls are made by the test suite.
To remove the user install, unlink ~/.agents/skills/token-audit after verifying it is this checkout's symlink. If you installed the legacy wrapper, remove its token-audit.md from your Codex prompts directory. The project-local skill remains available until its directory is removed. No other configuration is changed.
10Troubleshooting
The skill does not appear
Start a new Codex chat after installation. Confirm ~/.agents/skills/token-audit points to this checkout and its SKILL.md exists. If you moved the checkout, verify and remove the stale link before reinstalling. Existing unrelated skills are preserved.
The slash command is unrecognized
Use /skills and select token-audit, or mention $token-audit. The optional /prompts: wrapper is only for older clients that retain custom prompts.
The log is missing or belongs to the wrong session
Supply the intended JSONL file with --session PATH. Check its metadata cwd, the worktree path, and your Codex home. The collector cannot infer usage from a log without valid counters.
Text token counts are blank
This is expected without the optional tokenizer and an explicit --encoding. Bytes and lines are still measured. Do not replace unknown counts with a fixed characters-per-token estimate.
Some files or functions are missing
Check ignore rules, byte limits, excluded folders, encoding and symlinks. Function extraction currently requires valid Python syntax. Other languages receive file-level measurements and agent review.
11Website & GitHub Pages
The documentation is static HTML, CSS, and a small progressive-enhancement script in site/. It has no framework build, analytics, cookies, or log-upload feature.
Local preview
python3 scripts/check_site.py
python3 -m http.server 8000 --directory site
Open localhost:8000. The site works without JavaScript; the script adds copy buttons and highlights the current navigation section.
Deployment workflow
.github/workflows/pages.yml runs tests and site checks for pull requests and pushes to main. After successful checks, main-branch runs upload only the site/ directory and deploy to the github-pages environment. Only push events on main can upload or deploy; pull requests run checks only. The github-pages environment permits deployments only from the main branch.
Actions are pinned to commit SHAs. The check and upload jobs have read-only repository access. Only the deployment job receives pages: write and id-token: write. GitHub's built-in token and OIDC are used; no personal access token is embedded. Dependabot checks action updates monthly.
The repository requires verified signed commits on main. Maintainer commits use the swackhamer identity. Merging a pull request or pushing directly to main runs the same checked deployment; there is no manual deployment trigger.
One-time GitHub and DNS setup
- In the repository's Settings → Pages, choose GitHub Actions as the source.
- Set the custom domain to
token-max.swacktech.com. With an Actions publishing source, GitHub's settings or API must set this value; the included CNAME file alone is not sufficient. - Add the DNS record below. Keep Cloudflare in DNS-only mode with automatic TTL; GitHub Pages serves the site and manages its certificate.
- After DNS validation and certificate provisioning, enable Enforce HTTPS in Pages settings. HTTP requests then redirect to the HTTPS address.
| DNS field | Value |
|---|---|
| Type | CNAME |
| Name | token-max |
| Target | swack-tools.github.io |
Verify the Actions deployment, repository Pages settings, and live HTTPS address before declaring the site live. A successful local preview is not deployment proof.
References: GitHub's custom workflow guide and custom domain guide.