Release preview: the installation examples below target the forthcoming Gobstopper release. Check published releases before installing.

gobstopper

Automatic context compaction for coding-agent sessions — Codex and Claude Code today, any JSONL-transcript agent tomorrow.

gobstopper is the first cross-provider context compactor that preserves a content-addressed archive of every conversation state, compacts resumable Claude Code and Codex transcripts, and proves on real API calls that it reduces context without inventing answers.

Run it, and it watches your agent sessions. When a session's context crosses a configured threshold, gobstopper compacts it — using a strategy you choose per session, per provider, or per preset — and stores the exact pre- and post-state in the vault so you can audit what changed.

Why

Providers compact late. Codex fires auto-compaction at ~90% of the model's context window; Claude Code similar. On a 1M-token window that means every turn near the end of a long session costs ~900k input tokens — and on subscription plans, those tokens come out of your weekly allowance.

Context is a sawtooth problem. If you compact at trigger T and the summary floor is F, steady-state context occupancy per turn is roughly (T+F)/2:

policytriggerflooravg context/turnrelative occupancy
provider default (1M window)~900k~60k~480k1.0x
gobstopper default250k40k~145k~3.3x lower occupancy
gobstopper aggressive150k20k~85k~5.6x lower occupancy

This is an occupancy model, not a measured subscription savings claim. Actual token cost depends on cache hit rates, whether summary turns are billed, re-fetches caused by lost detail, and how often compaction itself runs. Anthropic documents that clearing tool results can invalidate the prompt cache. gobstopper reports measured file-byte changes and observed provider usage where available; it does not project dollar or quota savings into its public claims.

Infinite memory

Every compaction writes the pre- and post-state into a content-addressed vault (~/.local/share/gobstopper/vault/). The same record appears once across every version it participates in, so keeping every state does not explode storage.

gobstopper recall --query <q> turns that vault into agent-addressable memory: it searches every archived state-card digest, ranks results by query relevance, and returns the high-level state of the matching turns. The agent does not need to remember session IDs — it can ask for the last time it worked on a file, a goal, or a decision and get a ranked summary with a snapshot SHA it can show or diff.

gobstopper mcp exposes the same surface as a read-only Model Context Protocol server on stdio — tools list_sessions, recall, history, show, diff, plan, and verify. Register it once and the agent can query its own history mid-session:

claude mcp add gobstopper -- gobstopper mcp          # Claude Code
# ~/.codex/config.toml:  [mcp_servers.gobstopper] command = "gobstopper", args = ["mcp"]

Compaction itself isn't free — each cycle costs one large input call and risks losing detail — so strategy matters. That is the actual product here: not "compact earlier" but "compact with the right strategy at the right boundary."

Strategies

idkindwhat it does
auto (default)dynamicselects per-session from transcript composition: tool-heavy → cache_aware, chatty → structured, live/empty → sawtooth
sawtoothproviderfires the provider's own compaction early (thread/compact/start on Codex app-server; /compact or --autocompact on Claude)
elidetranscriptstubs stale tool outputs oldest-first until the floor; deterministic, no model call
cache_awaretranscript
compactedtranscript
scoredtranscript
structuredtranscriptplaceholder state-card digest (goal/decisions/files/todos); currently emits item labels, not a real summary. Safe for chat-only sessions but should not be mistaken for a semantic compressor
agentictranscriptreserved for a bounded editor-model backend; today preset.command is the only extension point and is treated as untrusted code

Custom strategies are userspace code: a preset can name a command that receives the normalized transcript as JSON on stdin and returns an edit plan on stdout, or install a versioned gobstopper-plugin.json bundle (see gobstopper plugin check). Host-side validation bounds every proposal: no edit can grow the transcript, leave protected recent output, bypass linkage checks, or exceed configured digest size.

Install & use

cargo install --git https://github.com/hraness/gobstopper gobstopper
# or from a checkout: cargo build --release

gobstopper detect                  # sessions, context sizes, lifetime burn
gobstopper plan <session>          # what would happen, under which strategy
gobstopper plan <session> --trigger 100000 --floor 310000   # tune the trade-off
gobstopper eval <session>          # every strategy side-by-side on temp copies
gobstopper apply <session>         # vault snapshot + produce validated fork (idle sessions)
gobstopper verify <session>        # resume-validity check (exit 1 on errors)
gobstopper fork <session>          # clone under a fresh session id + resume cmd
gobstopper undo <session>          # restore a pre-compaction snapshot into a new fork
gobstopper undo <session> --in-place   # restore exact bytes to the same session id
gobstopper vault                   # list snapshots in the undo vault
gobstopper install-hooks           # Claude + Codex compaction lifecycle hooks
gobstopper watch --dry-run         # the daemon path: poll, threshold, prepare copy
gobstopper explain                 # the occupancy math above
gobstopper recall --query <q>      # search state-card digests across all archived sessions
gobstopper history <session>       # every archived state of one session
gobstopper diff <sha-a> <sha-b>    # structural comparison of two vault snapshots
gobstopper bench                   # benchmark every strategy across discovered sessions
gobstopper tune <session>          # preview the adaptive trigger/floor for a session
gobstopper mcp                     # read-only MCP server: the vault as agent tools

Every apply/watch compaction snapshots the source transcript into a content-addressed vault (~/.local/share/gobstopper/vault/) and publishes the result as a separate, verified file. The original transcript is never overwritten by a standalone compaction run; live session surgery must be dispatched by the session owner. Each compaction appends a numeric record to events.jsonl in the gobstopper/compaction-events-v1 schema.

Config: ~/.config/gobstopper/config.toml

[policy]
strategy = "auto"            # sawtooth | elide | compacted | cache_aware | scored | structured | agentic
trigger_tokens = 250_000
floor_tokens = 40_000
adaptive = true              # derive trigger/floor per session — see `gobstopper tune`

[provider.codex]             # per-provider overrides
trigger_tokens = 200_000

[sessions."01a08d7c-…"]      # per-session overrides
strategy = "structured"
trigger_tokens = 120_000

[presets.deep-work]          # named presets, selectable via --preset
strategy = "elide"
trigger_tokens = 150_000

[presets.custom-script]      # userspace code preset
command = "python3 ~/bin/my_compactor.py"

Managed sessions (oompa profiles, sandboxed homes) use different roots: point gobstopper at them with --codex-home / --claude-home.

With adaptive = true, the effective trigger/floor are re-derived per session at each decision point: the trigger is capped at a quarter of the provider-advertised context window, backed off (bounded 2x) when recent compactions reclaimed too little to be worth a cycle, and tightened when most of the window is reclaimable tool output. The adjustment is deterministic and its reasons appear in plan output and telemetry. gobstopper tune <session> previews it.

The oompa seam

oompa never parses provider transcript files — that boundary stays intact. Integration is the policy-check subcommand: oompa already records token_usage events (totalTokens, modelContextWindow) in its neutral timeline, so it asks gobstopper what to do with them:

gobstopper policy-check --provider codex --context-tokens 300000 \
    --session-active --json
# {"action":"provider_compact","control":"thread/compact/start", ...}

oompa then invokes thread/compact/start on its own app-server connection, or launches Claude sessions with --autocompact <tokens>. Deeper transcript surgery (elide/structured) stays gobstopper-side, applied to idle sessions or at resume boundaries.

How providers compact today

levercodexclaude code
auto-compact thresholdmodel_auto_compact_token_limit (config; ≤90% of window)--autocompact <100k–1M> argv
on-demand triggerthread/compact/start (app-server v2)/compact [instructions]
compaction promptcompact_prompt config/compact instructions
tool output captool_output_token_limit
live usage streamthread/tokenUsage/updated notificationmessage.usage per turn
transcript store~/.codex/sessions/**/rollout-*.jsonl~/.claude/projects/*/*.jsonl

Codex persists compaction as a compacted rollout record carrying replacement_history — the provider's own resume-time context swap. gobstopper's transcript path respects that shape (elision inside replacement_history is supported); writing custom compacted records for fully custom summaries is the designed v0.2 path.

Layout

Live qualification

gobstopper has been live-qualified on real provider sessions. The most recent trial used a 333k-token Claude session, asked the same resume question under four conditions, and measured the tokens the provider actually consumed on the next turn:

conditioninput tokens on resumeoutput tokensrecalled the standing task?
none (original)312,7221,405yes — reported npm unification and stalled renames
gobstopper elide219,1671,052yes — same standing task, stalled renames
gobstopper compacted220,447621yes — same standing task from the state-card digest
claude --autocompact 10056,300416no — incorrectly claimed the renames were already done and published

gobstopper elide and compacted both cut the resume context by about 30% while keeping the answer accurate. Claude's native --autocompact 100 cut the resume context by ~82% but produced a confident, inaccurate summary of the session.

The same question was then asked on a 101k-token Codex session:

conditioninput tokens on resumeoutput tokensrecalled the standing task?
none (original)101,275244yes — Oh's memory benchmark and the 0.60 expansion gate
gobstopper elide57,98083yes — same 0.545 score and 0.60 gate
gobstopper compacted34,503159yes — same BEAM experiment and expansion gate

On Codex, compacted cut resume input tokens by 66% and elide cut them by 43%, both with accurate answers. There is no one-shot Codex native compact to compare against.

The same-session cache_aware A/B (339k-token Claude session, floor 310k, real provider cache counters):

conditioncache_readcache_creationfile-level prefix preservedcostaccurate?
none (original)10,010325,647$6.53yes
gobstopper cache_aware13,536258,517107,884 tokens$5.19yes
gobstopper compacted13,536257,5056,639 tokens$5.18yes

Honest result: cache_aware and compacted cost the same on the API — Claude Code's prompt-cache breakpoints sit at ~13.5k regardless of how much file-level prefix stays byte-identical. What cache_aware actually buys is 16x more preserved prefix at the transcript level, which is what keeps gobstopper diff audits small and Merkle-dedup efficient across repeated compactions. Choose it when auditability matters; choose compacted when you want the Codex-native record.

That is the difference gobstopper is built for: measured, auditable compaction that does not replace the transcript's actual state with a plausible invention. Every pre- and post-state is in the vault, so you can gobstopper diff the exact structural changes and decide which strategy to trust.

Status

Production-ready for idle and resume-boundary transcript compaction on Codex and Claude Code. Core detection, planning, and transcript surgery are covered by a property-tested suite plus the live qualifications above. apply and watch always snapshot the source into the content-addressed vault first; verify checks resume-validity, undo restores to a new fork (or --in-place back onto the same session id), and vault keeps every state. sawtooth can route to Codex's thread/compact/start over a private app-server connection when codex-cli is installed and trusted.

The structured and agentic strategies remain conservative placeholders or extension points, not proven semantic compressors. Custom strategy code is treated as untrusted and validated by the host before any transcript is written. See docs/design.md for the research record and docs/roadmap.md for the full plan.

License

MIT OR Apache-2.0