billion-context: Context compression for 100K windows, 5× fewer tokens, and long sessions
A bili plugin for compressing long AI coding sessions: 100K windows can handle billions of tokens while summaries stay reversible and cache-friendly.
GitHub ranxianglei/billion-context Updated 2026-10-11 Branch master Stars 782 Forks 74
TypeScript JavaScript Python SMT Context compression Token cost bili pi OpenCode Linux/macOS/Windows

🧭 Decision Guide

Try it if you

  • You run multi-day coding sessions with pi, OpenCode, or dsh and face a 100K context limit.
    The README's Why section and core description state that 100K context can support days, month-long sessions, and billions of tokens.
  • You want native client integration without URL edits, environment variables, or a fixed 8787 port.
    Option 1 says native plugins need no launcher, environment variables, fixed port, or URL edits; manual bili start owns 8787.
  • You use OpenCode and can set compaction.auto to false to prevent native and billion-context compression from stacking.
    The OpenCode section requires "compaction": { "auto": false }, otherwise it says double compression occurs.
  • You need cache-health visibility and can troubleshoot sessions using the 95–97% hit-rate target and /acp or /acp-cache.
    Cache health at a glance gives the 95–97% prefix-cache hit-rate figure and the /acp and /acp-cache diagnostics.

Skip it if you

  • You must keep both the native plugin and the standalone billion-context-pi or opencode-acp extension.
    The README explicitly says Native mode is mutually exclusive with billion-context-pi and opencode-acp, and the installer swaps entries.
  • You use codex but cannot start bili first or set HTTPS_PROXY yourself.
    The codex section requires bili start or a client lane proxy first; full compression also requires exporting HTTPS_PROXY yourself.
  • You plan to install the dsh or OpenCode plugin directly from a Git checkout instead of the npm-published package.
    The README says dsh and OpenCode Git checkouts have no published entry; dsh also fails to load the bundle without dist/.
  • You require a standard license such as MIT, but the project metadata only labels the license as Other.
    The project metadata lists the license as Other; the supplied material does not include the full license text.

Requirements

  • npm is required; the README installation command is npm install -g billion-context --prefix=~/.local.
  • On Linux/macOS, the README recommends the user-level ~/.local prefix; on Windows it uses %APPDATA%\npm, avoiding sudo.
  • Native plugins support pi, omp, OpenCode 1.x/2.x, dsh, kimi, hermes, and zcode.
  • Native dsh and OpenCode installation requires the npm-published form; Git checkouts have no published entry.
  • codex requires a reachable bili proxy; full compression also requires setting HTTPS_PROXY yourself.

First step (verbatim from README)

npm install -g billion-context --prefix=~/.local

Watch out

  • Native installation swaps standalone-extension entries, with the original configuration snapshotted in .bili-bak.
    The Native mode notes say the installer swaps entries and snapshots the original config.
  • pi native installation enables acp_delegate, acp_delegate_wait, and acp_delegate_cancel subagents by default.
    The pi notes list these three built-in subagents and reference a separate disabling section.
  • OpenCode can double-compress if auto-compaction is not disabled.
    The OpenCode native-installation instructions explicitly require compaction.auto to be false.
  • codex MCP tools depend on a reachable proxy; without one, tools/list fails with -32003.
    The codex section lists BILI_MCP_PROXY, the live-instance record, the 8787 default, and the -32003 error.
  • Manual bili start uses 8787; automatic lanes start in the 18787 zone and increment on collisions.
    The Ports section distinguishes the 8787 user zone from the 18787 self-managed zone.

Alternatives

  • Host built-in summarizer:It is a better fit when you only need the client's built-in summarization and do not need billion-context's incremental, reversible, prefix-cache-friendly mechanism.
    Why
  • billion-context-pi:It is a better fit when you already use pi's standalone in-process extension and do not plan to switch to bili native mode.
    Option 1 — Native plugin
  • opencode-acp:It is a better fit when you already use OpenCode's standalone in-process extension and need to retain that configuration path.
    Option 1 — Native plugin

Not stated in the README

  • The supplied README excerpt does not specify a supported Node.js version.
  • The supplied README excerpt does not list providers, model names, or compression-quality differences by model.
  • The supplied README excerpt does not provide real benchmarks for token cost, latency, or throughput before and after compression.
  • The supplied README excerpt does not specify the CPU, memory, or disk requirements for running the proxy.
  • The project metadata lists the license as Other, but the supplied material does not include the full license terms or restrictions.
  • The supplied README excerpt does not describe complete Windows installation, proxy, or client coverage.
  • The material does not establish a causal link between 782 stars, weekly Trending status, and any specific feature or release.

💡 Deep Analysis

6
No I use a 100K-context model for multi-day debugging and rely on subagents inheriting the parent session. Can I use the claimed 5x token savings and 95–97% cache hit rate for budgeting?
For: A heavy AI-agent user running long debugging and subagent-collaboration sessions on a 100K-context model who cares about prefix-cache hit rate and token cost

No, you should not use the claimed 5x savings or 95–97% hit rate as a budgeting guarantee; they are project targets or health indicators whose results depend on the model, cache, and session content.

  • The README describes “5× fewer tokens”; its cache section defines 95–97% as a healthy-session prefix-cache hit rate and says compression itself costs no more than 2%.
  • It also lists upstream cache TTL expiry, model switches, and bili bugs as causes of lower hit rates, so the figures are not an environment-independent SLA.
  • The project supports derived child sessions inheriting the parent’s compressed context, which suits subagent collaboration, but information completeness still depends on summarization.
  • Long sessions rely on incremental, layered, reversible compression; the project also explicitly cannot exceed the model or provider’s per-request context limit.

Use these figures to understand the design goals, not to calculate your bill or debugging quality directly. Complete logs, exact tool output, and non-summarizable state may not remain fully visible after compression.

  • README '# billion-context': '5× fewer tokens'
  • README 'Cache health at a glance': a healthy session keeps a 95–97% prefix-cache hit rate; compression itself costs ≤2%
  • README 'Cache health at a glance': causes include upstream cache TTL expiry, model switch, and a bili bug
  • README 'Derived (child) sessions inherit the parent's compressed context'
  • Project insights 'usage_limitations': it cannot exceed the model or provider’s per-request context limit
Not stated in the README:The README does not provide the distribution of actual token savings across models, providers, or task types.;It does not provide accuracy, added latency, or cache-hit data after subagents inherit compressed context.;It does not specify the exact fields or automation format of `/acp` and `/acp-cache` output.
Yes I use pi, omp, and OpenCode 1.x/2.x for large refactors and want month-long sessions to stay within a 100K context window. Is billion-context suitable for direct integration?
For: An AI coding engineer using pi, omp, and OpenCode 1.x/2.x for large refactors who wants month-long sessions to remain within a 100K context window

Yes, because this is the project’s primary long-session use case, but OpenCode requires special handling to avoid competing compression layers.

  • The README explicitly targets “small context windows (100K is enough)” and “month-long single sessions (billions of tokens)”.
  • Compression is incremental, reversible, and layered: summaries are written in small ranges and can be decompressed on demand, which fits large refactors better than one full summary at the context limit.
  • pi, omp, and OpenCode 1.x/2.x all have native plugin installation paths; the OpenCode installer also disables native auto-compaction.
  • The README warns that enabling OpenCode’s native auto-compaction together with billion-context causes double compression.

It still cannot exceed the provider’s per-request context limit, and retention of exact refactoring constraints depends on summary quality.

  • README '# billion-context': 'small context windows (100K is enough) · 5× fewer tokens · month-long single sessions'
  • README 'Why': 'compression here is incremental, reversible, and prefix-cache friendly'
  • README 'Option 1 — Native plugin': supported for pi, omp, and opencode (1.x and 2.x)
  • README 'Option 1 — Native plugin': OpenCode native auto-compaction is disabled; enabling both causes double compression
npm install -g billion-context --prefix=~/.local
Not stated in the README:The README does not provide correctness, latency, or token-saving comparisons for the same large refactor across pi, omp, and OpenCode 1.x/2.x.;It does not quantify how much complete build logs, exact code fragments, or implicit business constraints survive compression.
Yes I install billion-context globally with npm on Linux/macOS and run multiple pi, OpenCode, and dsh lanes. Does its installation and port management fit these constraints?
For: A personal AI coding developer on Linux or macOS using multiple client lanes who wants to avoid global-install permission errors and port conflicts

Yes, because the README explicitly addresses user-level npm installation and port isolation across multiple lanes, although you still need to verify PATH and the active instance state.

  • On Linux/macOS, the recommended --prefix=~/.local avoids sudo and EACCES caused by an old root-owned prefix; bili is installed under ~/.local/bin.
  • Manual bili start owns port 8787, while native hooks and launcher lanes use a self-managed zone beginning at 18787; collisions increment automatically.
  • During an upgrade restart, the proxy waits up to five seconds for the previous build to drain and reuses the same lane port when possible; it hops and logs only when the port is genuinely occupied.
  • This fits parallel client use, but stale processes, a manually started 8787 service, or abnormal exits can still cause a client to connect to the wrong instance.

It addresses common installation and port-lifecycle problems, but does not remove the operational cost of proxy processes, logs, and upgrades.

  • README 'Install': `npm install -g billion-context --prefix=~/.local`; avoid `sudo`
  • README 'Quickstart': `bili start` owns `8787`; lanes start at `18787`
  • README 'Quickstart': collisions hop +1; upgrade restart waits up to 5s and reuses the same port
  • Project insights 'common_pitfalls': stale processes and a manual 8787 service can lead to the wrong instance
npm install -g billion-context --prefix=~/.local
Not stated in the README:The README does not specify resource usage, maximum concurrency, or an upper port-range limit for many simultaneous lanes.;It does not describe automatic cleanup behavior after abnormal proxy exits outside the documented Windows notes.
It depends I use only OpenCode 2.x and must avoid double compression and configuration overwrites. Should I use billion-context’s native plugin or keep OpenCode’s built-in compression?
For: An AI coding engineer maintaining an OpenCode 2.x configuration who cannot accept native and external compression running at the same time

It depends: choose billion-context when you need long sessions, stable cache prefixes, and reversible recovery; keep OpenCode’s native compression when it is sufficient and sessions do not approach the context limit.

  • Native installation writes the plugin into OpenCode’s real configuration and disables compaction.auto, showing that the two compression mechanisms should not run together.
  • The README requires a manual configuration backup before installation; removal stores configuration snapshots in .bili-bak.
  • The project claims month-long single sessions and a 95–97% prefix-cache hit rate, but the result depends on model switches and upstream cache TTL.
  • With a git checkout rather than the npm package, the README says OpenCode has no published entry and may not load through its native Npm.add path.

The deciding factor is therefore not the OpenCode version alone, but whether you need a long-lived, recoverable context lifecycle and can accept configuration changes.

  • README 'Option 1 — Native plugin': set `"compaction": { "auto": false }`
  • README 'Option 1 — Native plugin': keep a manual backup of the file first
  • README 'Cache health at a glance': a healthy session keeps a 95–97% prefix-cache hit rate
  • README 'Option 1 — Native plugin': a git checkout has no published entry
bili plugin install opencode
Not stated in the README:The README does not compare quality or cost between OpenCode native compression and billion-context on the same model and session.;It does not state whether every OpenCode 2.x minor version has received equivalent stability validation.
No I need to integrate Claude and Codex in an enterprise environment handling sensitive code, but network policy restricts local proxies and deployment requires review of licenses, logs, and data flows. Is billion-context suitable?
For: A technical lead using AI coding clients such as Claude and Codex in an enterprise or sensitive-code environment with network-isolation and license-review constraints

No, not for direct deployment, unless enterprise security and compliance explicitly permit a local proxy that rewrites model requests and accept the licensing and session-storage implications.

  • The project operates as a local proxy between the client and model provider, observing, rewriting, or routing requests; that may conflict with network isolation, endpoint controls, or sensitive-code policies.
  • Codex integration usually involves starting a proxy, setting HTTPS_PROXY, or using bili codex, increasing enterprise networking and troubleshooting surface area.
  • The project data marks the license as Other, and the README mentions attribution requirements beyond MIT; closed-source integration, redistribution, and commercial deployment cannot be reviewed as MIT-only.
  • The proxy creates processes, ports, logs, configuration, and session files; the project insights call for review of request routing, log storage, and data compliance.

Claude and Codex may be technically integrable, but compliance prerequisites outweigh compression benefits here. The README provides no guarantee of enterprise auditing, data isolation, or zero-logging mode.

  • Project insights 'usage_limitations': the project observes, rewrites, or routes model requests at a local proxy layer
  • Project insights 'usage_limitations': enterprise security, network isolation, or sensitive-code environments may disallow this architecture
  • Project data: license is `Other`
  • README 'Attribution requirement': additional attribution requirements exist beyond MIT
  • Project insights 'usage_limitations': the proxy introduces processes, ports, logs, and configuration files
Not stated in the README:The README does not provide a zero-logging mode, enterprise audit interface, session-file encryption, or data-residency controls.;The exact scope of the full license text and additional attribution obligations is not included in the provided excerpt.;It does not state whether enterprise authentication, TLS-inspecting proxies, or offline operation in isolated networks are supported.
It depends I need to integrate Codex, Hermes, and Claude. Codex cannot fully inject proxy routing through plugin installation alone, Hermes uses a Python plugin, and Claude has a marketplace flow. Can billion-context be deployed consistently across them?
For: An AI-agent infrastructure engineer using clients with different integration capabilities, including Codex, Hermes, and Claude, who needs unified proxy routing and context compression

It depends: all three clients are covered, but the unified element is the context-management goal, not an identical installation or routing procedure.

  • Claude can be installed through the marketplace and then configured with /billion-context:bili-setup; the Hermes installer copies a Python plugin into ~/.hermes/plugins/billion-context/.
  • The launcher list includes bili codex, bili hermes, and bili claude, so each has a project-defined integration path.
  • Codex cannot obtain complete proxy routing from plugin installation alone; the project insights state that it usually requires starting bili, setting HTTPS_PROXY, or using bili codex.
  • Client configuration, authentication, session formats, and subagent behavior can differ, and the README does not promise feature equivalence.

You can standardize the operational entry point and compression layer, but not assume one plugin configuration covers all three clients; Codex proxy reachability is the critical deployment dependency.

  • README 'Option 1 — Native plugin': Claude marketplace and `/billion-context:bili-setup`
  • README 'Option 1 — Native plugin': Hermes uses a Python plugin copied to `~/.hermes/plugins/billion-context/`
  • README 'Option 2 — Launcher': includes `bili codex`, `bili hermes`, and `bili claude`
  • Project insights 'common_pitfalls': Codex usually needs `HTTPS_PROXY`, a running bili proxy, or `bili codex`
npm install -g billion-context --prefix=~/.local
Not stated in the README:The README does not provide a compatibility matrix for Codex, Hermes, and Claude with the same proxy and provider.;It does not state the required Python version or supported Hermes client-version range.

✨ Highlights

  • 100K context can support month-long sessions with billions of tokens
  • Incremental, reversible compression targets 5× fewer tokens while preserving the cache prefix
  • Native plugins support pi, OpenCode, dsh, kimi, hermes, and more
  • Healthy sessions report a 95–97% prefix-cache hit rate

🔧 Engineering

  • bili offers three integration modes: native plugin, Launcher, and /bili/ URL prefix
  • bili plugin install can configure clients including pi, omp, and OpenCode
  • Layered summaries can be decompressed on demand, unlike an irreversible host summarizer

⚠️ Risks

  • Native mode is mutually exclusive with billion-context-pi and opencode-acp
  • codex still requires bili to be running; unreachable proxies make tools/list return -32003
  • OpenCode requires auto-compaction to be disabled to avoid double compression
  • A dsh Git checkout has no published entry; missing dist/ prevents the bundle from loading

👥 For who?

  • Developers running long coding sessions with pi, OpenCode, or dsh
  • Teams constrained by 100K windows and token costs that need billion-token sessions
  • Users who want bili's native plugin to avoid fixed ports and URL edits