Ponytail: An AI-agent rules plugin that makes Claude Code write less code
A minimization plugin for Claude Code and Codex that reuses native features without dropping safety checks.
GitHub DietrichGebert/ponytail Updated 2026-09-04 Branch main Stars 159.6K Forks 8.6K
JavaScript Python Claude Code Codex AI agents YAGNI FastAPI React

🧭 Decision Guide

Try it if you

  • You use Claude Code on a FastAPI+React repository and often see overbuilt features such as date pickers.
    The README “Numbers” section says the benchmark edits full-stack-fastapi-template and reduces the date picker from 404 to 23 lines.
  • You use Codex and want to inspect the current diff with /ponytail-review.
    The README “Commands” section lists /ponytail-review and says Codex invokes it as @ponytail-review.
  • You need to retain security, error handling, and accessibility while reducing agent output.
    The README “Numbers” section reports 100% safety for ponytail and explicitly excludes these concerns from reduction.

Skip it if you

  • Your agent is GPT-5.5 and task cost or latency is highly sensitive.
    The README says the reasoning model can move in the opposite direction on GPT-5.5, increasing cost and latency.
  • You only need single-shot prompt generation rather than real agent sessions such as Claude Code.
    The README says the older 80–94% figure came from single-shot generation and was partly affected by prose-padding in the baseline.
  • Your environment lacks Node and requires Claude Code hooks to execute normally every time.
    The README “Install” section requires Node on PATH; otherwise always-on activation stays quiet instead of running normally.

Requirements

  • {'text': 'Claude Code and Codex plugins need node on PATH; Nix/nvm users must expose it to the non-interactive shell.', 'text_en': 'Claude Code and Codex plugins need node on PATH; Nix/nvm users must expose it to the non-interactive shell.'}
  • {'text': 'After Codex installation, run codex, open /hooks, and review and trust the two lifecycle hooks.', 'text_en': 'After Codex installation, run codex, open /hooks, and review and trust the two lifecycle hooks.'}
  • {'text': 'The correctness benchmark invokes python3 or python; CSV checks require pandas installed locally.', 'text_en': 'The correctness benchmark invokes python3 or python; CSV checks require pandas installed locally.'}

First step (verbatim from README)

/plugin marketplace add DietrichGebert/ponytail

Watch out

  • Claude Code installation requires sending two separate /plugin commands.
    The README “Claude Code” section explicitly says, “You have to send two separate prompts.”
  • Run node scripts/uninstall.js before removing the plugin, or the script will be deleted with it.
    The README “Uninstall” section requires running the script before the host remove command.
  • Codex requires a new thread, and its desktop app must be restarted after installation.
    The README “Codex” section says to start a new thread; the desktop app requires a restart.

Alternatives

  • caveman:Use it when you want to shrink the agent's prose rather than the code it actually builds.
    README “FAQ” and “Numbers”
  • YAGNI + one-liners prompt:Use it when you do not want to install a plugin and only need a prompt-based reduction approach.
    README “Numbers”

Not stated in the README

  • {'text': 'The README does not specify minimum versions for Claude Code, Codex, or other hosts.', 'text_en': 'The README does not specify minimum versions for Claude Code, Codex, or other hosts.'}
  • {'text': 'The README does not provide complete benchmark results across models and repository sizes.', 'text_en': 'The README does not provide complete benchmark results across models and repository sizes.'}
  • {'text': 'The README does not describe the exact behavior or permission scope of the two Node lifecycle hooks.', 'text_en': 'The README does not describe the exact behavior or permission scope of the two Node lifecycle hooks.'}
  • {'text': 'The README does not provide adoption figures or production cases corresponding to the 124,088 stars.', 'text_en': 'The README does not provide adoption figures or production cases corresponding to the 124,088 stars.'}

💡 Deep Analysis

7
Yes I maintain the JavaScript rule source while keeping Claude Code, Codex, Cursor, and OpenClaw rules aligned. After changing skill text, how do I prevent rule-copy drift and stop stale skills from being published?
For: A maintainer changing the JavaScript rule source, publishing six OpenClaw skills, and supporting multiple AI hosts

Yes, because the project already treats multi-host synchronization, generation, and testing as part of the development workflow rather than relying on maintainer memory.

  • The README requires node scripts/check-rule-copies.js after changing the compact rule text, keeping copies for different agents aligned.
  • OpenClaw’s .openclaw/skills/ is generated from skills/; after changing a skill, it must be regenerated, and the test suite fails when the generated output is stale.
  • The project also provides npm test and can publish all six OpenClaw skills to ClawHub. Publishing uses the version from package.json and supports --dry-run for previewing.
  • This shared-rules plus host-adapter structure is appropriate for maintaining both plugin-based and rules-file-based integrations.

The directly executable starting point is to check rule copies and then run the test suite; if a skill changed, the OpenClaw generation step is also required.

  • Development: run `node scripts/check-rule-copies.js` and `npm test` after changing compact rule text
  • Development: `.openclaw/skills/` is generated from `skills/`; run `node scripts/build-openclaw-skills.js` after changing a skill
  • Development: the test suite fails when OpenClaw output is stale
  • Architecture: host adapters plus shared rule content cover plugin-based and rules-file-based tools
node scripts/check-rule-copies.js
Not stated in the README:The README does not specify precedence when rules conflict across AI hosts.;It does not describe rollback or withdrawal procedures after a failed or incorrect ClawHub publication.
No I maintain high-reliability systems containing complex caches, concurrency control, and extensible infrastructure. Even if an AI agent thinks a 120-line cache class can be shortened, I require explicit domain abstractions. Is Ponytail suitable for global enablement?
For: An infrastructure developer maintaining high-reliability systems, complex caches, and concurrency control code who does not want architectural abstractions weakened for fewer lines

No, not for global enablement, because the project explicitly identifies high-reliability systems, complex caches, concurrency control, and extensible infrastructure as cases where complexity should not automatically be treated as a problem.

  • Ponytail preserves security validation, error handling, data integrity, and accessibility, but it still directs the agent to prefer existing capabilities, native features, or the minimum implementation.
  • The README FAQ uses a 120-line cache class as an example: if the user insists, the agent will build it, but its default position questions whether that complexity is necessary.
  • The usage limitations explicitly say the project is not suitable when complexity itself is not the problem, including complex caches, concurrency control, strong-audit systems, high-reliability systems, and long-lived platforms needing domain abstractions.
  • Ponytail changes agent decisions and generated code; it is not a compiler, static analyzer, or runtime security system and cannot replace architectural review or quality controls.

It can be scoped to simple internal tools or routine CRUD areas, but should not be the default rule for these infrastructure modules.

  • FAQ: `What if I really need the 120-line cache class?`
  • usage_limitations: not suitable for high-reliability systems, strong-audit systems, complex caches, concurrency control, extensible infrastructure, or long-lived platforms needing explicit domain abstractions
  • value_proposition: minimization does not mean deleting security validation, error handling, data integrity, or accessibility
  • usage_limitations: the project is not a compiler, static analyzer, or runtime security system
Not stated in the README:The README does not explain how to disable the rules by directory or module.;It provides no measurements of defect rate, review time, or architectural changes for complex infrastructure scenarios.
Yes I need to reproduce the comparison on the FastAPI + React repository with 12 feature tasks, Haiku 4.5, and n=4, while separating Ponytail’s code reduction from caveman’s response compression. Where should I start?
For: An evaluation engineer using Promptfoo to compare the no-skill baseline, Ponytail, and caveman, focusing on Haiku 4.5 code volume and safety

Yes, it is suitable for reproduction, but the result should be treated as a benchmark under specific conditions rather than a universal performance promise.

  • The corrected agentic benchmark uses a real FastAPI + React repository, 12 feature tickets, the same agent with and without the skill, Haiku 4.5, and n=4, scoring the resulting git diff.
  • The reported Ponytail results are -54% LOC, -22% tokens, -20% cost, -27% time, and 100% safety. The README also explains that the older 80–94% figure was partly inflated by conversational padding in single-shot generation.
  • Caveman reduces what the agent says, while Ponytail reduces what the agent builds. Their targets differ, so they can be compared without attributing code-byte changes to caveman.
  • The README provides the Promptfoo configuration entry point, but the supplied material does not include complete environment-locking instructions for reproducing the full agentic benchmark.

Run the original README command first, then verify the model, repository commit, task list, and scoring scripts against the report.

  • Numbers: the corrected agentic benchmark uses FastAPI + React, 12 tasks, Haiku 4.5, n=4, and scores the `git diff`
  • Numbers: Ponytail reports -54% LOC, -22% tokens, -20% cost, -27% time, and 100% safety
  • FAQ: caveman shrinks what the agent says; ponytail shrinks what it builds
  • Older single-shot numbers: reproduction command is `npx promptfoo eval -c benchmarks/promptfooconfig.yaml`
npx promptfoo eval -c benchmarks/promptfooconfig.yaml
Not stated in the README:The README does not provide the complete agentic-benchmark command, dependency versions, or hardware/API environment lock.;It does not clarify whether the Promptfoo command directly reproduces the corrected agentic benchmark rather than the older single-shot experiment.
Yes I use Claude Code with Haiku 4.5 on a FastAPI + React repository and want to reduce diffs, tokens, and API cost across 12 kinds of routine feature tasks; is Ponytail suitable for adding directly to my workflow?
For: A developer using Claude Code with Haiku 4.5 on a real FastAPI + React repository who wants to reduce generated code and cost

Yes, it is suitable because the README’s measured scenario closely matches my model, frameworks, and task profile, although the reported percentages are not guarantees.

  • On a real FastAPI + React repository, across 12 feature tasks with Haiku 4.5 and n=4, Ponytail reduced LOC by 54%, tokens by 22%, cost by 20%, and time by 27% on average.
  • It does not merely ask the model to produce shorter text. It checks whether the feature is needed, looks for an existing implementation, and then prefers the standard library, native platform features, and installed dependencies.
  • The same benchmark reports 100% safety, while the rules explicitly preserve trust-boundary validation, data-loss handling, security, and accessibility.

The README also says gains are near zero when code is already minimal, and some models may become slower because of additional reasoning. Treat the figures as benchmark-specific evidence.

  • Numbers: real FastAPI + React repository, 12 feature tasks, Haiku 4.5, n=4
  • Numbers: ponytail -54% LOC, -22% tokens, -20% cost, -27% time, 100% safe
  • How it works: checks necessity, reuse, standard library, native platform features, and installed dependencies in order
  • How it works: trust-boundary validation, data-loss handling, security, and accessibility are never on the chopping block
Not stated in the README:The README does not report results for my specific Claude Code version, repository size, or task complexity.;It does not describe compatibility details with my existing tests, static analysis, and code review workflow.
Yes I use Cursor, Windsurf, and Cline to build a React internal tool with frequent date-picker and color-picker tasks. I want the agent to avoid installing flatpickr or writing wrapper components; is Ponytail suitable through a rules file?
For: A frontend developer building a React internal tool with Cursor, Windsurf, or Cline who wants native HTML controls preferred over new component dependencies

Yes, especially for well-bounded internal-tool features, because Ponytail’s rules-file integration and examples target exactly this kind of overbuilding.

  • The README lists Cursor, Windsurf, and Cline as hosts that can use a copied rules file, without requiring changes to the runtime architecture of the project being edited.
  • Its date-picker example reduces installing flatpickr, creating a wrapper component, and adding styles to native <input type="date">. The color-picker task is also among the benchmark cases with the largest reduction in overbuilding.
  • The decision ladder tells the agent to read affected code and trace the real flow before preferring existing repository code, the standard library, native platform features, and installed dependencies.
  • This does not mean every React interaction should become a native control. The project notes that browser consistency, design-system requirements, advanced accessibility, and internationalization may justify a more complex implementation.

Use it as an agent decision layer, not as a replacement for a component library.

  • Host coverage: Cursor, Windsurf, and Cline use a rules file or AGENTS.md
  • Before / after: the date picker changes from flatpickr, a wrapper component, and a stylesheet to `<input type="date">`
  • How it works: the agent reads affected code and traces the real flow before preferring reuse, standard library, native platform features, and installed dependencies
  • user_experience.common_pitfalls: native controls may sacrifice complex interaction, browser consistency, design systems, or advanced accessibility
Not stated in the README:The README does not specify the exact destination path for the copied rules file or rule-precedence behavior across host versions.;It provides no validation results for my design system, browser support matrix, or internationalization requirements.
It depends I mainly maintain JavaScript code, but Ponytail’s correctness benchmark invokes Python for email and CSV checks. My machine may only have `python3`, and pandas may be missing. Will that block testing or skill usage?
For: A JavaScript project maintainer using Python helper checks for email and CSV logic on a machine where pandas may not be installed

It depends: missing Python or pandas mainly affects the related correctness checks and does not imply that the entire skill becomes unusable.

  • The README says the correctness benchmark spawns Python for email and CSV checks, trying python3 before python.
  • CSV checks require pandas to be installed locally, so a JavaScript-and-Shell-only environment may not run those checks completely.
  • JavaScript is the main language, but the project data also reports 109,292 lines of Python, indicating that the helper checks are part of the repository rather than an unrelated external step.
  • The README says the agent skills can run without a configuration file, so the absence of Python should not be inferred to disable every Ponytail feature.

If you need the complete test suite, verify that either python3 or python is available and that pandas is installed. If you only need the agent rules, the impact is narrower.

  • Development: the correctness benchmark spawns Python for email and CSV checks, trying `python3` before `python`
  • Development: CSV checks need `pandas` installed locally
  • Project data: JavaScript is the main language and Python accounts for 109,292 lines
  • FAQ: no config file is required; nothing is required
npm test
Not stated in the README:The README does not specify a minimum pandas version or an installation command.;It does not state whether missing pandas causes failure, skipping, or degraded checks.
It depends I use Codex and provide Node.js through nvm or Nix; since the Codex integration runs two Node.js lifecycle hooks, can I treat Ponytail as an always-on agent-governance plugin?
For: A developer running the Codex plugin from a non-interactive shell with Node.js managed by nvm or Nix who must review lifecycle hooks

It depends: default enablement is appropriate only if the non-interactive shell can find Node.js and you have reviewed and trusted the Codex hooks.

  • The README says the Claude Code and Codex plugins run two small Node.js lifecycle hooks. node must be on the non-interactive shell’s PATH; otherwise the skills still work, but always-on activation remains silent.
  • The project changes agent behavior through plugins, skills, commands, and lifecycle logic. It is not a runtime security system, so it should not be treated as a reliable security control if a hook silently fails.
  • The uninstall table lists codex plugin remove ponytail and notes that mode state and configuration may remain outside the plugin directory. The cleanup script should be run before removing the plugin.

Because the Codex installation section is omitted from the supplied README, I cannot confirm the complete install command or default hook permissions from the available material.

  • Install: Claude Code and Codex plugins run two Node.js lifecycle hooks; node must be on the non-interactive shell PATH
  • Install: if node is unavailable, skills still work but always-on activation stays quiet
  • Uninstall: Codex uses `codex plugin remove ponytail`
  • Usage limitations: the project is not a compiler, static analyzer, or runtime security system
Not stated in the README:The supplied README does not include the complete Codex installation command.;It does not specify the exact Codex hook permissions, trigger timing, or failure-handling behavior.

✨ Highlights

  • Claude Code benchmark averages 54% less code
  • 100% safety across 12 FastAPI+React tasks
  • Supports Codex, Claude Code, and Pi plugins
  • GPT-5.5 may become slower and more expensive

🔧 Engineering

  • Uses a YAGNI ladder to prioritize native capabilities
  • The /ponytail-review command audits diffs and returns a delete-list
  • Claude Code and Codex installs use two Node lifecycle hooks

⚠️ Risks

  • The benchmark covers only 12 Haiku 4.5 tasks
  • The ladder may increase cost and latency on GPT-5.5
  • Codex requires reviewing and trusting two lifecycle hooks
  • Uninstalling leaves config and Claude state files

👥 For who?

  • Teams maintaining FastAPI+React projects with Claude Code
  • Agent developers using Codex, Pi, or Devin CLI
  • Open-source maintainers reviewing over-engineered diffs