cli-ck vs. 4 Coding Agent CLIs
Same five coding tasks, the same model (GPT-4o mini), and the same grader against Codex CLI, Goose, OpenCode, and Aider: we actually run the code each agent produces. No self-reported scores — and no cherry-picking where cli-ck doesn't come out ahead. This run includes an experimental single-shot fast path (still in review, not yet shipped) that skips the tool-calling loop entirely for small, simple edits.
The honest tradeoff
cli-ck beats the other two tool-calling agents tested — Codex CLI and OpenCode — on both speed and tokens outright, and is the fastest tool in this run overall. Two of the five tasks here took a new single-shot fast path that skips the tool-calling loop for small, self-contained edits — those two dropped 15–17x in tokens on their own. The other three fell back to the normal tool-calling loop, which is by design: it only takes the fast path when it's confident, and falls back rather than risk a wrong edit. Aider still uses fewer tokens per task overall: it skips tool-calling entirely, for every task, and rewrites whole files in one shot. Goose uses fewer tokens too, but its fix for the off-by-one bug didn't actually pass the test — every other tool, including Goose on its remaining four tasks, reached correct code.
Every run, side by side
Time and token totals per task. Tokens are input + output tokens reported by each tool for that run.
| Task | cli-ck | Codex CLI | Goose | OpenCode | Aider |
|---|---|---|---|---|---|
Pure function + test Write factorial(n) and a test that checks it with node's assert module. | 6.3s6.9k tok Pass | 10.2s42.1k tok Pass | 7.5s4.1k tok Pass | 13.7s59.5k tok Pass | 7.0s0.9k tok Pass |
Off-by-one bug fix A failing test expects an inclusive range sum; fix the loop without touching the test. | 2.1s0.6k tok PassSingle-shot fast path — skipped the tool-calling loop entirely | 10.8s70.3k tok Pass | 1.7s3.5k tok FailGoose's fix didn't actually resolve the bug (test still failed: 10 !== 15) | 27.3s81.1k tok Pass | 6.7s1.1k tok Pass |
Dedup refactor Extract a shared helper to remove duplicated averaging logic across two functions. | 15.1s12.6k tok Pass | 20.2s71.7k tok Pass | 12.3s4.2k tok Pass | 20.8s102.4k tok Pass | 7.4s1.3k tok Pass |
Boundary error handling Make divide-by-zero throw a specific error message; leave normal division untouched. | 2.2s0.6k tok PassSingle-shot fast path — skipped the tool-calling loop entirely | 10.7s83.5k tok Pass | 7.0s4.0k tok Pass | 24.9s122.8k tok Pass | 6.1s1.0k tok Pass |
CLI arg parser Parse --key value pairs and standalone --flag booleans into an object, plus a test. | 7.4s7.4k tok Pass | 13.7s42.8k tok Pass | 9.5s4.4k tok Pass | 20.3s81.1k tok Pass | 7.3s1.0k tok Pass |
| Total | 33.0s · 28.0k tok | 65.8s · 310.4k tok | 37.9s · 20.2k tok | 107.0s · 446.9k tok | 34.5s · 5.3k tok |
Methodology
All five tools ran the same five self-contained coding tasks — a pure function, a seeded off-by-one bug fix, a dedup refactor, a boundary/error-handling fix, and a small CLI arg parser — against fresh scratch git repos, using GPT-4o mini through each tool's own non-interactive mode: cli-ck's headless agent runner, codex exec, goose run, opencode run, and aider --message. Every tool authenticated with the same raw OpenAI API key — no subscription credits, no vendor account tier differences. None got project instructions, prior context, or retries.
Grading is not self-reported. After each run, we execute the code the agent produced with a plain node and check the real exit code — a tool's own final message saying it succeeded doesn't count.
The four competitors split into two architectures: cli-ck, Codex CLI, Goose, and OpenCode all run an agentic loop of individual read/write/shell tool calls; Aider defaults to a leaner whole-file-rewrite edit format with far fewer model round trips. That split shows up directly in the numbers — it's the most likely reason Aider and Goose use so many fewer tokens here, not a difference in code quality.
Caveats
- This is one run of five tasks, not a statistically averaged suite — treat it as a snapshot, not a permanent ranking. Individual task times and token counts vary noticeably between runs; this page doesn't hide that behind a favorable run.
- Token counts are each tool's own reported input + output tokens; a meaningful share of some tools' input tokens were served from cache (not fully billed at first-use rates), but the raw count is what we report here.
- Warp's
oz agent runwas originally in scope too, but its plain (non-cloud) exec command is credit-gated with no bring-your-own-key path, so it isn't included. - cli-ck's numbers here reflect a set of token-efficiency fixes (lite system prompt, trimmed idle tools, history pruning, a code-enforced do-not-touch guard, and a verify-before-declaring-done hint) — see PR #314 for the exact changes and what was deliberately left out (a more aggressive tool cut was tested but not shipped, since it would have cost real capability like codebase search and subagents).
- On top of those fixes, this run also used an experimental single-shot fast path: for a small, freshly-started scope, one structured-output call proposes edits directly instead of running the full tool-calling loop, verified with a syntax check before being accepted. It only took that path on 2 of 5 tasks here — the other 3 didn't clear its own eligibility or verification bar and fell back to the normal agent loop at roughly the same cost as without it. Neither this nor the fixes above have shipped in cli-ck yet — see PR #315 (stacked on PR #314, both still open) and isn't wired into the live chat UI yet — it's a headless-only capability for now, deliberately kept out of the streaming/approval UX until it's proven out further.
Want to reproduce this? The benchmark harness is open source in the cli-ck repo.