cli-ck
Benchmarks

cli-ck vs. 4 Coding Agent CLIs

Same five coding tasks, the same model (GPT-4o mini), and the same grader against Codex CLI, Goose, OpenCode, and Aider: we actually run the code each agent produces. No self-reported scores — and no cherry-picking where cli-ck doesn't come out ahead. This run includes an experimental single-shot fast path (still in review, not yet shipped) that skips the tool-calling loop entirely for small, simple edits.

Run August 22, 2026GPT-4o mini on every tool
5/5
cli-ck tasks passed
3.24x faster
vs. OpenCode
91% fewer tokens
vs. Codex CLI
Head to head

The honest tradeoff

cli-ck beats the other two tool-calling agents tested — Codex CLI and OpenCode — on both speed and tokens outright, and is the fastest tool in this run overall. Two of the five tasks here took a new single-shot fast path that skips the tool-calling loop for small, self-contained edits — those two dropped 15–17x in tokens on their own. The other three fell back to the normal tool-calling loop, which is by design: it only takes the fast path when it's confident, and falls back rather than risk a wrong edit. Aider still uses fewer tokens per task overall: it skips tool-calling entirely, for every task, and rewrites whole files in one shot. Goose uses fewer tokens too, but its fix for the off-by-one bug didn't actually pass the test — every other tool, including Goose on its remaining four tasks, reached correct code.

cli-ck vs. Codex CLI
agentic tool-calling
1.99x faster
speed
91% fewer tokens
tokens
5/5 pass
cli-ck vs. Goose
agentic tool-calling
1.15x faster
speed
1.4x more tokens
tokens
4/5 pass
cli-ck vs. OpenCode
agentic tool-calling
3.24x faster
speed
94% fewer tokens
tokens
5/5 pass
cli-ck vs. Aider
whole-file edit
1.05x faster
speed
5.3x more tokens
tokens
5/5 pass
Per-task results

Every run, side by side

Time and token totals per task. Tokens are input + output tokens reported by each tool for that run.

Taskcli-ckCodex CLIGooseOpenCodeAider
Pure function + test
Write factorial(n) and a test that checks it with node's assert module.
6.3s6.9k tok
Pass
10.2s42.1k tok
Pass
7.5s4.1k tok
Pass
13.7s59.5k tok
Pass
7.0s0.9k tok
Pass
Off-by-one bug fix
A failing test expects an inclusive range sum; fix the loop without touching the test.
2.1s0.6k tok
Pass
Single-shot fast path — skipped the tool-calling loop entirely
10.8s70.3k tok
Pass
1.7s3.5k tok
Fail
Goose's fix didn't actually resolve the bug (test still failed: 10 !== 15)
27.3s81.1k tok
Pass
6.7s1.1k tok
Pass
Dedup refactor
Extract a shared helper to remove duplicated averaging logic across two functions.
15.1s12.6k tok
Pass
20.2s71.7k tok
Pass
12.3s4.2k tok
Pass
20.8s102.4k tok
Pass
7.4s1.3k tok
Pass
Boundary error handling
Make divide-by-zero throw a specific error message; leave normal division untouched.
2.2s0.6k tok
Pass
Single-shot fast path — skipped the tool-calling loop entirely
10.7s83.5k tok
Pass
7.0s4.0k tok
Pass
24.9s122.8k tok
Pass
6.1s1.0k tok
Pass
CLI arg parser
Parse --key value pairs and standalone --flag booleans into an object, plus a test.
7.4s7.4k tok
Pass
13.7s42.8k tok
Pass
9.5s4.4k tok
Pass
20.3s81.1k tok
Pass
7.3s1.0k tok
Pass
Total33.0s · 28.0k tok65.8s · 310.4k tok37.9s · 20.2k tok107.0s · 446.9k tok34.5s · 5.3k tok

Methodology

All five tools ran the same five self-contained coding tasks — a pure function, a seeded off-by-one bug fix, a dedup refactor, a boundary/error-handling fix, and a small CLI arg parser — against fresh scratch git repos, using GPT-4o mini through each tool's own non-interactive mode: cli-ck's headless agent runner, codex exec, goose run, opencode run, and aider --message. Every tool authenticated with the same raw OpenAI API key — no subscription credits, no vendor account tier differences. None got project instructions, prior context, or retries.

Grading is not self-reported. After each run, we execute the code the agent produced with a plain node and check the real exit code — a tool's own final message saying it succeeded doesn't count.

The four competitors split into two architectures: cli-ck, Codex CLI, Goose, and OpenCode all run an agentic loop of individual read/write/shell tool calls; Aider defaults to a leaner whole-file-rewrite edit format with far fewer model round trips. That split shows up directly in the numbers — it's the most likely reason Aider and Goose use so many fewer tokens here, not a difference in code quality.

Caveats

Want to reproduce this? The benchmark harness is open source in the cli-ck repo.