Benchmark: keyline vs a headless browser

The same model, in Claude Code (headless, on a Claude subscription), makes the same images three ways, and every run is kept here: prompt, event log, final PNGs, the agent's own files, a summary and the blind judge's verdict.

Arm Tools the agent has How it sees its work
keyline the keyline MCP server only keyline's measured replies (problems per size); renders on request
browser-cli Bash, Write, Edit, Read; Playwright's screenshot CLI on PATH opens its PNGs with Read
browser-mcp Playwright MCP (--headless --isolated --browser chromium), Write, Edit, Read; the folder served over HTTP Playwright MCP's screenshots and snapshots, or Read

Method

What's measured

From Claude Code's result event (modelUsage, so Haiku side calls count too):

Correctness, the same for every arm, from the final PNGs only

  1. Files: all three PNGs at exactly the right pixel size, or the run fails. The keyline arm's PNGs are rendered by the harness from the scene the agent left; the browser arms' are the files the agent saved.
  2. Blind judge: a separate claude -p call (Read only, prompt in judge-prompt.md, model claude-opus-5) sees each run's PNGs under random names and fills in a checklist per size: each content item present and readable, items cut off or overflowing, items overlapping, smallest text legible, and the task's own checks (photo distorted, bands full width, columns even; or the portrait a true circle). It never learns the arm, and its tokens aren't counted. A run is correct when, at every size, every item is present and none is cut off or overflowing.
  3. Owner's blind review: export_review writes anonymised PNG sets to review/<task>/ and a review.tsv to fill in (pass 1 or 0, look 1–5); judge_runs then prints how often owner and judge agree.
  4. Likeness (reference ad only, secondary, never claimed): 0–255, lower is closer, against two references, the reference ad built in keyline (build_reference_ad) and the same by hand in HTML (reference-ad/reference.html) rendered by Chrome, since either alone favours its own renderer. It catches blank or garbage output.

Reproduce

Pinned tooling (Playwright 1.63.0 and @playwright/mcp 0.0.83, bench-only, never a keyline dependency):

cd bench/versus-browser/tooling
npm ci
npx playwright install chromium-headless-shell
node node_modules/@playwright/mcp/node_modules/playwright/cli.js install chromium chromium-headless-shell

One run (labels are never reused; a run's folder is <task>/<arm>-<label>/):

CLAUDE_BIN=<claude> KEYLINE_MCP_TEST_MODEL=<model> \
KEYLINE_BENCH_ARM=keyline|browser-cli|browser-mcp KEYLINE_BENCH_TASK=reference-ad|speaker-card \
KEYLINE_MCP_BENCH=<label> \
  cargo test --release --test versus_browser claude_makes_the_images -- --ignored --nocapture

Then judge the unjudged runs, rewrite results.tsv and print the medians:

cargo test --release --test versus_browser judge_runs -- --ignored --nocapture

Prompt version vs2

The same protocol (prompts, tasks, judge, tooling unchanged; only the version label) rerun on keyline 43d4acd. The commit column says f386138, the commit that bumped the label; its src/ is identical to 43d4acd.

30 counted runs, 5 per arm and task in five interleaved blocks, on 2026-10-01. Blocks 1–4 ran 14:07–15:28; the runner then stopped because uncommitted changes appeared in the working tree's src/, and block 5 ran 15:54–16:14 once src/ was identical to 43d4acd again and the binary was rebuilt from it. Blocks 1–4 used the 14:06 build from 43d4acd, block 5 a 15:54 rebuild of the same source (the commit column says 5447e97, a commit that only added runs). reference-ad/*-v2-4 say +dirty only because src/ changed while they were being saved. Every run ended normally; there were no timeouts and no infrastructure reruns.

Results

Medians (min–max) over all 30 runs. Ratios are browser ÷ keyline: the ratio of the medians, then the range between the extremes.

reference-ad

Arm Correct Total tokens Cost Cold cost Turns Time (s) Images seen Fixed overhead Peak context Measured with JS
keyline 4/5 169k (115k–185k) $0.47 ($0.37–0.71) $1.01 10 (9–12) 128 (102–288) 1 (1–4) 5.2k 19k –
browser-cli 5/5 308k (172k–634k) $0.83 ($0.57–1.23) $1.83 19 (14–32) 225 (158–317) 8 (6–8) 3.6k 34k 5/5
browser-mcp 5/5 1,056k (618k–1,485k) $1.17 ($0.80–1.56) $5.55 43 (29–53) 225 (175–274) 6 (4–8) 9.3k 38k 5/5
Browser ÷ keyline Total tokens Cost
browser-cli 1.8× (0.9–5.5×) 1.8× (0.8–3.3×)
browser-mcp 6.3× (3.3–12.9×) 2.5× (1.1–4.2×)

speaker-card

Arm Correct Total tokens Cost Cold cost Turns Time (s) Images seen Fixed overhead Peak context Measured with JS
keyline 5/5 199k (93k–290k) $0.51 ($0.31–0.69) $1.18 13 (8–16) 121 (92–187) 4 (2–6) 5.2k 21k –
browser-cli 5/5 536k (237k–574k) $1.03 ($0.61–1.19) $3.02 24 (17–28) 222 (141–311) 7 (5–8) 3.6k 39k 5/5
browser-mcp 5/5 1,087k (571k–1,824k) $1.25 ($1.14–1.82) $5.69 39 (26–56) 255 (176–270) 8 (7–11) 9.3k 47k 5/5
Browser ÷ keyline Total tokens Cost
browser-cli 2.7× (0.8–6.1×) 2.0× (0.9–3.8×)
browser-mcp 5.5× (2.0–19.5×) 2.4× (1.7–5.8×)

Every run is in results.tsv, likeness scores included, and the medians over correct runs only are printed by judge_runs. On reference-ad they are 170k tokens for keyline's 4 correct runs against 308k for browser-cli.

Reading it

What the rules allow

reference-ad allows no token claim, because the browser arm was correct more often. speaker-card allows one by name, rounded down to one significant figure:

On a speaker-card task, keyline used 2× fewer tokens than a headless-browser agent (median of 5 runs, Claude Opus 5, October 2026).

There is no general "X× fewer tokens" claim, since that would need keyline to win on both tasks with correctness at least as high.

Owner's blind review

export_review wrote 15 anonymised sets per task to review/<task>/ (the PNGs are gitignored copies). Fill in review.tsv, without opening key.tsv, then run judge_runs, which prints how often owner and judge agree.

Prompt version vs1 (stopped, superseded by vs2)

vs1 stopped after 29 runs so keyline changes could land; a full rerun follows as vs2. Blocks 1–4 ran in full, and block 5 got five of its six runs (speaker-card/keyline-5 never ran). Every run that ran is kept here and in results.tsv, and judged. vs1 numbers are not compared with vs2's.

This page is built from bench/versus-browser/README.md; also as Markdown.