ScaffBench

Measuring coding agents on real fullstack scaffolding tasks — time, tokens, cost, and whether the result actually builds.

ScaffBench 3

Full pass rate

150k100k50k0k0%25%50%75%100%cheap + reliable ↗Avg output tokens per scaffoldOx Alpha Free High15%33.8kOpus 5 High31%108.8kGPT-5.6 Sol High44%28.7kFable 5 Low23%22.0kFable 5 High31%58.8k

Leaderboard

Scored across 13 specs on a clean machine. Higher pass rate is better; lower cost, time, and tokens are better.

ModelPassCoreWiredTimeAvg costOut tokStepsLoC
1GPT-5.6 Sol[high]legacy
44%67%96%10.0m$2.1528.7k215.2k
2Fable 5[high]legacy
31%31%95%10.2m$4.9458.8k48
3Opus 5[high]legacy
31%77%96%23.8m$8.34108.8k1089.5k
4Fable 5[low]legacy
23%38%95%4.0m$2.1222.0k37
5Ox Alpha Free[high]
15%62%98%29.2m$0.0033.8k31.2k
0%20%40%60%80%100%

Give your agent the fast path.

One MCP server, every spec-to-scaffold tool the benchmark used. Pick your agent, paste, done.

all supported clients
$ claude mcp add --transport stdio better-fullstack -- npx -y create-better-fullstack@latest mcp

run in your terminal

GitHub Sponsors