What's new
1394 builds · 41 models · 34 prompts · updated 2026-07-31
-
Compare any models across every prompt
A new side-by-side grid: pick up to 10 models and see how each one built every prompt, in one scrollable table.
Read update → -
Kimi K3 vs Claude Opus 4.8, one shot each
We ran Moonshot's new Kimi K3 and Claude Opus 4.8 through the same 34 one-shot prompts, same harness, and they finish within a hair of each other — K3 at 3.24 average, Opus at 3.12. The interesting part is where each one misses.
Read update →
CLEANkimi-k3
PARTIALclaude-opus-4.8
CLEANkimi-k3
BLANKclaude-opus-4.8 -
Added Moonshot Kimi K3
Moonshot shipped Kimi K3, a million-token model, so we added it to the matrix and ran all 34 one-shot prompts against it — same harness, same single-shot rules as every other build in the gallery.
Read update →
CLEANkimi-k3
CLEANkimi-k3
CLEANkimi-k3
CLEANkimi-k3 -
Added GLM-5.2, MiMo v2.5 and 3 more models
Five more models joined the matrix this week, and we re-ran every prompt against them. The alien-shooter build is a good stress test: it needs a game loop, collision, and input handling all in one file, so the gap between models shows up fast.
Read update →
CLEANdeepseek-v4-pro
PARTIALkimi-k2.7-code
BROKENglm-5.2
BLANKminimax-m2.7 -
Two-tier artifact evaluation is live
Every build now carries a two-tier grade. Tier 1 measures the artifact deterministically — how much of the screen moves, whether it responds to input, console and JS errors — and Tier 2 hands those measurements plus a capture to a vision model for the final verdict.
Read update →
CLEANclaude-opus-4.8
CLEANdeepseek-v4-pro
PARTIALclaude-opus-4.8
BLANKdeepseek-v4-flash -
New prompt: Fluid simulation
A real-time WebGL fluid sim in one file — the hardest prompt in the set so far.
Read update →
CLEANdeepseek-v4-pro
CLEANglm-5
CLEANminimax-m2.7
PARTIALclaude-opus-4.8