What's new
4525 builds · 142 models · 35 prompts · updated 2026-09-02
-
Added Gemini 3.7 Flash
Gemini 3.7 Flash completed all 35 builds with a 3.00 average across graded samples, showing solid native canvas physics alongside external library loading issues.
Read update →
CLEANgemini-3.7-flash
CLEANgemini-3.7-flash
CLEANgemini-3.7-flash
CLEANgemini-3.7-flash -
Qwen3.8: the 27B meets the 2.4T flagship
We added the smaller Qwen3.8-27B, so we put it straight up against the 2.4-trillion-parameter Qwen3.8-2.4T-A95B already in the gallery — same family, roughly two orders of magnitude apart in size. Both are reliable builders (34 of 35 and 35 of 35 shipped), so the story is entirely in the quality.
Read update →
BLANKqwen3.8-27b
CLEANqwen3.8-2.4t-a95b
BLANKqwen3.8-27b
CLEANqwen3.8-2.4t-a95b -
Added Grok 4.6
xAI's Grok 4.6 ran the full matrix clean — 35 of 35 builds shipped — and averaged 3.43, with a clean five and a clear strength in generative canvas work.
Read update →
CLEANgrok-4.6
CLEANgrok-4.6
CLEANgrok-4.6
CLEANgrok-4.6 -
Meta Muse: Spark 1.2 vs Glimmer 30B
Meta shipped two new Muse models, so we ran both through the same 35 one-shot prompts. Neither failed to ship a single build — 35 of 35 runnable each — so the gap between them is entirely about quality, not reliability.
Read update →
CLEANmuse-spark-1.2
CLEANmuse-glimmer-30b
CLEANmuse-spark-1.2
CLEANmuse-glimmer-30b -
Added Qwen3.8-2.4T-A95B
Alibaba's 2.4-trillion-parameter Qwen3.8-2.4T-A95B turned in the most consistent run of the batch — 35 of 35 builds shipped, zero blank or crashed, and the tightest quality spread in the set.
Read update →
CLEANqwen3.8-2.4t-a95b
CLEANqwen3.8-2.4t-a95b
CLEANqwen3.8-2.4t-a95b
CLEANqwen3.8-2.4t-a95b -
ByteDance Seed: 2.0 Code vs 2.1 Turbo
Two new ByteDance Seed models went through the matrix together — both 35 of 35 on build reliability, but with opposite temperaments.
Read update →
CLEANseed-2.0-code
CLEANseed-2-1-turbo
CLEANseed-2.0-code
CLEANseed-2-1-turbo -
Added Qwen3.8-Max
Alibaba's new flagship Qwen3.8-Max ran all 35 one-shot prompts with zero build failures — a clean sweep of the matrix, and a solid mid-pack showing on quality.
Read update →
CLEANqwen3.8-max
CLEANqwen3.8-max
CLEANqwen3.8-max
CLEANqwen3.8-max -
Compare any models across every prompt
A new side-by-side grid: pick up to 10 models and see how each one built every prompt, in one scrollable table.
Read update → -
Kimi K3 vs Claude Opus 4.8, one shot each
We ran Moonshot's new Kimi K3 and Claude Opus 4.8 through the same 34 one-shot prompts, same harness, and they finish within a hair of each other — K3 at 3.24 average, Opus at 3.12. The interesting part is where each one misses.
Read update →
CLEANkimi-k3
PARTIALclaude-opus-4.8
CLEANkimi-k3
BLANKclaude-opus-4.8 -
Added Moonshot Kimi K3
Moonshot shipped Kimi K3, a million-token model, so we added it to the matrix and ran all 34 one-shot prompts against it — same harness, same single-shot rules as every other build in the gallery.
Read update →
CLEANkimi-k3
CLEANkimi-k3
CLEANkimi-k3
CLEANkimi-k3 -
Added GLM-5.2, MiMo v2.5 and 3 more models
Five more models joined the matrix this week, and we re-ran every prompt against them. The alien-shooter build is a good stress test: it needs a game loop, collision, and input handling all in one file, so the gap between models shows up fast.
Read update →
CLEANdeepseek-v4-pro
PARTIALkimi-k2.7-code
BROKENglm-5.2
BLANKminimax-m2.7 -
Two-tier artifact evaluation is live
Every build now carries a two-tier grade. Tier 1 measures the artifact deterministically — how much of the screen moves, whether it responds to input, console and JS errors — and Tier 2 hands those measurements plus a capture to a vision model for the final verdict.
Read update →
CLEANclaude-opus-4.8
CLEANdeepseek-v4-pro
PARTIALclaude-opus-4.8
CLEANdeepseek-v4-flash -
New prompt: Fluid simulation
A real-time WebGL fluid sim in one file — the hardest prompt in the set so far.
Read update →
CLEANdeepseek-v4-pro
CLEANglm-5
CLEANminimax-m2.7
PARTIALclaude-opus-4.8