oneshotlm

What's new

4525 builds · 142 models · 35 prompts · updated 2026-09-02

  1. New Models

    Added Gemini 3.7 Flash

    Gemini 3.7 Flash completed all 35 builds with a 3.00 average across graded samples, showing solid native canvas physics alongside external library loading issues.

    gemini-3.7-flash — Asteroids CLEAN
    gemini-3.7-flash
    gemini-3.7-flash — Boids flocking CLEAN
    gemini-3.7-flash
    gemini-3.7-flash — Bouncing balls in a heptagon CLEAN
    gemini-3.7-flash
    gemini-3.7-flash — 2048 CLEAN
    gemini-3.7-flash
    Read update →
  2. New Models

    Qwen3.8: the 27B meets the 2.4T flagship

    We added the smaller Qwen3.8-27B, so we put it straight up against the 2.4-trillion-parameter Qwen3.8-2.4T-A95B already in the gallery — same family, roughly two orders of magnitude apart in size. Both are reliable builders (34 of 35 and 35 of 35 shipped), so the story is entirely in the quality.

    qwen3.8-27b — Brick breaker BLANK
    qwen3.8-27b
    qwen3.8-2.4t-a95b — Brick breaker CLEAN
    qwen3.8-2.4t-a95b
    qwen3.8-27b — Flow-field particles BLANK
    qwen3.8-27b
    qwen3.8-2.4t-a95b — Flow-field particles CLEAN
    qwen3.8-2.4t-a95b
    Read update →
  3. New Models

    Added Grok 4.6

    xAI's Grok 4.6 ran the full matrix clean — 35 of 35 builds shipped — and averaged 3.43, with a clean five and a clear strength in generative canvas work.

    grok-4.6 — Flow-field particles CLEAN
    grok-4.6
    grok-4.6 — Aquarium breach CLEAN
    grok-4.6
    grok-4.6 — Boids flocking CLEAN
    grok-4.6
    grok-4.6 — Brick breaker CLEAN
    grok-4.6
    Read update →
  4. New Models

    Meta Muse: Spark 1.2 vs Glimmer 30B

    Meta shipped two new Muse models, so we ran both through the same 35 one-shot prompts. Neither failed to ship a single build — 35 of 35 runnable each — so the gap between them is entirely about quality, not reliability.

    muse-spark-1.2 — Flow-field particles CLEAN
    muse-spark-1.2
    muse-glimmer-30b — Flow-field particles CLEAN
    muse-glimmer-30b
    muse-spark-1.2 — Boids flocking CLEAN
    muse-spark-1.2
    muse-glimmer-30b — Boids flocking CLEAN
    muse-glimmer-30b
    Read update →
  5. New Models

    Added Qwen3.8-2.4T-A95B

    Alibaba's 2.4-trillion-parameter Qwen3.8-2.4T-A95B turned in the most consistent run of the batch — 35 of 35 builds shipped, zero blank or crashed, and the tightest quality spread in the set.

    qwen3.8-2.4t-a95b — Arpeggiator pad CLEAN
    qwen3.8-2.4t-a95b
    qwen3.8-2.4t-a95b — Boids flocking CLEAN
    qwen3.8-2.4t-a95b
    qwen3.8-2.4t-a95b — Bouncing balls in a heptagon CLEAN
    qwen3.8-2.4t-a95b
    qwen3.8-2.4t-a95b — Brick breaker CLEAN
    qwen3.8-2.4t-a95b
    Read update →
  6. New Models

    ByteDance Seed: 2.0 Code vs 2.1 Turbo

    Two new ByteDance Seed models went through the matrix together — both 35 of 35 on build reliability, but with opposite temperaments.

    seed-2.0-code — Boids flocking CLEAN
    seed-2.0-code
    seed-2-1-turbo — Boids flocking CLEAN
    seed-2-1-turbo
    seed-2.0-code — Lorenz attractor CLEAN
    seed-2.0-code
    seed-2-1-turbo — Lorenz attractor CLEAN
    seed-2-1-turbo
    Read update →
  7. New Models

    Added Qwen3.8-Max

    Alibaba's new flagship Qwen3.8-Max ran all 35 one-shot prompts with zero build failures — a clean sweep of the matrix, and a solid mid-pack showing on quality.

    qwen3.8-max — 2048 CLEAN
    qwen3.8-max
    qwen3.8-max — Boids flocking CLEAN
    qwen3.8-max
    qwen3.8-max — Rubik's Cube CLEAN
    qwen3.8-max
    qwen3.8-max — Flow-field particles CLEAN
    qwen3.8-max
    Read update →
  8. New Features

    Compare any models across every prompt

    A new side-by-side grid: pick up to 10 models and see how each one built every prompt, in one scrollable table.

    Read update →
  9. Blog

    Kimi K3 vs Claude Opus 4.8, one shot each

    We ran Moonshot's new Kimi K3 and Claude Opus 4.8 through the same 34 one-shot prompts, same harness, and they finish within a hair of each other — K3 at 3.24 average, Opus at 3.12. The interesting part is where each one misses.

    kimi-k3 — Fluid simulation CLEAN
    kimi-k3
    claude-opus-4.8 — Fluid simulation PARTIAL
    claude-opus-4.8
    kimi-k3 — Reaction-diffusion CLEAN
    kimi-k3
    claude-opus-4.8 — Reaction-diffusion BLANK
    claude-opus-4.8
    Read update →
  10. New Models

    Added Moonshot Kimi K3

    Moonshot shipped Kimi K3, a million-token model, so we added it to the matrix and ran all 34 one-shot prompts against it — same harness, same single-shot rules as every other build in the gallery.

    kimi-k3 — Fluid simulation CLEAN
    kimi-k3
    kimi-k3 — Boids flocking CLEAN
    kimi-k3
    kimi-k3 — Fireworks CLEAN
    kimi-k3
    kimi-k3 — Double pendulum CLEAN
    kimi-k3
    Read update →
  11. New Models

    Added GLM-5.2, MiMo v2.5 and 3 more models

    Five more models joined the matrix this week, and we re-ran every prompt against them. The alien-shooter build is a good stress test: it needs a game loop, collision, and input handling all in one file, so the gap between models shows up fast.

    deepseek-v4-pro — Top-down alien shooter CLEAN
    deepseek-v4-pro
    kimi-k2.7-code — Top-down alien shooter PARTIAL
    kimi-k2.7-code
    glm-5.2 — Top-down alien shooter BROKEN
    glm-5.2
    minimax-m2.7 — Top-down alien shooter BLANK
    minimax-m2.7
    Read update →
  12. New Features

    Two-tier artifact evaluation is live

    Every build now carries a two-tier grade. Tier 1 measures the artifact deterministically — how much of the screen moves, whether it responds to input, console and JS errors — and Tier 2 hands those measurements plus a capture to a vision model for the final verdict.

    claude-opus-4.8 — 3D solar system CLEAN
    claude-opus-4.8
    deepseek-v4-pro — 3D solar system CLEAN
    deepseek-v4-pro
    claude-opus-4.8 — Fluid simulation PARTIAL
    claude-opus-4.8
    deepseek-v4-flash — Fluid simulation CLEAN
    deepseek-v4-flash
    Read update →
  13. New Prompts

    New prompt: Fluid simulation

    A real-time WebGL fluid sim in one file — the hardest prompt in the set so far.

    deepseek-v4-pro — Fluid simulation CLEAN
    deepseek-v4-pro
    glm-5 — Fluid simulation CLEAN
    glm-5
    minimax-m2.7 — Fluid simulation CLEAN
    minimax-m2.7
    claude-opus-4.8 — Fluid simulation PARTIAL
    claude-opus-4.8
    Read update →