Qwen3.8: the 27B meets the 2.4T flagship
We added the smaller Qwen3.8-27B, so we put it straight up against the 2.4-trillion-parameter Qwen3.8-2.4T-A95B already in the gallery — same family, roughly two orders of magnitude apart in size. Both are reliable builders (34 of 35 and 35 of 35 shipped), so the story is entirely in the quality.
The flagship leads on average, 3.57 to 3.0, but the gap is lopsided rather than uniform. On more than a dozen prompts the two land at the very same grade — boids, the double pendulum, fireworks, Game of Life, Lorenz, matrix rain, a click-to-rotate Rubik's cube — mostly both at four stars. Where the 2.4T pulls away is the hard render-and-game prompts the 27B can't hold together: it crashes brick-breaker, the particle flow field, and reaction-diffusion to black or broken screens where the flagship ships clean fours. Scale mostly buys it *consistency* on the failure-prone prompts, not a higher ceiling.
The surprise runs the other way, too. The little 27B beats its giant sibling on fluid-simulation — a working dye solver where the 2.4T washed the canvas out — and edges ahead on asteroids, a force-directed graph, and Tetris. On the prompts both can do, the 27B keeps pace for a fraction of the size and about a fifth the cost per run ($0.04 vs $0.19 average). The head-to-head prompts are below; each build links to its live artifact and the grade behind it.