01 · FULL ESSAY
My rat beat Qwen, DeepSeek, Kimi and Gemma at 118M
Everyone publishes their architecture. Almost nobody reimplements anyone else's. So I did all of them — Qwen3.5's deltanet stack, DeepSeek's compressed attention, GPT-OSS's MoE, Kimi's delta layers, GLM's indexer, Gemma's sliding windows, plus Mamba hybrids, conv challengers and half a dozen shape ablations — each rebuilt faithfully from public configs and modeling code, each cut to the same ~118M budget, each trained on the same batches with the same optimizer for 1,000 steps.
Same budget. Same batches. No excuses left. That's the entire trick of a fair fight, and almost nobody runs one — because running one means you might lose to someone else's idea.
The setup
Thirty-five architectures, one harness. Every entry keeps the same skeleton — RoPE positions, RMSNorm, a tied output head — so the only thing being measured is the idea under test: the attention variant, the wiring, the normalization. Embeddings stay tied because untying them would eat the whole budget and turn the benchmark into a parameter-count contest.
What won
QK-post at 2.2733. My own QK-norm plus post-norm stack — stability stacked on stability. Second place is a two-way tie between my remy-post and remy-sand at 2.28, one hundredth behind. Fourth and fifth: also mine. The first outside entry, Gemma-4, lands sixth at 2.38. DeepSeek's cut diverged outright. Qwen3.5 limped home at 3.53, barely above the vanilla baseline it should have lapped.
The honest caveat, because there is one: 118M parameters and 1,000 steps is a sprint, not a marathon. Late scalers win sprints differently than marathons. The zoo doesn't prove my stack wins at 7B. It proves something arguably more useful — that at equal budgets, with zero tuning advantage, the subtractive idea learns faster than everything the frontier labs shipped.
What it says
Labs publish architectures the way chefs publish recipes — precisely enough to admire, never expecting you to cook it. But configs are configs and math is math. I cooked eleven of them in a month, at eighteen, on rented GPUs, and the things I invented for fun took the top five.
The 1.20B training now is the marathon version of this sprint. Same rat, full distance. We'll see.