BLOG

Notes from
the training run.

01

My rat beat Qwen, DeepSeek, Kimi and Gemma at 118M →

How reimplementing eleven frontier architectures from their configs — faithfully, at equal budgets — ended with my own variants sweeping the top five. And what it says about who actually understands these models.

full essay8 min
02

Beating Muon with one state →

AdamW keeps two moments per parameter. Muon keeps one plus an AdamW fallback. Terry keeps one, period — and wins on loss. The full story, with the part I'm not publishing.

full essay6 min
03

Tying Tiktoken at 3× the vocab →

A byte-level BPE trained in 20 minutes on a CPU box matches OpenAI's production encodings on our distribution. Tokenizers are underrated leverage.

full essay5 min
04

Nobody reads the code they run →

Modern ML runs on import statements — ten thousand dependencies nobody chose. Laziness compounds, until memory walls and compute bills make thinking for yourself pay again.

full essay4 min

01 · FULL ESSAY

My rat beat Qwen, DeepSeek, Kimi and Gemma at 118M

Everyone publishes their architecture. Almost nobody reimplements anyone else's. So I did all of them — Qwen3.5's deltanet stack, DeepSeek's compressed attention, GPT-OSS's MoE, Kimi's delta layers, GLM's indexer, Gemma's sliding windows, plus Mamba hybrids, conv challengers and half a dozen shape ablations — each rebuilt faithfully from public configs and modeling code, each cut to the same ~118M budget, each trained on the same batches with the same optimizer for 1,000 steps.

Same budget. Same batches. No excuses left. That's the entire trick of a fair fight, and almost nobody runs one — because running one means you might lose to someone else's idea.

The setup

Thirty-five architectures, one harness. Every entry keeps the same skeleton — RoPE positions, RMSNorm, a tied output head — so the only thing being measured is the idea under test: the attention variant, the wiring, the normalization. Embeddings stay tied because untying them would eat the whole budget and turn the benchmark into a parameter-count contest.

What won

QK-post at 2.2733. My own QK-norm plus post-norm stack — stability stacked on stability. Second place is a two-way tie between my remy-post and remy-sand at 2.28, one hundredth behind. Fourth and fifth: also mine. The first outside entry, Gemma-4, lands sixth at 2.38. DeepSeek's cut diverged outright. Qwen3.5 limped home at 3.53, barely above the vanilla baseline it should have lapped.

The honest caveat, because there is one: 118M parameters and 1,000 steps is a sprint, not a marathon. Late scalers win sprints differently than marathons. The zoo doesn't prove my stack wins at 7B. It proves something arguably more useful — that at equal budgets, with zero tuning advantage, the subtractive idea learns faster than everything the frontier labs shipped.

What it says

Labs publish architectures the way chefs publish recipes — precisely enough to admire, never expecting you to cook it. But configs are configs and math is math. I cooked eleven of them in a month, at eighteen, on rented GPUs, and the things I invented for fun took the top five.

The 1.20B training now is the marathon version of this sprint. Same rat, full distance. We'll see.

02 · FULL ESSAY

Beating Muon with one state

Every popular optimizer is a memory hog and nobody talks about it. AdamW keeps two full-precision states per parameter — mean and variance. On a 1B model that's 8 gigabytes of numbers that aren't even the model, just bookkeeping about the model. Muon, the cool new thing, keeps one state for matrices and quietly falls back to AdamW's two for everything else. So the scoreboard reads: two states, or one-and-secretly-two.

Terry keeps one momentum state per parameter. Dense matrices get grafted, orthogonalized momentum — Newton-Schulz, the same math Muon uses, with magnitudes read off the data — and tables get factored adaptive moments: full adaptivity at two tiny vectors instead of two full states. Half the memory of AdamW, and it doesn't just match. It wins. 7M transformer, 1000 steps, identical everything: Terry 2.93, Muon 3.24, AdamW 3.53.

The spike

Then we scaled to 140M and it exploded. Loss hit 46 by step 50 — same disease, bigger patient: full learning rate on noise, orthogonalized into big confident wrong directions. The fix was the warmup ramps plus learning the actual lesson, that 7M settings don't transfer: retuned down, no explosion at any learning rate from 0.003 to 0.008. Then a 15-lane shootout on an H100: Terry 3.98, Muon 4.09, with 11 of 14 Terry configs beating Muon and the ablations falling behind it. Kill the graft or the table-nesterov and you lose to Muon. Every piece earns its place.

The part I'm not publishing

There is a second variant. Same loss on the transformer bench, noise — at a fraction of the memory. The public tables show its loss column and nothing else, because how it saves is the entire moat and it stays closed until there's an NDA on the table. Guess the family all you want. The recipe isn't in the numbers.

Named after Terry Davis, who wrote everything from scratch. So did I.

03 · FULL ESSAY

Tying Tiktoken at 3× the vocab

Everyone uses Tiktoken and nobody asks whether they should. It's OpenAI's encoding, trained on the internet, 100 to 200 thousand tokens of generalist compromise. If you're training your own model on your own distribution, a custom tokenizer is nearly free leverage — a few hours on a CPU box — and almost nobody takes it. So I built Timmy.

Byte-level BPE, 65,536 vocab, trained in 19.9 minutes on a rented CPU: half a million education docs plus every clean conversation I'd written for the model. The head-to-head against the production encodings: 171 total tokens vs 168 and 170. Effectively tied — at a third of the vocabulary — while winning outright on plain English, chat, and reasoning traces, which is to say, on everything the model will actually do.

The deliberate loss

Two gaps, both on purpose. Code loses, because the training mix was chat-heavy and code has its own statistical texture. Math loses, because single digits are split apart — ten tokens for a phone number instead of three — which trades compression for arithmetic generalization. A model that sees "4" and "2" as atoms reasons about numbers better than one that memorized "42" as a word. That trade is worth it every time.

Why it matters

A smaller vocab is a smaller embedding table, and the embedding table is the single biggest matrix in a small model — 134M parameters at 64k×2048 that I then tie to the output head and delete twice. Tokenizer, architecture, and optimizer aren't three decisions. They're one budget, and Timmy spends its share wisely.

04 · FULL ESSAY

Nobody reads the code they run

Before any of the transformer work, there was Opus: a general-purpose systems language with its own codegen and its own linker, zero dependencies. It links in milliseconds. MSVC takes a second. That gap isn't specialization — Opus does the same job. It's what happens when every line exists because it's needed instead of because a committee, a backlog, and three decades of backwards compatibility put it there.

Modern machine learning runs on import statements. Import the model. Import the tokenizer. Import the optimizer, the scheduler, the trainer, the evaluator, the thing that logs to the thing. An entire profession of library monkeys assembling other people's code into shapes they couldn't build and couldn't debug — script kiddies with salaries and "Software Engineer" in their bios, shipping pip install as a skillset. Underneath: slop all the way down. Slow code written by titled engineers, stacked on slower code written by other titled engineers, none of it ever read by the people whose GPUs burn running it. The title is doing load-bearing work everywhere. Strip the titles and look at the code, and half the field's infrastructure is indefensible.

Laziness compounds

Every abstraction you don't understand is a tax you can't see. The default AdamW loop burns twice the memory it needs because everyone copied the same two-state update without asking what the second state buys. Framework training loops re-materialize logits nobody looks at. Linkers built for every program on earth take a second to link one. None of this is anyone's fault individually — it's what a field looks like when thinking for yourself stops paying, and a generation of engineers who never learned to read down the stack.

It pays now. Memory is the wall and compute is the bill, and both yield exclusively to people willing to open the black boxes. My optimizer exists because I refused to import one. My tokenizer exists because I refused the default. My linker exists because I got annoyed waiting. Annoyance, aimed precisely, is an engineering strategy — and it's one the import-and-pray crowd can't copy, because copying it would require reading something.

So no — I didn't write a compiler to be interesting. I wrote one because depending on forty thousand lines of other people's decisions felt insane, and it turns out that instinct, applied to transformers, beats Muon. The monkeys can keep the libraries. I'll keep the machines.