◈

Veri

from-scratch 1.20B. built at 18. no shortcuts.

transformer · tokenizer · optimizer · all hand-written · training live

ARCHITECTURE · MODEL EXPLORER

Every tensor, on the table.

real shapes · real param math · source lines · click any node

veri / block
click a node for details · ⤢ opens its view
dataresidualside input
24 layers

thick underline = sliding-2048 local · plain = full global · click to open block

Why it wins

QK-norm stabilizes late training (2.24 vs 2.42 in arch bench) · sliding halves attention on even layers free · diff heads subtract common-mode noise (2.28 fixed-zoo 2nd) · tied head saves 134M.

TERRY

One state. Orthogonalized.

matrices → grafted + orthogonalized momentum · tables → factored adaptive · still one state

Remy λ=0.62

Terry vs the world.

Loss · lower wins

7M · 1000 steps · same init, batches, seeds

terry2.93
muon3.24
adamw3.53
vs muon
−9.5%
vs adamw
−16.9%
private · 300-step
3.86 vs 3.84

terry adds factored adaptive tables, grafted matrices and table nesterov, swept across 8 waves of configs. sign-graft ablation flopped at 5.02.

Memory · optimizer states

live states · same init, batches, seeds

adamw56.3 MB
muon79.7 MB
terry28.4 MB
private4.0 MB
140M · H100 · 400 steps · same packs, same seeds

113M bf16 · loss and live states, one setup — states in each opt's native dtype

private3.9561.0 MB states
terry3.95226.9 MB states
muon4.09380.4 MB states
adamw4.59452.0 MB states

Everyone was invited.

35 archs · 118M · 1000 steps · one harness. click any row for the recipe.