Two days ago, AxionLab built a complete Python Jupyter Notebook for a 5M parameters model and told me to train it on a RTX 5090 on Runpod and it instantly ran - almost like a single-shot of an LLM if you know what I mean 😂

Five million parameters. GatedDeltaNet under the hood. About 300 million tokens of training data. That is not a typo. It's real 😭

While a lot of models in the community think that you'll need to throw like several billion tokens on a 5M just to look competitive, this little thing walked into the hard benchmarks tests - PIQA, HellaSwag, ARC-Easy, ARC-Challenge - and nearly sat down at the same table as CMA-8M, Qana-mini-5M, and GPT-S2-5M.

It's not a lie - it's true. Keep reading and you'll find out. :D

The setup, no fluff

Architecture: GatedDeltaNet
Size: 5M parameters
Data: ~300M tokens
Age: trained two days ago
Mood: slightly feral 😂

Most 5M class models you actually respect were fed some billion tokens. We gave this one a small fraction of that and it still was amazingly competitive!

That's the whole thing.

Why GatedDeltaNet hits different

Transformers are the default for a reason. They also spend like it 😭

GatedDeltaNet is a linear-ish sequence model with a gated delta rule: memory that updates, forgets on purpose, and does not make you pay quadratic rent for every extra token. At 5M params that is not a nice to have, it's almost obligatory.

Fast mixing. Controlled retention. Recurrence that actually remembers what it should. The kind of architecture that makes a small model feel bigger than the spreadsheet says.

If you have been waiting for a reason to care about gated delta style sequence models outside of papers: this is one. 🧠

Keep reading!

The hard boards

We care about the benches that do not clap for you.

PIQA --> physical commonsense. Does the model know how the world actually behaves?

HellaSwag --> completion that looks easy to humans and eats small models alive.

ARC-Easy and ARC-Challenge --> science questions, including the ones that are supposed to hurt.

Against CMA-8M, Qana-mini-5M, and GPT-S2-5M - names the community already treats as near the front of this weight class (soon not any longer 😏🔥) - our 5M GatedDeltaNet closed most of the distance.

Wait, what? Read that again. Smaller or matched size. Far less data. Same brutal evals. Almost there. Amazing 🤩

When compute is scarce, data efficiency is not a footnote. It is the product.

300M vs several billion is not a flex. It is the point.

Training on ~300M tokens is a constraint we chose to respect, not a bug we will quietly patch in the appendix.

If an architecture only looks good after you drown it in tokens, you did not find a better model. You found a more expensive average 😭

GatedDeltaNet at this scale is saying something louder:

less data + strong structure --> real reasoning signal

That arrow is the whole lab thesis. ✨

The real benchmark numbers

Here you can see the model performing. Remember: ~300M tokens. Not some billions of it!

ModelTrain tokensARC-EasyARC-ChallengeHellaSwagPIQA
Supra-5M-GatedDeltaNet~300 😏33.29%17.83%26.10%54.19%
fromziro/Qana-mini-5M~21B 😭34.97%23.21%27.60%57.18%
AxiomicLabs/GPT-S2-5M~75B 😭33.92%22.87%27.87%57.56%
User01110/CMA-8M~21B 😭35.35%23.29%28.19%58.22%

This is the whole thing. All values are acc_norm. And we're not lying, faking or anything else the values. All is 1:1 spit out from lm_eval. 1:1. For you. Read it again and again and enjoy how good GatedDeltaNet is :D

What we are not saying

We are not pretending 5M params solved intelligence. Surely not 😂

We are not shipping a silent "trust us" leaderboard with mystery evals.

We are not done. And something improved will come VERY soon! (stay tuned)...

The model is two days old. The architecture is clearly not. That combination is what made this model really great.

What is next

More tokens, carefully. Not a mindless 10x dump.

Better data mix. Tighter evals. The same spine.

If GatedDeltaNet can almost hang with the community SOTA club on a small fraction of the usual diet, we want to know where the ceiling actually is - not where the default recipe says it should be. And we'll find it out - for you 🤗

One last thing

SupraLabs exists to find the models that punch up.

This one punches. Defintely.

GatedDeltaNet. 5M. ~300M tokens. Two days old. Already bothering the names people quote.

#research #small-model #gated-delta-net #GDN #edge-ai #tinyml