TinyLM Lab
How much can a 27-million-parameter language model learn when it is trained from random numbers on children's stories? And how do you tell real ability apart from memorization?
Six findings, each backed by a file in the repo
The standard test set leaks
6,582 of the 21,989 stories in TinyStories' official validation split have an exact copy in the training split. Scores reported on that split partly measure memory, including this project's own first run.
A clean test score: 1.384
After removing every copied story, the model scores 1.384 nats per token (perplexity 3.99) on 15,388 untouched stories, evaluated once, after model selection.
It writes new stories, not copies
Its stories are no closer to the training set than real held-out stories are. It reuses more stock phrases, but no long passages: the longest run of words shared with the closest training story is 17.
RoPE matters most; RMSNorm doesn't
Across 15 parameter-matched runs, rotary positions cut loss by 0.136. RMSNorm changed nothing measurable. SwiGLU and grouped-query attention each helped a little, consistently across seeds.
Fluent, but it loses the plot
It forgets constraints ("couldn't swim"), contradicts its prompt, fixes broken glass, and ends long-context stories after a sentence or two. Every failure shown is a real, unedited output.
At this size, the CPU wins
Generating one story at a time, a laptop CPU with a KV cache (89 tokens/s) beats the laptop GPU (~70 tokens/s), where the cache makes no difference at all.
A Llama-style decoder, shrunk to 27M parameters
The architecture follows the recipe used by Llama, Mistral and Qwen: pre-norm RMSNorm, rotary position embeddings, a gated SwiGLU feed-forward layer, grouped-query attention (8 query heads sharing 2 key/value heads), no biases, and an output layer that reuses the input embedding.
The tokenizer and every weight start from scratch. The production model uses Hugging
Face's LlamaForCausalLM class. A separate plain-PyTorch implementation of
every component reproduces its outputs exactly (difference of 0.0 in the tests) and
serves as the model family for the ablations.
Where the parameters go
Two thirds of the model is feed-forward. Tying the output layer to the embedding saves another 4.19M parameters.
TinyStories, and the leak in its test set
TinyStories is about 2.1 million short, simple stories generated by GPT-3.5 and GPT-4. It is narrow enough that a tiny model can write fluently, which makes the interesting questions about coherence and memorization.
A leakage check found that 29.9 % of the official validation stories appear, word for word, in the training split. The test set used here removes them, plus 19 stories that duplicate the validation slice. After that, the test set shows the same leakage profile as validation, which is held out from training by a content hash.
| Split | Role | Stories | Tokens | Dropped by packing |
|---|---|---|---|---|
| train | gradient updates | 2,109,002 | 464,403,968 | 221 |
| validation | choosing the checkpoint | 10,474 | 2,311,680 | 3 |
| test | one final evaluation | 15,388 | 3,140,096 | 391 |
The byte-level tokenizer averages 4.09 characters per token, and a story averages 220 tokens (95th percentile: 434), so a 512-token window holds about 95 % of stories whole. Stories are packed back to back, and only the last partial block of each split is dropped.
One clean run, 12,000 steps
AdamW with a cosine learning-rate schedule, 72 stories of 512 tokens per step, and bf16 mixed precision. The run took 6.3 hours on a 4 GB laptop GPU at about 21,100 tokens per second. Validation loss improved at all 48 checks, so the selected model is the final one.
The training loop is written out by hand so that a stopped run resumes exactly: a test checks that 5 steps + stop + 7 steps equals 12 uninterrupted steps, bit for bit. The project's first notebook run had resumed a short run as a long one, which put a visible jump in its loss curve.
What it scores, and what it writes
| Measure | Split | Loss | Perplexity | Note |
|---|---|---|---|---|
| Final score | test (15,388 stories) | 1.3842 | 3.99 | evaluated once, after selection |
| Model selection | validation (10,474) | 1.3895 | 4.01 | used to choose the checkpoint |
| First notebook run | official validation | 1.4148 | 4.12 | 30 % of that split was training data; different tokenizer |
Perplexity only compares runs that share a tokenizer and test set, so the first notebook run is listed for context, not as a baseline.
A sample, exactly as generated
Once upon a time, a small rabbit named Benny lived in a big forest. Benny loved to eat carrots and play with his friends. One day, Benny saw a big, scary wolf. The wolf wanted to eat Benny for dinner. Benny was very scared and didn't know what to do. Benny's friend, a wise old owl, saw that Benny was scared and asked him what was wrong. Benny told the owl about the wolf. The owl said, "Don't worry, Benny. I will protect you. I will protect you." The owl flew up high and scared the wolf away. Benny was very happy and thanked the owl. Benny's friends came to his house and they all had a big feast together. Benny was no longer scared of the wolf and he was grateful to his friends for protecting him.
Greedy decoding, so it reproduces exactly. A full arc with problem, helper, resolution and ending, plus two flaws: a repeated sentence, and thanks given to "his friends" when the owl did the protecting.
New stories, built from familiar phrases
A fluent story could be stitched together from training data. Every one of the model's 92 test-suite stories was checked against all 464 million training tokens and its nearest training story. So were 200 real held-out stories, because TinyStories is formulaic enough that even unseen text overlaps with training data.
The only flagged output matched a stock opening ("there was a little girl named Lily. She loved to play…") and diverged right after it. Not tested here: deliberately prompting with the start of real training stories to try to extract them.
Which parts of the modern recipe matter?
Five variants, each with the same 26.7M parameters (only the feed-forward width changes), trained on the same 24.6M tokens in the same order, with the same code. Each was run three times with different seeds, which makes 15 runs. Each bar adds one change to the bar before it.
- RoPE: −0.136. By far the biggest effect, about 50× the spread between seeds.
- RMSNorm: +0.004. No measurable quality gain. Its usual selling point is speed.
- SwiGLU: −0.029 and GQA: −0.033. Small, but in the same direction for all three seeds. GQA's gain comes from moving attention parameters into a wider feed-forward layer at a fixed total.
These are short runs (about 6 % of the main budget), and the rankings could change with longer training.
Where it goes wrong
Everything below is an unedited output from the fixed prompt suite. Weak outputs are part of the result.
Contradicts the prompt
"…he really loved trains." → "But I don't like trains," Leo said.
1 of 4 runs. Local context ("Leo was scared") beats the opening line.
Forgets a rule instantly
"You can't play on the swing." … She listened to her mom. She played on the swing.
Both halves are likely on their own; the rule isn't tracked.
Breaks a stated fact
Max was a dog who could not swim … "They swam to the boat."
2 of 4 runs drop the constraint within ~100 tokens.
Loops
"That's not a fish. That's a fish. It's a fish. It's not a fish. It's a fish."
Greedy decoding repeats 4-grams twice as often as sampling.
Impossible physics
The glass broke … soon the glass was fixed.
"Broken thing → mom fixes it" is a common script.
Gives up on long context
After a ~200-token setup about a lost turtle, it writes 33 tokens and ends the story. The lost turtle finds a shell and shows it to grandpa.
All 4 runs end within 33–51 tokens.
Rare words fall apart
"archaeologist" → "The archae arched around the room… the archaeist was so happy"
Rare subword pieces with weak embeddings.
No world knowledge
"What is the capital of France? Answer: Neverse!"
Expected: it only knows children's stories, and it is not an assistant.
Robust where it should be
"Once upon a time" and "Once upon a time," give the same story.
Simple cause and effect is often right, and most stories reach an ending.
Measured on one laptop
An RTX 3050 Ti Laptop GPU (4 GB) and an Intel Core i5-11300H. No other hardware was measured, so no numbers here are extrapolated to it.
What this does not show
- Nothing about general language ability or world knowledge: the data is simple, synthetic, English-only stories.
- One full-size run with one seed. Variance comes only from the short ablation runs.
- Memorization was tested on the model's own stories, not with deliberate extraction attempts.
- Automatic text metrics are diagnostics, not quality scores. No human or LLM ratings are reported yet.
- Not an assistant: no instruction tuning, no safety tuning.
Where to take it next
- Longer ablations, to see whether the SwiGLU and GQA gains hold past early training.
- Deduplicate the training set (15 % of stories appear two or three times) and measure the effect.
- Extraction tests: prompt with the first 32 tokens of real training stories.
- Tokenizer sizes: 4K vs 8K vs 16K vocabularies at a fixed parameter budget.
- A small scaling study (about 7M → 27M → 90M parameters) before claiming anything about scaling.
- Quantization to int8/int4, scored on the same test set and prompt suite.
- Longer context, measured on the long-context prompt where the model fails today.
Every number above comes from a script
git clone https://github.com/sudhanshumukherjeexx/tiny-language-models
cd tiny-language-models
uv sync --extra cu130 # or --extra cu126 / --extra cpu
uv run python scripts/train_tokenizer.py --config configs/tiny-27m.yaml
uv run python scripts/prepare_data.py --config configs/tiny-27m.yaml
uv run python scripts/train.py --config configs/tiny-27m.yaml
uv run python scripts/evaluate.py --config configs/tiny-27m.yaml
uv run python scripts/check_memorization.py --config configs/tiny-27m.yaml \
--generations runs/tiny-27m/evaluation/generation_samples.jsonl
uv run python scripts/run_ablation.py --seeds 0 1 2
The charts on this page are drawn from site/data.js,
which site/build_data.py generates from the result files in
results/.