A small-language-model study

TinyLM Lab

How much can a 27-million-parameter language model learn when it is trained from random numbers on children's stories? And how do you tell real ability apart from memorization?

26.7M parameters 8K byte-level BPE 512-token context 442M training tokens 1 laptop GPU
1.384
test loss (perplexity 3.99) on 15,388 held-out stories
29.9%
of TinyStories' official validation split is copied from its training split
−0.194
loss gained by the modern recipe over a GPT-2-style baseline, same size
6.3 h
full training run on an RTX 3050 Ti Laptop GPU (4 GB)
What we found

Six findings, each backed by a file in the repo

The standard test set leaks

6,582 of the 21,989 stories in TinyStories' official validation split have an exact copy in the training split. Scores reported on that split partly measure memory, including this project's own first run.

A clean test score: 1.384

After removing every copied story, the model scores 1.384 nats per token (perplexity 3.99) on 15,388 untouched stories, evaluated once, after model selection.

It writes new stories, not copies

Its stories are no closer to the training set than real held-out stories are. It reuses more stock phrases, but no long passages: the longest run of words shared with the closest training story is 17.

RoPE matters most; RMSNorm doesn't

Across 15 parameter-matched runs, rotary positions cut loss by 0.136. RMSNorm changed nothing measurable. SwiGLU and grouped-query attention each helped a little, consistently across seeds.

Fluent, but it loses the plot

It forgets constraints ("couldn't swim"), contradicts its prompt, fixes broken glass, and ends long-context stories after a sentence or two. Every failure shown is a real, unedited output.

At this size, the CPU wins

Generating one story at a time, a laptop CPU with a KV cache (89 tokens/s) beats the laptop GPU (~70 tokens/s), where the cache makes no difference at all.

The model

A Llama-style decoder, shrunk to 27M parameters

The architecture follows the recipe used by Llama, Mistral and Qwen: pre-norm RMSNorm, rotary position embeddings, a gated SwiGLU feed-forward layer, grouped-query attention (8 query heads sharing 2 key/value heads), no biases, and an output layer that reuses the input embedding.

The tokenizer and every weight start from scratch. The production model uses Hugging Face's LlamaForCausalLM class. A separate plain-PyTorch implementation of every component reproduces its outputs exactly (difference of 0.0 in the tests) and serves as the model family for the ablations.

Where the parameters go

Two thirds of the model is feed-forward. Tying the output layer to the embedding saves another 4.19M parameters.

One decoder block, repeated 8 times. Residual connections wrap each sub-layer.
Data

TinyStories, and the leak in its test set

TinyStories is about 2.1 million short, simple stories generated by GPT-3.5 and GPT-4. It is narrow enough that a tiny model can write fluently, which makes the interesting questions about coherence and memorization.

A leakage check found that 29.9 % of the official validation stories appear, word for word, in the training split. The test set used here removes them, plus 19 stories that duplicate the validation slice. After that, the test set shows the same leakage profile as validation, which is held out from training by a content hash.

Test stories with an exact (normalized) copy in the training split.
Share of each story's 13-token sequences found anywhere in the 464M-token training set. Some overlap is normal, because the stories are formulaic.
SplitRoleStoriesTokensDropped by packing
traingradient updates2,109,002464,403,968221
validationchoosing the checkpoint10,4742,311,6803
testone final evaluation15,3883,140,096391

The byte-level tokenizer averages 4.09 characters per token, and a story averages 220 tokens (95th percentile: 434), so a 512-token window holds about 95 % of stories whole. Stories are packed back to back, and only the last partial block of each split is dropped.

Training

One clean run, 12,000 steps

AdamW with a cosine learning-rate schedule, 72 stories of 512 tokens per step, and bf16 mixed precision. The run took 6.3 hours on a 4 GB laptop GPU at about 21,100 tokens per second. Validation loss improved at all 48 checks, so the selected model is the final one.

The training loop is written out by hand so that a stopped run resumes exactly: a test checks that 5 steps + stop + 7 steps equals 12 uninterrupted steps, bit for bit. The project's first notebook run had resumed a short run as a long one, which put a visible jump in its loss curve.

Training loss (solid) and validation loss (dashed) per step. The y-axis is cropped: loss starts at 9.12, which is what random guessing over 8,192 tokens scores.
Results

What it scores, and what it writes

MeasureSplitLossPerplexityNote
Final scoretest (15,388 stories)1.38423.99evaluated once, after selection
Model selectionvalidation (10,474)1.38954.01used to choose the checkpoint
First notebook runofficial validation1.41484.1230 % of that split was training data; different tokenizer

Perplexity only compares runs that share a tokenizer and test set, so the first notebook run is listed for context, not as a baseline.

A sample, exactly as generated

Once upon a time, a small rabbit named Benny lived in a big forest. Benny loved to eat carrots and play with his friends. One day, Benny saw a big, scary wolf. The wolf wanted to eat Benny for dinner. Benny was very scared and didn't know what to do. Benny's friend, a wise old owl, saw that Benny was scared and asked him what was wrong. Benny told the owl about the wolf. The owl said, "Don't worry, Benny. I will protect you. I will protect you." The owl flew up high and scared the wolf away. Benny was very happy and thanked the owl. Benny's friends came to his house and they all had a big feast together. Benny was no longer scared of the wolf and he was grateful to his friends for protecting him.

Greedy decoding, so it reproduces exactly. A full arc with problem, helper, resolution and ending, plus two flaws: a repeated sentence, and thanks given to "his friends" when the owl did the protecting.

Memorization

New stories, built from familiar phrases

A fluent story could be stitched together from training data. Every one of the model's 92 test-suite stories was checked against all 464 million training tokens and its nearest training story. So were 200 real held-out stories, because TinyStories is formulaic enough that even unseen text overlaps with training data.

The model reuses more short phrases than human-written stories, but fewer long passages. Its median similarity to the nearest training story is the same as for real held-out stories (0.255 vs 0.250).

The only flagged output matched a stock opening ("there was a little girl named Lily. She loved to play…") and diverged right after it. Not tested here: deliberately prompting with the start of real training stories to try to extract them.

Ablations

Which parts of the modern recipe matter?

Five variants, each with the same 26.7M parameters (only the feed-forward width changes), trained on the same 24.6M tokens in the same order, with the same code. Each was run three times with different seeds, which makes 15 runs. Each bar adds one change to the bar before it.

Mean validation loss with ±1 standard deviation; circles are individual seeds. Lower is better.
  • RoPE: −0.136. By far the biggest effect, about 50× the spread between seeds.
  • RMSNorm: +0.004. No measurable quality gain. Its usual selling point is speed.
  • SwiGLU: −0.029 and GQA: −0.033. Small, but in the same direction for all three seeds. GQA's gain comes from moving attention parameters into a wider feed-forward layer at a fixed total.

These are short runs (about 6 % of the main budget), and the rankings could change with longer training.

Failure analysis

Where it goes wrong

Everything below is an unedited output from the fixed prompt suite. Weak outputs are part of the result.

Contradicts the prompt

"…he really loved trains." → "But I don't like trains," Leo said.

1 of 4 runs. Local context ("Leo was scared") beats the opening line.

Forgets a rule instantly

"You can't play on the swing." … She listened to her mom. She played on the swing.

Both halves are likely on their own; the rule isn't tracked.

Breaks a stated fact

Max was a dog who could not swim … "They swam to the boat."

2 of 4 runs drop the constraint within ~100 tokens.

Loops

"That's not a fish. That's a fish. It's a fish. It's not a fish. It's a fish."

Greedy decoding repeats 4-grams twice as often as sampling.

Impossible physics

The glass broke … soon the glass was fixed.

"Broken thing → mom fixes it" is a common script.

Gives up on long context

After a ~200-token setup about a lost turtle, it writes 33 tokens and ends the story. The lost turtle finds a shell and shows it to grandpa.

All 4 runs end within 33–51 tokens.

Rare words fall apart

"archaeologist" → "The archae arched around the room… the archaeist was so happy"

Rare subword pieces with weak embeddings.

No world knowledge

"What is the capital of France? Answer: Neverse!"

Expected: it only knows children's stories, and it is not an assistant.

Robust where it should be

"Once upon a time" and "Once upon a time," give the same story.

Simple cause and effect is often right, and most stories reach an ending.

Performance

Measured on one laptop

An RTX 3050 Ti Laptop GPU (4 GB) and an Intel Core i5-11300H. No other hardware was measured, so no numbers here are extrapolated to it.

Generating one story at a time. On the CPU the KV cache gives a 6× speedup; on the GPU it makes no difference, because a 27M model is too small to keep the GPU busy.
Training speed. Mixed precision is about 2.1× faster than fp32. Past 4 GB of memory, Windows quietly spills to system RAM and speed collapses instead of failing.
Limitations

What this does not show

  • Nothing about general language ability or world knowledge: the data is simple, synthetic, English-only stories.
  • One full-size run with one seed. Variance comes only from the short ablation runs.
  • Memorization was tested on the model's own stories, not with deliberate extraction attempts.
  • Automatic text metrics are diagnostics, not quality scores. No human or LLM ratings are reported yet.
  • Not an assistant: no instruction tuning, no safety tuning.
Future scope

Where to take it next

  • Longer ablations, to see whether the SwiGLU and GQA gains hold past early training.
  • Deduplicate the training set (15 % of stories appear two or three times) and measure the effect.
  • Extraction tests: prompt with the first 32 tokens of real training stories.
  • Tokenizer sizes: 4K vs 8K vs 16K vocabularies at a fixed parameter budget.
  • A small scaling study (about 7M → 27M → 90M parameters) before claiming anything about scaling.
  • Quantization to int8/int4, scored on the same test set and prompt suite.
  • Longer context, measured on the long-context prompt where the model fails today.
Reproduce it

Every number above comes from a script

git clone https://github.com/sudhanshumukherjeexx/tiny-language-models
cd tiny-language-models
uv sync --extra cu130            # or --extra cu126 / --extra cpu
uv run python scripts/train_tokenizer.py --config configs/tiny-27m.yaml
uv run python scripts/prepare_data.py    --config configs/tiny-27m.yaml
uv run python scripts/train.py           --config configs/tiny-27m.yaml
uv run python scripts/evaluate.py        --config configs/tiny-27m.yaml
uv run python scripts/check_memorization.py --config configs/tiny-27m.yaml \
    --generations runs/tiny-27m/evaluation/generation_samples.jsonl
uv run python scripts/run_ablation.py    --seeds 0 1 2

The charts on this page are drawn from site/data.js, which site/build_data.py generates from the result files in results/.