Article

Build a Mini GPT Language Model from Scratch

GPT-style models do one job, over and over: look at the words so far, then predict the next word. This project builds that job by hand in PyTorch. The result is a mini language model trained on ten short sentences about tea, rain, festivals, and cricket.

The code lives in Build a Mini GPT Model From Scratch Using PyTorch. There are two files. transformer_blocks.py holds the attention and transformer pieces. demo.py holds the sentences, the model, the training loop, and the text generator.

The whole path

Sentences→ Word IDs→ Embeddings→ 2 transformer blocks→ Next-word scores→ A new word

1. Start with a tiny library of sentences

The training text is a Python list of ten lines, joined into one long string. A space is added after each line so the sentences stay separate when they are joined.

corpus = [
      "hello friends how are you",
      "the tea is very hot",
      "my name is Aarohi",
      ...
  ]
  text = " ".join(sentence + " " for sentence in corpus)

That text is the model’s entire world. It can learn patterns that appear here, such as "the tea is …" or "it is raining in …". It has no other knowledge.

2. Turn every word into a number

Neural nets work on numbers. The code splits the text on spaces, keeps the unique words, and gives each word an ID.

  • word2idx turns a word into a number, like "tea" → 7.
  • idx2word turns that number back into a word when the model writes.
  • vocab_size is how many different words exist. With this list, that is about 41 words.

The full text then becomes one long tensor of IDs, called data. Because the vocabulary is built with a set, the ID numbers can change from one run to the next. The meaning of the mapping stays the same: each distinct word gets one slot.

A useful picture

"the tea is very hot" might become something like [3, 7, 1, 12, 9]. The exact numbers depend on the shuffle of the vocabulary. Training only cares that the same word always keeps the same ID inside one run.

3. The lesson is always "what comes next?"

The model never sees a question-and-answer pair. It sees a short window of words and is asked to predict the word that follows each position.

The window length is block_size = 6. A helper called get_batch picks 16 random starting points in the ID list. For each start i:

  • x is words i through i + 5 — the input.
  • y is words i + 1 through i + 6 — the targets, shifted one step forward.

If the input is the tea is very hot today, the targets are tea is very hot today morning. Position 0 should predict "tea". Position 1 should predict "is". Every position is a next-word question. That is why one short snippet teaches six predictions at once.

4. The small settings

Name in codeValueWhat it means
block_size6The model reads six words at a time.
embedding_dim32Each word is represented by 32 numbers.
n_heads2Two attention heads look at the context in parallel.
n_layers2Two transformer blocks are stacked.
lr0.001AdamW step size.
epochs1500Number of training updates.
batch size16Sixteen random windows per update.

With a vocabulary of about 41 words, the whole network has roughly 28,000 parameters. A modern GPT has billions. The shape of the computation is the part worth learning here.

5. TinyGPT, from the outside

The class TinyGPT is the full model. Its forward method does five things:

  1. Look up a token embedding for every word ID.
  2. Look up a position embedding for "1st word, 2nd word, …" and add it.
  3. Pass that mix through two transformer blocks.
  4. Apply a final layer norm.
  5. Project each position to a score for every word in the vocabulary.

Those scores are called logits. If targets were provided, the loss is cross-entropy: a measure of how surprised the model was by the true next word. Lower loss means the correct word got a higher score.

Inside one forward pass

Linear head → one score per vocabulary word
Final layer norm
Transformer block 2
Transformer block 1
Token embedding + position embedding

6. Embeddings: a word becomes a movable list of numbers

nn.Embedding(vocab_size, 32) is a lookup table. Row 7 is the current meaning of whatever word has ID 7. At the start those rows are random. Training nudges them so words that play similar roles end up with related vectors.

Order matters too. "hot tea" and "tea hot" use the same words in a different place. position_embedding is a second table with one row per slot in the window: positions 0 through 5. The model adds the word vector and the position vector. After that addition, each token carries both "which word I am" and "where I sit".

The position table only has six rows. That is why generation later keeps only the latest six tokens. The model has no learned vector for a seventh position.

7. One attention head: who should I listen to?

SelfAttentionHead lets every word gather information from earlier words. For the incoming vectors it builds three views with linear layers and no bias:

  • Query — what this word is looking for.
  • Key — what each word offers as a label.
  • Value — the content that word can pass along.

Each head is smaller than the full embedding. With 32 dimensions and 2 heads, head_size is 16. Queries, keys, and values in one head live in that 16-number space.

The attention scores are the match between queries and keys:

scores = query @ key.transpose(-2, -1) / sqrt(embedding_dim)
  scores = mask_future(scores)
  weights = softmax(scores)
  output = weights @ value

Dividing by the square root of the embedding size keeps the raw matches from becoming huge before softmax. Softmax turns each row into weights that add up to 1. The output is a weighted mix of the value vectors. A word that gets weight 0.8 contributes most of its value; a word that gets weight 0.05 barely contributes.

The triangle mask

The head stores a lower-triangular matrix of ones, named tril, with register_buffer. It is part of the module, and it is fixed. It is never updated by the optimizer.

Any score above the diagonal — a look into the future — is replaced with negative infinity. Softmax then gives those positions weight 0. When the model is predicting the word after "the tea", it may use "the" and "tea". It may not use "is", because "is" is the answer it is supposed to guess.

That one-way view is what makes this a causal, or decoder-style, language model. The same rule GPT uses when it writes left to right.

8. Two heads, then a small private network

MultiHeadAttention runs two heads on the same input and concatenates their outputs. A final linear layer, proj, mixes them back to 32 dimensions. One head can learn to track nearby grammar. The other can learn a different habit, such as "festival words follow festival words". With only two heads and a tiny corpus, those habits stay simple, and the idea is already the real one.

After attention, FeedForward works on each position on its own. It expands 32 numbers to 128, applies ReLU, then shrinks back to 32. Attention mixes words together. The feed-forward layer is where each position thinks about what it just heard.

Linear(32 → 128) → ReLU → Linear(128 → 32)

The factor of four — 32 to 128 — is the usual transformer width expansion, kept here even though the model is tiny.

9. A transformer block keeps the original message

Class Block is one full layer. It uses pre-norm residuals, which means normalize first, transform, then add the result back onto the original input:

x = x + attention(layer_norm_1(x))
  x = x + feed_forward(layer_norm_2(x))

One block

Add the feed-forward result back
Feed-forward: 32 → 128 → 32
Layer norm
Add the attention result back
Multi-head causal attention
Layer norm

The skip connection matters. The block adds a correction on top of x. Information from earlier layers can pass forward even if a particular attention pattern is still learning. Layer norm keeps the scale of those vectors steady so training is less jumpy.

TinyGPT stacks two of these blocks in an nn.Sequential. A final layer norm cleans the representation, and self.head maps 32 dimensions to one logit per vocabulary word.

10. Training is a short loop

The optimizer is AdamW with learning rate 1e-3. Each of the 1500 steps does the same five lines of work:

  1. Draw a batch of 16 windows.
  2. Run the model and read the loss.
  3. Zero the old gradients.
  4. Backpropagate the loss.
  5. Update every embedding and every weight.
for step in range(1500):
      xb, yb = get_batch()
      logits, loss = model(xb, yb)
      optimizer.zero_grad()
      loss.backward()
      optimizer.step()

Every 300 steps the code prints the loss. On this corpus the number should fall. The model is memorizing the local habits of these sentences. A falling loss here means "I am getting better at this tiny book," which is exactly the right goal for the demo.

11. Writing is sampling, one word at a time

generate starts from a context tensor. The demo uses the single word "hello" and asks for 15 new tokens.

Each round:

  1. Keep only the last six tokens, because that is the position table’s width.
  2. Run the model.
  3. Take the logits at the final position only. Earlier positions were needed as context; the new word comes from the last row.
  4. Turn logits into probabilities with softmax.
  5. Draw one word with torch.multinomial, like a weighted die.
  6. Append that word and repeat.

Sampling means the same prompt can produce different continuations. A word with probability 0.7 is likely, and a word with probability 0.05 can still appear. After 15 draws, the IDs are mapped back through idx2word and printed as a sentence.

What a good run looks like

You should see phrases that echo the corpus: tea, rain, Delhi, festivals, cricket, names. You should also see odd jumps. Forty-one words and ten sentences are enough to show the mechanism. They are a small sample of English.

12. Walk one prediction by hand

Suppose the context is hello friends how and the model is about to choose the fourth word.

  1. "hello", "friends", and "how" become three rows of 32 numbers, plus their position rows.
  2. In each attention head, "how" forms a query and compares it with the keys of "hello", "friends", and itself. Future positions are masked.
  3. The values of those earlier words are mixed using the attention weights.
  4. Both heads are joined, then the feed-forward layer updates the vector for "how".
  5. The second block repeats that listen-then-think pattern.
  6. The head turns the final vector into about 41 scores. Softmax makes them probabilities. Multinomial picks one ID, maybe the ID for "are".

The next round’s context is hello friends how are, trimmed to six tokens if it ever grows longer. That is the entire generation algorithm.

13. How this sits next to a full GPT

This mini modelA full GPT-style model
Whole words, split on spacesSubword tokens from a tokenizer such as byte-pair encoding
About 40 words of vocabularyTens of thousands of tokens
Context of 6 tokensThousands of tokens of context
2 layers, 2 heads, width 32Dozens of layers, many heads, width in the thousands
ReLU in the feed-forward layerOften GELU or a gated activation
Ten handmade sentencesA very large text collection
Plain learned position rowsOften rotary or other relative position methods

The shared core is unchanged: embeddings, causal multi-head attention, a position-wise feed-forward network, residual connections, layer norm, next-token cross-entropy, and autoregressive sampling.

14. Small details that make the code easier to trust

  • Device check. The script prints the PyTorch version and whether a GPU is visible. The tensors in this demo are created on the default device, so a CUDA printout is informational unless you move the model and data yourself.
  • Attention scale. The code divides by the square root of C, the full embedding width (32). Many textbooks divide by the square root of the head size (16). Both keep the scores in a sane range. The textbook choice matches the original scaled dot-product formula a little more closely.
  • No dropout. This demo always uses every connection. Larger models often drop random units during training so they rely on more than one path.
  • Loss shape. Logits of shape (batch, time, vocab) are flattened to (batch × time, vocab) so one cross-entropy call scores every position in the batch.
  • Capital letters count. "Aarohi", "Delhi", and "Mumbai" are their own vocabulary entries. "the" and "The" would be different words if both appeared.

15. The idea to keep

A language model is a next-word machine. This one learns that job on a handful of sentences, with two small transformer blocks you can read from top to bottom. Attention decides which earlier words matter. The mask keeps the answer hidden. The feed-forward layer updates each word after it has listened. Training rewards the correct next ID. Generation rolls a weighted die and feeds the result back in.

Once those moves are familiar, a larger GPT is the same diagram with a bigger vocabulary, a longer memory, more blocks, and much more text.

Walkthrough of the public lecture code in codewithaarohi/Build-a-Mini-GPT-Model-From-Scratch-Using-PyTorch: demo.py and transformer_blocks.py.

Download ANSNEW APP For Ads Free Experiences!
✖
Yamin Hossain Shohan
Software Engineer, Researcher & Digital Creator

I’m a researcher, software engineer and digital creator focused on applying technology and creative problem-solving to build useful tools, explore new ideas and create engaging digital content.

Copyright Disclaimer

✖

All the information is published in good faith and for general information purpose only. We does not make any warranties about the completeness, reliability and accuracy of this information. Any action you take upon the information you find on ansnew.com is strictly at your own risk. We will not be liable for any losses and/or damages in connection with the use of our website. Please read our complete disclaimer. And we do not hold any copyright over the article multimedia materials. All credit goes to the respective owner/creator of the pictures, audios and videos. We also accept no liability for any links to other URLs which appear on our website. If you are a copyright owner or an agent thereof, and you believe that any material available on our services infringes your copyrights, then you may submit a written copyright infringement notification using the contact details

(0) Comments on "Build a Mini GPT Language Model from Scratch"

* Most comments will be posted if that are on-topic and not abusive