Article series

Notes From a Software Engineer Learning LLMs

I am using GPT-2 as a model I can trace and test while I build a clearer picture of how LLMs work.

Ongoing · 3 articles published

Start with Part 1

Visual notes

A precise system view that will grow only as the related articles are published.

Text becomes embeddings, passes through twelve GPT-2 Small decoder blocks, and uses softmax plus sampling or argmax to select the next token ID.
From text to the next token IDRead the related article

Card 1 of 2

Published articles

The complete published sequence. Unpublished work does not appear here.

  1. Part 1GPT-2, One Token at a TimeI use GPT-2 as a black box to trace one prompt into token IDs and logits before implementing the loop that generates text one token at a time.6 min read
  2. Part 2Understanding GPT-2 TokenizationI expected a space to be its own token. GPT-2 showed me why hello and a space before hello have different token IDs and how it handles a word outside its vocabulary.4 min read
  3. Part 3Understanding GPT-2 Token EmbeddingsI knew an embedding layer turned token IDs into vectors. I wanted to find out whether GPT-2 learned those vectors with the rest of the model or used an embedding created elsewhere.4 min read