Article series
Notes From a Software Engineer Learning LLMs
I am using GPT-2 as a model I can trace and test while I build a clearer picture of how LLMs work.
Ongoing · 3 articles published
Start with Part 1Visual notes
A precise system view that will grow only as the related articles are published.

Card 1 of 2
Published articles
The complete published sequence. Unpublished work does not appear here.
- Part 1GPT-2, One Token at a TimeI use GPT-2 as a black box to trace one prompt into token IDs and logits before implementing the loop that generates text one token at a time.6 min read
- Part 2Understanding GPT-2 TokenizationI expected a space to be its own token. GPT-2 showed me why hello and a space before hello have different token IDs and how it handles a word outside its vocabulary.4 min read
- Part 3Understanding GPT-2 Token EmbeddingsI knew an embedding layer turned token IDs into vectors. I wanted to find out whether GPT-2 learned those vectors with the rest of the model or used an embedding created elsewhere.4 min read