artificial-intelligence

Starting With GPT-2: One Token at a Time

By Ganessh Kumar6 min read
  • gpt-2
  • llm
  • language-models

I wanted to understand large language models by getting my hands dirty, not by only reading about them. While looking for a model I could inspect end to end, GPT-2 felt like the right starting point: simple enough to run and study, but complete enough to observe the next-token interface used during generation.

GPT-2 also gives me a practical route into the Transformer ideas introduced by Attention Is All You Need. I can first make its one-token-at-a-time generation loop concrete, then open its attention, embeddings and other internal parts before comparing them with newer models.

Before starting, I already had the broad mental model. Give the model some tokens, ask it for the next token, append that token and ask again. What remained unclear was hidden inside the phrase "some deep learning."

In this article, I am keeping that deep learning machinery as a black box. I want to trace one prompt through GPT-2, inspect its next-token scores and make the input and output contract concrete.

The prompt becomes token IDs, GPT-2 returns logits and one token is chosen.
The black-box contract for this article: token IDs go in, vocabulary scores come out and one token is selected.

Start With a Prompt

I will use this prompt throughout the article:

The meaning of life is

Here is the smallest useful GPT-2 setup:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("gpt2")
model = AutoModelForCausalLM.from_pretrained("gpt2")
model.eval()

prompt = "The meaning of life is"
input_ids = tokenizer(prompt, return_tensors="pt")["input_ids"]

with torch.no_grad():
    logits = model(input_ids).logits

The two from_pretrained calls hide some useful setup. If the files are not already cached, the Transformers library downloads the GPT-2 tokenizer files, configuration and pretrained weights from the Hugging Face Hub. The library contains the Python implementation of the GPT-2 architecture. The downloaded weights contain the values GPT-2 learned during training.

AutoModelForCausalLM reads the configuration, creates that architecture and loads the weights into it. It does not train GPT-2 on my machine.

The tokenizer turns the prompt into five token IDs:

[464, 3616, 286, 1204, 318]

Decoded separately, those IDs represent:

"The"
" meaning"
" of"
" life"
" is"

The spaces are part of the decoded token pieces. I am leaving the tokenization details for the next article. Here, GPT-2 receives these integers rather than the Python string.

The input_ids tensor has shape [1, 5]: one prompt with five token positions.

Using BB for batch size and TT for sequence length:

input_ids∈NB×T\texttt{input\_ids} \in \mathbb{N}^{B \times T}

For this prompt, B=1B = 1 and T=5T = 5.

Look at the Logits

The model returns a logits tensor with this shape:

[1, 5, 50257]

Using Z\mathbf{Z} for the logits tensor and VV for vocabulary size, the general form is:

Z=GPT2⁡(input_ids),Z∈RB×T×V\mathbf{Z} = \operatorname{GPT2}(\texttt{input\_ids}), \qquad \mathbf{Z} \in \mathbb{R}^{B \times T \times V}

The base GPT-2 configuration uses V=50,257V = 50{,}257. For this run, the shape means one prompt, five token positions and 50,257 vocabulary scores at each position.

I had expected one set of next-token scores. Instead, GPT-2 returned one set at every input position.

Each score answers a narrow question:

How suitable is this vocabulary token as the token that follows this position?

These scores are called logits. They are not probabilities yet and they do not need to add up to 1.

Why Take the Last Position?

When generating text, we want the prediction made after GPT-2 has seen the complete prompt. That is the final position:

next_token_logits = logits[0, -1, :]

Its shape is:

[50,257]

In this slice:

  • 0: the first prompt in the batch.
  • -1: the final token position.
  • :: every vocabulary score.

In mathematical notation, that slice is one vector in RV\mathbb{R}^{V}:

znext=Z0,−1,:\mathbf{z}_{\text{next}} = \mathbf{Z}_{0,-1,:}

For the five-token prompt, GPT-2 gives us five different prediction points:

PositionPrefix available to GPT-2What its logits predict
0"The"The token after "The"
1"The meaning"The token after "The meaning"
2"The meaning of"The token after "The meaning of"
3"The meaning of life"The token after "The meaning of life"
4 or -1"The meaning of life is"The new token after the full prompt

The first four positions predict tokens inside a sequence we already supplied. They matter during training, when GPT-2 learns from every position. During generation, position -1 is the one that continues beyond the supplied prompt.

Turn the Scores Into Probabilities

Softmax turns the logits into a probability distribution:

probabilities = torch.softmax(next_token_logits, dim=-1)

For vocabulary entry ii, the calculation is:

pi=ezi∑j=1Vezjp_i = \frac{e^{z_i}}{\sum_{j=1}^{V} e^{z_j}}

Here, ziz_i is the logit for token ii, while the denominator includes all VV vocabulary logits.

The resulting 50,257 probabilities add up to approximately 1.

For this prompt, one checked run produced these leading candidates:

" not"   0.108805
" to"    0.082914
" the"   0.051284
" that"  0.048244
" a"     0.039578

Those numbers are from one checked run. They can change with the model revision and software environment.

Pick the Next Token

The simplest decoding rule is greedy decoding:

next_token_id = torch.argmax(
    next_token_logits,
    dim=-1,
)

This chooses the highest-scoring token. I do not need softmax before argmax because softmax preserves the ordering of the logits.

For the sample above, the selected token is ID 407, which decodes to " not".

Appending it changes the input from:

The meaning of life is

to:

The meaning of life is not

Repeat

To generate more than one token, we repeat the same operation in a loop:

generated_ids = input_ids

for _ in range(30):
    with torch.no_grad():
        logits = model(generated_ids).logits

    next_token_logits = logits[:, -1, :]
    next_token_id = torch.argmax(
        next_token_logits,
        dim=-1,
        keepdim=True,
    )

    generated_ids = torch.cat(
        [generated_ids, next_token_id],
        dim=1,
    )
GPT-2 runs, one token is chosen and appended and the loop starts again.
Generation repeats the same contract after appending each selected token.

This code reruns GPT-2 over the whole sequence on every pass. Production inference can reuse earlier work, but that optimization does not change the next-token contract.

Greedy decoding can become repetitive because it always takes the locally highest-scoring option. Sampling and temperature would change the choice, not the loop.

The Black Box Is Still Closed

I stopped at these two shape transformations:

token IDs⏟B×T→GPT-2logits⏟B×T×V\underbrace{\text{token IDs}}_{B \times T} \xrightarrow{\text{GPT-2}} \underbrace{\text{logits}}_{B \times T \times V}

For generation, I kept the scores at the final position and applied a decoding rule:

Z:,−1,:⏟B×V→decoding rulenext token IDs⏟B×1\underbrace{\mathbf{Z}_{:,-1,:}}_{B \times V} \xrightarrow{\text{decoding rule}} \underbrace{\text{next token IDs}}_{B \times 1}

What I Took Away

My starting model was right but incomplete. GPT-2 does not return only one set of scores for the whole prompt. It returns scores at every input position. The generation loop takes the last set, picks one token and runs again.

The next article in the series will open the first part of that black box: how the tokenizer turns text into the token IDs GPT-2 receives.

Sources

Keep reading

I write about full-stack development, developer tooling and the problems I run into building things. Here is everything else on the blog.