Understanding GPT-2 Tokenization
- gpt-2
- tokenization
- llm
Series contents
Part 2 · 3 published- Part 1GPT-2, One Token at a Time
- Part 2Understanding GPT-2 TokenizationCurrent
- Part 3Understanding GPT-2 Token Embeddings
In the first GPT-2 article, I treated the tokenizer as a black box. Text went in and token IDs came out.
I expected the space before a word to become its own token. Instead, "hello" and " hello" each became one token with a different ID.
I also wanted to see what GPT-2 does when a misspelled or unseen word is missing from its vocabulary.
A Space Is Not Always a Separate Token
I used the GPT-2 tokenizer from Hugging Face Transformers:
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("gpt2")
for text in ["hello", " hello"]:
token_ids = tokenizer.encode(text)
tokens = tokenizer.convert_ids_to_tokens(token_ids)
print(repr(text), token_ids, tokens)
'hello' [31373] ['hello']
' hello' [23748] ['Ġhello']
Both inputs remained one token. Their IDs changed from 31373 to 23748. Decoding 23748 returned " hello" with the space intact.
The Ġ in Ġhello is not a character from the input. GPT-2's tokenizer uses a reversible mapping that gives every byte a visible Unicode representation. In that internal representation, Ġ stands for the byte value 32, which is an ASCII space.
"hello" → [104, 101, 108, 108, 111]
" hello" → [32, 104, 101, 108, 108, 111]
^^
space
GPT-2's initial text matching pattern can group one leading space with the letters that follow it. Byte pair encoding then applies learned merge rules to that group. Its vocabulary contains both hello and Ġhello. Each input becomes one token.
Between Text and Token IDs
BPE stands for byte pair encoding. It starts with small symbols and uses a ranked list of learned merges. When adjacent pieces have a merge rule, the tokenizer joins them. Frequent byte sequences can eventually become one vocabulary token while less common sequences remain split across several tokens.
This process works with byte sequences, not meanings. The tokenizer does not know that hello is a greeting. It knows that the bytes behind hello and hello follow merge paths that end at two different vocabulary entries.
The merge list is fixed after tokenizer training. Encoding the same text with the same GPT-2 tokenizer is deterministic.
A Word Outside the Vocabulary
For my second question, I used my name without a space:
text = "ganesshkumar"
print(text in tokenizer.get_vocab())
print(tokenizer.encode(text))
print(tokenizer.convert_ids_to_tokens(tokenizer.encode(text)))
False
[1030, 408, 71, 74, 44844]
['gan', 'ess', 'h', 'k', 'umar']
Decoding all five IDs returned ganesshkumar exactly.
GPT-2 does not need one vocabulary entry for every name, misspelling or identifier. If the complete text is absent, BPE leaves it as smaller pieces. GPT-2's BPE has a base representation for all 256 byte values. Valid UTF-8 text can be represented even when the whole word has never become one token.
The same process handles source code identifiers, new words and names. Learned merge rules determine the split.
The base GPT-2 tokenizer has 50,257 vocabulary entries. ID 31373 points to "hello", ID 23748 points to " hello" and ID 1030 points to "gan". Their numeric size says nothing about how the tokens are related.
This is the same 50,257 dimension I saw in the logits earlier. It contains one score for each vocabulary entry.
Tokenization also determines T, the sequence length. "hello" occupies one position. "ganesshkumar" occupies five positions in this checked run. GPT-2 receives those five IDs before any attention or other model computation begins.
What I Took Away
I started with a model of separate words and spaces: a word would become a token and its leading space would be another token. GPT-2's tokenizer does not keep that boundary.
A token can contain a leading space. A word can be one token or several reusable pieces. If the whole word is missing, the tokenizer can continue down to smaller byte pieces and still reconstruct the original text.
The output is still a list of arbitrary integer IDs. The next question is how GPT-2 turns those IDs into vectors that the model can actually work with.