Embeddings — LLMs Part 5
The Definition
An embedding is a list of numbers that represents one token’s meaning. That list is also called a vector — an ordered list of numbers, nothing more.
Here meaning has a narrow sense: how a token gets used — the kinds of sentences it appears in — not anything about the real world.
The same token always gets the same vector, no matter where it appears. Special tokens get vectors too — the end-of-sequence token from Part 3 is one of them.
Every token in the vocabulary (Part 3) gets its own vector. Picture the embedding table — one row per token: a token’s position in the vocabulary is its row number, which is exactly its token ID. Turning a token into its vector is a lookup, not a calculation: the network stores one vector per vocabulary entry — these stored vectors are part of the network’s own parameters (Part 2 covered weights and biases; embeddings are the other kind) — and retrieves the right one. Part 4’s tokenizer spells every input from known pieces, with no generic “unknown” token — so every ID the lookup ever receives already has its own row.
A few illustrative rows make this concrete — invented numbers, not an actual model’s real weights. The table itself stores only the vector column; the row number is the token ID, not a separately stored value, and the token spelling isn’t stored in this table at all — that mapping lives in the vocabulary (Part 3/Part 4) instead. “Row” and “Token” are shown here only so the table is readable:
| Row (= token ID) | Token (not stored here — shown for reference) | Vector |
|---|---|---|
| 1 | cat | 0.22, 0.68, … |
| 2 | dog | 0.42, 0.80, … |
| 3 | car | 0.82, 0.22, … |
A real table has one row per vocabulary entry — 50,257 of them for GPT-2 — and hundreds or thousands of numbers per row, not two; this is trimmed down to show the shape:
flowchart LR
T["token"] -->|"position in<br/>vocabulary"| ID["token ID"]
ID -->|"lookup"| V["vector<br/>(embedding)"]
accTitle: Looking up a token's embedding
accDescr: A token maps to its token ID by position in the vocabulary, and that ID is used to look up its vector, the embedding, from a stored table.
Where This Fits in the Pipeline
This lookup isn’t a separate phase — it’s the network’s own first move, every single time. It happens right after tokenization and right before the network’s first layer (Part 2) does anything with the result, for every token: the ones in the original prompt, and every token the model generates and feeds back in during the loop Part 3 described:
flowchart TB
TOK["prompt tokens<br/>(Part 3)"] -->|"whole prompt<br/>at once"| EMB["embedding lookup<br/>(this post)"]
EMB -->|"one vector<br/>per token"| NET["network layers<br/>(Part 2)"]
NET -->|"one final vector"| SCORE["scored against<br/>every row<br/>(this post)"]
SCORE -->|"highest score wins"| ID["winning row:<br/>next token ID"]
ID -->|"one ID<br/>at a time"| EMB
ID -->|"looked up in<br/>vocabulary (Part 3/4)"| TXT["next token<br/>(text)"]
accTitle: Where embedding lookup sits in the pipeline
accDescr: Prompt tokens are looked up all at once, one vector per token; the vectors feed the network's layers; the layers produce one final vector, scored against every row to find the highest-scoring row; that winning row is the next token's ID, fed back into the lookup one ID at a time to continue generation, and separately looked up in the vocabulary to produce the actual output text.
The lookup only cares which token it is, not where in the sentence it appears. The same token gets the same vector whether it’s the first word or the last — word order gets added afterward, in a future post.
One spelling gets one row even when it has two meanings: bank the riverbank and bank the money store start from the same vector. That starting vector doesn’t stay frozen, though. Each layer revises it using the surrounding tokens. A future post covers that mechanism — which is also how the two banks get told apart downstream.
Training, covered in a future post, is what shapes the vectors themselves — it adjusts them the same way it adjusts any other parameter. By the time the model is generating text, the vectors are already fixed, and this step is retrieval, nothing more.
Many models, including GPT-2, use that same table a second time. Scoring every possible next token (Part 1) means comparing the network’s output to every vocabulary entry’s vector — and rather than learning a whole separate set of vectors for that, the model reuses the same embedding table, applied in reverse: rows read out as comparisons instead of looked up by ID. Each row gets one score for how well it matches the output — computed by multiplying the output’s coordinates with the row’s matching coordinates and adding the results up — and the highest score wins. Reusing one table for both jobs is called weight tying; it saves a whole second table’s worth of parameters.
That scoring isn’t a loop over rows, one at a time — it’s a single matrix multiplication. The embedding table is already a matrix, one row per vocabulary entry; multiplying that matrix by the output vector produces exactly one number per row in one operation, and each of those numbers is precisely the row-by-row multiply-and-add score described above. A tiny, illustrative version, reusing the 3-row table above and a made-up 2-number output vector [0.4, 0.8]:
| Row | Vector | Score (vector × output, added up) |
|---|---|---|
cat | 0.22, 0.68 | 0.22×0.4 + 0.68×0.8 = 0.63 |
dog | 0.42, 0.80 | 0.42×0.4 + 0.80×0.8 = 0.81 |
car | 0.82, 0.22 | 0.82×0.4 + 0.22×0.8 = 0.50 |
All three scores come out of that one multiplication; dog wins because its vector was the closest match to the made-up output vector. A real model runs this same arithmetic across all 50,257 rows and 768 numbers per row at once — the scale changes, the operation doesn’t.
The winning row’s number is the model’s actual output here, not text — row 2 scoring highest doesn’t hand the model the word “dog” directly, it hands it “row 2.” Since the embedding table never stored the token spelling in the first place, turning that row number into the text dog is a separate step done afterward, by looking the ID up in the tokenizer’s vocabulary (Part 3/Part 4) — the same vocabulary as the embedding table’s rows, used in the opposite direction from how tokenizing an input works: text in, ID out there; ID in, text out here.
How It Works
Each vector has many numbers in it, and real models vary a lot: GPT-2 uses 768, Llama 2’s 7-billion-parameter model uses 4,096, and GPT-3’s largest model uses 12,288. Each number is one coordinate in a space with that many dimensions — called a vector space. Two tokens with similar meaning end up with similar vectors: close together in that space. Two unrelated tokens end up far apart. Closeness isn’t something the model sees — it’s something it computes, with the same multiply-and-add as scoring: multiply the two vectors’ matching coordinates and add the results up. A big total means close together; a small one means far apart.
Multiply vocabulary size by dimension count and the scale becomes concrete: GPT-2’s vocabulary (Part 4) has 50,257 entries, each needing its own 768-number vector — about 38.6 million numbers in this one table alone, before the rest of the network’s own layers (Part 2) add any parameters of their own.
Simplified to two coordinates so it can be drawn, the idea looks like this. Read only the distances between the dots:
%%{init: {"quadrantChart": {"pointTextFontSize": 20}}}%%
quadrantChart
title Relative distances only
x-axis " " --> " "
y-axis " " --> " "
cat: [0.22, 0.68]
dog: [0.42, 0.8]
car: [0.82, 0.22]
accTitle: cat and dog close together, car far apart
accDescr: A simplified 2-coordinate plot: cat and dog as nearby points, car as a distant point. Only the relative distances between points carry meaning.
cat and dog land near each other; car lands far from both. A real embedding uses hundreds or thousands of coordinates, not two, which is what lets it capture far more relationships than “close” or “far” in a simple picture can show.
One thing the simplified picture can’t show honestly: unlike its x and y axes, a real embedding’s individual coordinates don’t each stand for a nameable property like size or color. No single number says “how much of a pet this is.” Meaning lives in the whole vector together — spread across every coordinate at once, not pinned to any one of them.
Where the Closeness Comes From
No one, and nothing, explicitly decides that cat and dog should be close. The vectors start as random numbers, the same way every other parameter in the network does (Part 2) — before training, cat and car are no closer together than cat and dog are. Training, covered in a future post, is the only thing that ever touches these numbers again, and closeness isn’t something it’s ever directly told to produce — it’s a side effect of the specific task those vectors get trained on.
The classic approach, covered in full in a future post, trains each token’s vector to help predict the words that tend to appear near it in real text. cat and dog show up in a lot of the same kinds of sentences — “I fed my __,” “the __ ran across the yard” — so the easiest way for training to get better at predicting those surrounding words for both tokens is to give them similar vectors, so the same downstream computation works for either one. car never shows up in those same contexts, so nothing pulls its vector anywhere near them.
This is sometimes called the distributional hypothesis: words used in similar contexts tend to mean similar things. Training doesn’t know anything about meaning directly — it only ever reduces prediction error on whatever task it’s given — but reducing prediction error on a context-prediction task happens to produce vectors that line up with meaning as a side effect.
Why Embeddings Matter
Once meaning is a list of numbers, “how similar are these two tokens” becomes a plain math question — the distance between two vectors — instead of something a program has to be told by hand.
This pays off directly in generation. Raw token IDs (Part 3) carry no relationship information at all — if cat happened to be token 4187 and dog token 9341, those numbers alone would say nothing about each other. Embeddings fix that: whatever the network has learned about handling cat in a given context transfers to dog almost for free, since their nearby vectors mean the same downstream computation treats them similarly. Everything downstream in the network, including attention, covered in a future post, operates on these vectors, not on the original text — so that transfer happens automatically, every layer, without the model needing to have seen every word combination during training to produce coherent output. These same vectors also get reused outside the model entirely — for finding documents by meaning, covered in a future post.
Next: Paper: word2vec — the paper that showed meaning could be learned this way. (Coming soon.)
This post uses “training” and “attention” loosely — covered in future posts,.