Inside LLMs: How Embeddings, Attention, and Feedforward Networks Work
- Mark Chomiczewski
- 7 October 2026
- 0 Comments
You’ve probably wondered how a chatbot can write poetry in the style of Shakespeare or debug your Python code without blinking. It feels like magic, but it’s actually math-specifically, a specific arrangement of neural network components called the Transformer architecture. Since Google Brain introduced this concept in 2017 with the paper "Attention is All You Need," it has become the backbone of every major AI model, from GPT-4 to Llama 3. But what’s actually happening inside that black box?
At its core, a Large Language Model (LLM) isn’t one giant brain; it’s a pipeline of three distinct jobs: turning words into numbers (Embeddings), figuring out which words matter most to each other (Attention), and processing those relationships to create new meaning (Feedforward Networks). If you understand these three pieces, you understand how modern AI thinks. Let’s break them down without the heavy academic jargon.
The Foundation: Turning Words into Vectors
Computers don’t understand words; they understand numbers. The first step in any LLM is converting text into a format the machine can crunch. This starts with tokenization, where sentences are chopped into small units called tokens. For example, the word "unbelievable" might be split into "un", "believ", and "able." Once we have tokens, we need to give them meaning.
This is where embeddings come in. An embedding is essentially a list of numbers-a vector-that represents a word’s meaning. In models like BERT-base, these vectors are 768 dimensions long. In GPT-3, they’re massive at 12,288 dimensions. Why so many? Because each dimension captures a different nuance of meaning. One dimension might track whether a word is positive or negative, another might track if it’s about animals, and another might handle grammatical tense.
| Model | Embedding Dimension | Vocabulary Size | Primary Use Case |
|---|---|---|---|
| BERT-base | 768 | 30,522 | Understanding & Classification |
| GPT-3 | 12,288 | 50,257 | Text Generation |
| Llama 2 | 4,096 | 32,000 | Open Source Research |
Here’s the cool part: semantically similar words end up close together in this high-dimensional space. Researchers have famously demonstrated that if you take the vector for "King," subtract the vector for "Man," and add the vector for "Woman," you get a result very close to the vector for "Queen." This geometric relationship allows the model to grasp analogies and semantic shifts without being explicitly programmed with grammar rules.
But wait-if all words are just vectors, how does the model know the order? "Dog bites man" means something totally different than "Man bites dog." Transformers process all tokens in parallel, not sequentially like older Recurrent Neural Networks (RNNs). To solve this, we add positional embeddings. These are special vectors added to the word embeddings that encode position information. Without them, the model would see a bag of words rather than a sentence.
The Heartbeat: How Attention Works
If embeddings provide the raw data, attention provides the context. This is the component that made Transformers famous. Before 2017, models had to look at words one by one, forgetting earlier parts of a long sentence as they moved forward. Attention lets the model look at the entire sentence at once and decide which parts are relevant to each other.
Think of it like reading a mystery novel. When you reach the final chapter, you might flip back to page 10 to remember a clue. Attention mechanisms do this automatically for every word. Technically, this works using three vectors for every token: Query (Q), Key (K), and Value (V).
- Query: What am I looking for?
- Key: What do I contain?
- Value: What information do I provide if matched?
The model calculates an attention score by comparing the Query of one word with the Keys of all other words. A high score means the two words are strongly related. For instance, in the sentence "The animal didn't cross the street because it was too tired," the word "it" needs to figure out who it refers to. The attention mechanism gives a high score between "it" and "animal," effectively linking them.
To make this even more powerful, modern LLMs use Multi-Head Attention. Instead of one single way of looking at relationships, the model splits the work into multiple "heads." GPT-3 uses up to 96 heads. One head might focus on grammatical structure, another on semantic similarity, and another on long-range dependencies. This parallel processing allows the model to capture complex linguistic patterns simultaneously.
However, attention comes at a cost. The computational complexity grows quadratically with sequence length. Processing a 32,768-token sequence requires over a billion calculations just for attention. This is why running large models locally on your laptop can be tough-you often need specialized hardware like NVIDIA A100 GPUs with 80GB of VRAM to handle the memory load efficiently.
The Processor: Feedforward Networks
Once the attention mechanism has gathered context, the information needs to be processed and transformed. This is the job of the Feedforward Network (FFN), also known as a Multilayer Perceptron (MLP). While attention handles the relationships between words, the FFN handles the internal computation of each word’s representation independently.
Imagine attention as a library where you find all the books relevant to your topic. The FFN is then the desk where you sit down, read those books, synthesize the information, and write your notes. Structurally, an FFN consists of two linear transformations with a non-linear activation function in between. Typically, the first layer expands the dimensionality by four times (e.g., from 768 to 3,072 units in BERT), applies a GELU activation function, and then projects it back down.
Why expand and then shrink? This expansion creates a bottleneck that forces the model to learn compressed, efficient representations of knowledge. Recent research suggests that feedforward networks act as the model’s "memory." They store factual knowledge learned during training. When you ask an LLM "Who wrote Hamlet?", the attention mechanism finds the words "Hamlet" and "wrote," but the FFN retrieves the association with "Shakespeare" from its stored weights.
Optimizations here are critical. Because FFNs account for 30-40% of inference time, researchers have developed techniques like Mixture of Experts (MoE). MoE splits the FFN into multiple smaller networks and only activates a few of them for each token, saving massive amounts of compute power without sacrificing performance.
Putting It All Together: The Transformer Block
These three components don’t work in isolation. They are stacked in repeating units called Transformer Blocks. A typical block follows this flow:
- Input: Token embeddings + Positional encodings.
- Self-Attention: Contextualizes each token based on others.
- Add & Norm: Adds the original input back (residual connection) and normalizes the output to stabilize training.
- Feedforward Network: Processes the contextualized information.
- Add & Norm: Another residual connection and normalization.
- Output: Passed to the next block.
A model like GPT-3 stacks 96 of these blocks deep. Each layer refines the understanding slightly more. Early layers might focus on syntax and local grammar, while deeper layers handle abstract reasoning and long-term narrative consistency. This depth is what allows LLMs to perform tasks they weren’t explicitly trained for, a phenomenon known as emergent capabilities.
Real-World Implications and Pitfalls
Understanding these components helps explain why LLMs behave the way they do. For developers, the biggest pain point is often implementing custom attention masks. If you’re fine-tuning a model for a specific domain, you might need to restrict attention so the model doesn’t "peek" at future words during generation. Getting this wrong leads to poor coherence.
Another common issue is numerical instability in embeddings. If the initialization isn’t right, or if the learning rate is too high, the vector spaces can collapse, causing the model to forget distinctions between similar words. Techniques like Rotary Position Embedding (RoPE), used in Llama models, help mitigate some of these issues by encoding relative positions directly into the attention calculation, allowing models to generalize better to longer sequences than they saw during training.
For businesses, the choice between autoregressive models (like GPT) and autoencoding models (like BERT) depends entirely on the task. GPT uses causal attention-it can only look backward-which makes it great for generating text. BERT uses bidirectional attention-it looks both ways-making it superior for classification and search relevance. Knowing which architecture fits your problem saves weeks of development time.
Frequently Asked Questions
Why are embeddings important in LLMs?
Embeddings convert discrete words into continuous vector spaces, allowing the model to understand semantic relationships. By placing similar words close together in high-dimensional space, the model can generalize from known examples to unseen ones, enabling it to grasp concepts like synonyms and analogies without explicit programming.
What is the difference between self-attention and multi-head attention?
Self-attention computes relationships between all tokens in a sequence simultaneously. Multi-head attention runs multiple self-attention operations in parallel, each with different learned parameters. This allows the model to attend to different types of information at once, such as syntactic structure in one head and semantic meaning in another, providing a richer representation.
How do feedforward networks contribute to knowledge storage?
Research suggests that feedforward networks act as key-value memories within LLMs. While attention retrieves context, the FFN processes this context and stores factual associations learned during training. Expanding the dimensionality in the FFN allows the model to hold a vast amount of parametric knowledge, which is retrieved when generating responses.
Why is attention computationally expensive?
Attention mechanisms have quadratic complexity relative to sequence length. Every token must compare itself with every other token in the sequence. As the context window grows (e.g., from 4k to 128k tokens), the number of computations increases exponentially, requiring significant GPU memory and processing power, which is why optimizations like FlashAttention are critical.
Can LLMs exist without transformers?
Yes, but they were less effective. Previous architectures like RNNs and LSTMs processed text sequentially, leading to vanishing gradient problems and difficulty capturing long-range dependencies. The transformer architecture, with its parallelizable attention mechanism, solved these bottlenecks, enabling the scale and performance seen in modern LLMs.