Inside the Transformer Training Loop

Neo: I know kung fu.
Morpheus: Show me.
(Neo steps forward, his body moving in a blur of perfect, calculated strikes. He doesn't have to think about it. The knowledge is crystallized in his muscles. He just executes.)
Morpheus: (Lowering his hands, a slight smile) Impressive. But do you know how you know it, Neo? You didn't spend ten years in a dojo. You didn't bleed on the mat. You simply... downloaded it. Every block, every strike, is just an optimized path frozen in memory.
Welcome to the construct.
Act I: The Illusion of Memory
Neo: Morpheus, look. I downloaded a 7 Billion parameter model that can rewrite Shakespeare and code an entire operating system. I've been tearing apart its directory on my hard drive.
But there's nothing in here. No database. No Wikipedia dump. No web scraper. No folder called "knowledge."
Just one single file. model_parameters.ckpt. 1.24 gigabytes of pure, chaotic decimal points.
W1: 0.1241, -0.9032, 0.5517... W2: -0.4428, 0.7763... — just noise.
How can it know so many things... if the data isn't actually there?
Morpheus: You are still looking at the machine through human eyes, Neo.
You assumed that to know a world, the system must carry a copy of that world inside it. You were looking for filing cabinets inside a mind.
Reset the simulator.
That file you are staring at doesn't contain the data. The file IS the data.
The text you think it is reading was burned to ash a long time ago. What you downloaded is not a library, Neo. It is the fossilized echoes of a library. A crystallized memory of how words once flowed.
You are not looking at decimals. You are looking at frozen knowledge.
" References:
Radford, A., et al. (2019). Language Models are Unsupervised Multitask Learners. (The GPT-2 foundational paper proving next-token prediction inherently builds world knowledge and reasoning without external database lookups).
Kaplan, J., et al. (2020). Scaling Laws for Neural Language Models. (Demonstrating how cross-entropy loss scales predictably with compute, dataset size, and parameter count). "
Act II: The First Flight (The Forward Pass)
Neo: Okay, so there is no database. But how does the text enter the engine in the first place? If I feed it an entire book to learn from, does it read it page by page?
Morpheus: No. It devours it in massive blocks, Neo. We slice the book into thousands of numerical fragments called tokens, and we feed them into the network in massive arrays—chunks of 4,096 tokens at a time. The model doesn't just sit there waiting for the end of a sentence. It plays a game. It calculates a blind guess for what the next word should be for every single word in that block, all at the same time.
Neo: Wait, it guesses the next word for the entire block simultaneously? Even if it's the very first day of training?
Morpheus: Especially then. On day one, its internal dials are completely random. It looks at the first 4,096 tokens of your book, and it simultaneously blabs out 4,096 words of absolute, unbaked noise. That massive, parallel surge of text entering the machine and spitting out thousands of blind guesses is what we call the Forward Pass. It will repeat this surge, block by batch, until it consumes the entire book. One full pass through the entire book is one Epoch.
" References:
Vaswani, A., et al. (2017). Attention Is All You Need. (Introducing the self-attention mechanism and the Causal Attention Mask that prevents tokens from peeking at future positions during parallel training).
Sennrich, R., et al. (2015). Neural Machine Translation of Rare Words with Subword Units. (The foundational work on Byte-Pair Encoding (BPE) text tokenization). "
Act III: The Interrogation (The Loss Function)
Neo: So it sits there on day one, blindly throwing thousands of random words into the void. It’s a total mess. How does the system recognize the failure? Who stands over the machine and tells it that it got the answer wrong?
Morpheus: The simulation relies on a cold, mathematical evaluator, Neo. We call it the Loss Function. The network doesn’t just get a simple pass or fail grade. The Loss Function holds the actual target text—the original, uncorrupted book—and places it directly alongside the model's chaotic predictions. It doesn't just look at whether the guess was right or wrong; it measures the model's confidence.
Neo: Its confidence?
Morpheus: If the target text says the next word is "control", and the model assigned a 99% probability to "Bicycle", the Loss Function hits it with an exponential penalty. It transforms the model's profound confusion into a single, concrete number. The higher the number, the worse the catastrophe. The entire goal of the training loop is to force that single number down as close to zero as possible.
" References:
Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. (The foundational textbook detailing the mathematics of Cross-Entropy minimization and maximizing log-likelihood).
Shannon, C. E. (1948). A Mathematical Theory of Communication. (The original framework establishing information entropy, the mathematical bedrock underneath neural network error evaluation). "
Act IV: Rewinding the Simulation (Backpropagation)
Neo: Okay, so the model has a penalty score. It knows it messed up. But how does that score actually fix the machine? How do we change the numbers inside that massive file?
Morpheus: We trace the glitch back to its source, Neo. We take that loss score and send a mathematical shockwave backward through the entire architecture. The wave travels down the residual highways, in reverse through the MLP memory blocks and the Attention matrices.
It doesn't fix anything yet. It only asks one question at every single weight dial: "How much of this catastrophe was your fault?"
It calculates exactly how much each individual dial contributed to the bad guess. That error map — how guilty each weight is — is called the Gradient. ∂L/∂W.
Neo: So now we have a map of blame. How do we fix it?
Morpheus: Right behind the shockwave comes the Optimizer — the mechanical engineer of the system. It takes that map of guilt and subtly nudges every weight dial against the gradient. W ← W - η·∂L/∂W
It alters the mathematical gravity inside. By shifting those dials, it warps the vector space — pulling coordinates of similar concepts closer together, pushing unrelated ones apart. So the next time it sees that sentence, the gravitational pull bends away from "Bicycle" and straight toward "control."
" References:
Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. (The landmark paper establishing the calculus chain-rule architecture for modern neural net training).
Kingma, D. P., & Ba, J. (2014). Adam: A Method for Stochastic Optimization. (Introducing the Adam optimization framework used to smoothly step weight updates down the gradient descent slope). "
Act V: The Ultimate "Aha!" Moment (The Discard)
Neo: Hold on... let me follow this to the end. The text comes in, the model guesses, the evaluator grades it, the shockwave travels backward, and the internal weight dials shift to warp the space. What happens to the Wikipedia page or the book after that dial moves?
Morpheus: It is permanently deleted, Neo. It is purged from the system memory and thrown in the trash before the next block of text even loads.
Neo: Wait. It's completely gone? But if the text is thrown away, how does the model still remember the facts inside the book?!
Morpheus: Because the text was never meant to be stored. Think of a massive river flowing through a wild forest for the very first time. The water rushes in blindly, crashing against the dirt and stone, searching for a way out. But as millions of gallons of water keep surging through the exact same spot, the force of the water erodes the earth—carving out a deep, permanent canyon. Once the canyon is formed, that original water is gone, Neo. It flowed out to the ocean days ago. But the shape of the canyon remains. The model doesn’t contain the text. The model is the canyon that the text left behind.
Neo: So when I download an "Open Weights" model... I'm downloading the shape of the canyon.
Morpheus: Exactly. The billion-dollar training run is over, and the master canyon is locked in the vault. The dials are frozen. FROZEN: TRUE. No gradients. No backprop. Read-only.. But because the weights are "open," you hold the keys. You can keep 99% of the canyon completely frozen, unlock just a tiny fraction of the active weight dials at the banks, and divert a little bit of your own water through it to teach it new paths. You are no longer training the canyon, Neo. You are navigating it
" References:
Hu, E. J., et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. (The foundational research for parameter-efficient fine-tuning, explaining how to map weight matrices updates while freezing the primary base layers).
Touvron, H., et al. (2023). LLaMA: Open and Efficient Foundation Language Models. (Demonstrating optimization frameworks for public open-weights ecosystem architectures). "
Summary: The Bottom Line
Training an LLM isn't about saving files to a digital hard drive. It is a continuous, beautiful cycle: the model guesses (Forward Pass), gets graded (Loss), traces its errors backward (Backpropagation), and updates its internal structure (Optimizer)—raw text vanishes so that only the mathematical pathways of human thought remain.
Next time you watch a progress bar crawl during a training run, don't just see a loading screen. See Neo, watching the machine learn to see the code.
Keep coding, redpills.




