LLM Engineering From Scratch · Part 1
Why I Am Learning LLMs From Scratch (And Publishing All of It)
A staff engineer's 12-month plan to understand language models from the agent layer down to the kernels, written in public, grounded in real production data.
I have shipped a lot of software that uses language models. I built an agent platform that has run over a thousand real coding sessions. I can make these things work.
I could not tell you why.
Not really. Not in the way I can tell you why a database query is slow, or why a Kubernetes pod is stuck in CrashLoopBackOff. For those I have a mental model that goes all the way down. For language models I had a bag of habits: this prompt works better, that model is smarter, use a bigger context. That is not understanding. That is superstition with good uptime.
So I am fixing it. Over the next twelve months I am going to learn how these systems actually work, from the token all the way down to the GPU kernel, and I am going to publish every step here.
The specific problem with knowing "enough"
Here is a real thing that happened. I looked at my agent platform's bill: $10,239 across 1,098 sessions. I looked at the token counts and found that 97.5% of my input tokens were cache reads, which sounds excellent, and that I had also written 67 million tokens into the cache, which costs 1.25 times the normal rate.
Was that good? Was it terrible? I genuinely did not know. I did not know because I did not understand what a prompt cache actually is at the level of the serving system. I was reading a dashboard and guessing.
That is the gap. Not "can I build with this" but "can I reason about this". The second one is what lets you make a decision that is not a coin flip.
What I am actually going to do
The plan has three parts running at the same time.
Learn the theory properly. Not paper-skimming. Implement-it-or-you-did-not-read-it. Five areas:
- How a model is built. Tokenizers, attention, the transformer block, scaling laws.
- How it is trained at scale. Memory arithmetic, parallelism, kernels, floating point numerics.
- How it is taught to behave. Fine-tuning, RLHF, DPO, GRPO, verifiable rewards.
- How it is served fast. KV cache, batching, paged attention, speculative decoding.
- How you know it works. Evaluation statistics, agent benchmarks, interpretability, calibration.
Build the exercises. A byte-pair encoding tokenizer from nothing. A transformer with no library helpers. Train GPT-2 small and log every dollar. Write a fused kernel in Triton. Implement DPO from the loss function up. Serve a model and chart the latency curves. Train a sparse autoencoder and look inside a model's head.
Apply it to something real. This is the part I think matters most, and it is the part almost no learning plan has. I have a running agent platform producing real telemetry: 1,098 sessions, 152,359 conversation events, 906 human approval decisions, all with cost and outcome labels attached. That is not a toy dataset. Every concept I learn gets pointed at that data to see if it survives contact with reality.
The five things I am going to build
These are the applied projects, and they double as the reason to keep going when week nine gets boring.
A real benchmark for my own agent. Right now I change a prompt, run it once, and decide it is better. That is not evidence, it is vibes. I want confidence intervals and paired comparisons on a private task set drawn from my own history.
A prompt cache audit. Find every source of cache invalidation in my prompt assembly and fix it. This is the one with the fastest payback and it is the subject of post 14.
A failure predictor. 115 of my 1,098 sessions failed. I have the complete event trail leading to each one. Can a model predict a doomed session from its first few turns, so I can stop burning money on it? At $0.32 per turn, with some sessions running past a hundred turns, this pays for itself if it works at all.
My task board as a reinforcement learning environment. This one surprised me when I noticed it. My agent board already has objectively verifiable outcomes: tests pass, the build succeeds, the pull request merges. That is the definition of a reinforcement learning environment with verifiable rewards. I have been running one for a year without calling it that.
Teaching an agent when to ask for help. This is the capstone and the one I care about most. My platform asks humans to approve agent actions, and today that decision is a static rule: this tool is allowed, that one is not. It should be an uncertainty decision. I have 906 recorded human judgments about when an agent should not be trusted alone. That is a calibration dataset. With conformal prediction you can turn it into a real guarantee, something like: interrupt the human at most 15% of the time, while keeping the chance of an unrecoverable mistake under 2%. No assumptions about the distribution required. That would be a genuinely better product and a genuinely new result.
Why publish it
Three honest reasons.
Writing is the test. You can finish a lecture feeling like you understood it. You cannot finish a blog post feeling that way, because the gaps show up as sentences you cannot write. Every post in this series is a checkpoint I have to actually pass.
Public work compounds. I am aiming at a research or engineering role at one of the frontier AI labs. Nobody there will care about a certificate. They might care about a reproduction that works, a benchmark harness that is statistically honest, or a result nobody has published because nobody else had the dataset.
The explainers I wanted did not exist. Most writing about language models is either a press release or a paper. There is very little in between for the working engineer who wants the actual mechanism, in plain words, with real numbers. That is the gap I am writing into.
What this series will and will not be
It will be: plain English, short sentences, and a real number attached to every claim. Sources at the end of every post, all of them checked. Negative results published alongside the wins, because the failures are usually more informative.
It will not be: takes about AGI, model release commentary, or prompt engineering tips. Those are well covered by people who enjoy them more than I do.
The next post opens the books. Real production telemetry from a real agent platform: what a thousand agent sessions cost, where the money actually went, and the one number that turned out to be the most interesting thing in the whole dataset.
Sources
- Stanford CS336, Language Modeling from Scratch - the course this series leans on most heavily. Lectures and assignments are public. https://github.com/stanford-cs336 and https://online.stanford.edu/courses/cs336-language-modeling-scratch
- Andrej Karpathy, nanoGPT and minbpe - the reference implementations for building a transformer and a tokenizer from nothing. https://github.com/karpathy/nanoGPT and https://github.com/karpathy/minbpe
- Hugging Face, The Ultra-Scale Playbook - training large models across many GPUs. https://huggingface.co/spaces/nanotron/ultrascale-playbook
- Jacob Austin et al., How to Scale Your Model - the systems arithmetic behind training and serving. https://jax-ml.github.io/scaling-book/
- Stas Bekman, Machine Learning Engineering Open Book - the practical failure modes nobody writes papers about. https://github.com/stas00/ml-engineering
- Anastasios N. Angelopoulos and Stephen Bates, A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification, arXiv:2107.07511. The mathematical basis for the capstone project. https://arxiv.org/abs/2107.07511
- Richard Sutton, The Bitter Lesson. http://www.incompleteideas.net/IncIdeas/BitterLesson.html