<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Under Development Blog</title>
  <subtitle>Engineering notes from Under Development LLC on language models, agent systems and the infrastructure underneath them.</subtitle>
  <link href="https://undev.ai/blog/feed.xml" rel="self"/>
  <link href="https://undev.ai/blog/index.html"/>
  <id>https://undev.ai/blog/index.html</id>
  <updated>2026-08-26T06:54:28Z</updated>
  <author><name>Under Development LLC</name></author>
  <entry>
    <title>Why I Am Learning LLMs From Scratch (And Publishing All of It)</title>
    <link href="https://undev.ai/blog/why-im-learning-llms-from-scratch.html"/>
    <id>https://undev.ai/blog/why-im-learning-llms-from-scratch.html</id>
    <updated>2026-08-25T00:00:00Z</updated>
    <summary>A staff engineer&#x27;s 12-month plan to understand language models from the agent layer down to the kernels, written in public, grounded in real production data.</summary>
    <content type="html">&lt;p&gt;I have shipped a lot of software that uses language models. I built an agent platform that has run over a thousand real coding sessions. I can make these things work.&lt;/p&gt;
&lt;p&gt;I could not tell you why.&lt;/p&gt;
&lt;p&gt;Not really. Not in the way I can tell you why a database query is slow, or why a Kubernetes pod is stuck in &lt;code&gt;CrashLoopBackOff&lt;/code&gt;. For those I have a mental model that goes all the way down. For language models I had a bag of habits: this prompt works better, that model is smarter, use a bigger context. That is not understanding. That is superstition with good uptime.&lt;/p&gt;
&lt;p&gt;So I am fixing it. Over the next twelve months I am going to learn how these systems actually work, from the token all the way down to the GPU kernel, and I am going to publish every step here.&lt;/p&gt;
&lt;h2 id=&quot;the-specific-problem-with-knowing-enough&quot;&gt;The specific problem with knowing &quot;enough&quot;&lt;/h2&gt;
&lt;p&gt;Here is a real thing that happened. I looked at my agent platform&#x27;s bill: $10,239 across 1,098 sessions. I looked at the token counts and found that 97.5% of my input tokens were cache reads, which sounds excellent, and that I had also written 67 million tokens &lt;em&gt;into&lt;/em&gt; the cache, which costs 1.25 times the normal rate.&lt;/p&gt;
&lt;p&gt;Was that good? Was it terrible? I genuinely did not know. I did not know because I did not understand what a prompt cache actually is at the level of the serving system. I was reading a dashboard and guessing.&lt;/p&gt;
&lt;p&gt;That is the gap. Not &quot;can I build with this&quot; but &quot;can I reason about this&quot;. The second one is what lets you make a decision that is not a coin flip.&lt;/p&gt;
&lt;h2 id=&quot;what-i-am-actually-going-to-do&quot;&gt;What I am actually going to do&lt;/h2&gt;
&lt;p&gt;The plan has three parts running at the same time.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Learn the theory properly.&lt;/strong&gt; Not paper-skimming. Implement-it-or-you-did-not-read-it. Five areas:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;How a model is built.&lt;/strong&gt; Tokenizers, attention, the transformer block, scaling laws.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How it is trained at scale.&lt;/strong&gt; Memory arithmetic, parallelism, kernels, floating point numerics.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How it is taught to behave.&lt;/strong&gt; Fine-tuning, RLHF, DPO, GRPO, verifiable rewards.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How it is served fast.&lt;/strong&gt; KV cache, batching, paged attention, speculative decoding.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How you know it works.&lt;/strong&gt; Evaluation statistics, agent benchmarks, interpretability, calibration.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Build the exercises.&lt;/strong&gt; A byte-pair encoding tokenizer from nothing. A transformer with no library helpers. Train GPT-2 small and log every dollar. Write a fused kernel in Triton. Implement DPO from the loss function up. Serve a model and chart the latency curves. Train a sparse autoencoder and look inside a model&#x27;s head.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Apply it to something real.&lt;/strong&gt; This is the part I think matters most, and it is the part almost no learning plan has. I have a running agent platform producing real telemetry: 1,098 sessions, 152,359 conversation events, 906 human approval decisions, all with cost and outcome labels attached. That is not a toy dataset. Every concept I learn gets pointed at that data to see if it survives contact with reality.&lt;/p&gt;
&lt;h2 id=&quot;the-five-things-i-am-going-to-build&quot;&gt;The five things I am going to build&lt;/h2&gt;
&lt;p&gt;These are the applied projects, and they double as the reason to keep going when week nine gets boring.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A real benchmark for my own agent.&lt;/strong&gt; Right now I change a prompt, run it once, and decide it is better. That is not evidence, it is vibes. I want confidence intervals and paired comparisons on a private task set drawn from my own history.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A prompt cache audit.&lt;/strong&gt; Find every source of cache invalidation in my prompt assembly and fix it. This is the one with the fastest payback and it is the subject of post 14.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A failure predictor.&lt;/strong&gt; 115 of my 1,098 sessions failed. I have the complete event trail leading to each one. Can a model predict a doomed session from its first few turns, so I can stop burning money on it? At $0.32 per turn, with some sessions running past a hundred turns, this pays for itself if it works at all.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;My task board as a reinforcement learning environment.&lt;/strong&gt; This one surprised me when I noticed it. My agent board already has objectively verifiable outcomes: tests pass, the build succeeds, the pull request merges. That is the definition of a reinforcement learning environment with verifiable rewards. I have been running one for a year without calling it that.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Teaching an agent when to ask for help.&lt;/strong&gt; This is the capstone and the one I care about most. My platform asks humans to approve agent actions, and today that decision is a static rule: this tool is allowed, that one is not. It should be an uncertainty decision. I have 906 recorded human judgments about when an agent should not be trusted alone. That is a calibration dataset. With conformal prediction you can turn it into a real guarantee, something like: &lt;em&gt;interrupt the human at most 15% of the time, while keeping the chance of an unrecoverable mistake under 2%.&lt;/em&gt; No assumptions about the distribution required. That would be a genuinely better product and a genuinely new result.&lt;/p&gt;
&lt;h2 id=&quot;why-publish-it&quot;&gt;Why publish it&lt;/h2&gt;
&lt;p&gt;Three honest reasons.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Writing is the test.&lt;/strong&gt; You can finish a lecture feeling like you understood it. You cannot finish a blog post feeling that way, because the gaps show up as sentences you cannot write. Every post in this series is a checkpoint I have to actually pass.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Public work compounds.&lt;/strong&gt; I am aiming at a research or engineering role at one of the frontier AI labs. Nobody there will care about a certificate. They might care about a reproduction that works, a benchmark harness that is statistically honest, or a result nobody has published because nobody else had the dataset.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The explainers I wanted did not exist.&lt;/strong&gt; Most writing about language models is either a press release or a paper. There is very little in between for the working engineer who wants the actual mechanism, in plain words, with real numbers. That is the gap I am writing into.&lt;/p&gt;
&lt;h2 id=&quot;what-this-series-will-and-will-not-be&quot;&gt;What this series will and will not be&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;It will be:&lt;/strong&gt; plain English, short sentences, and a real number attached to every claim. Sources at the end of every post, all of them checked. Negative results published alongside the wins, because the failures are usually more informative.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;It will not be:&lt;/strong&gt; takes about AGI, model release commentary, or prompt engineering tips. Those are well covered by people who enjoy them more than I do.&lt;/p&gt;
&lt;p&gt;The next post opens the books. Real production telemetry from a real agent platform: what a thousand agent sessions cost, where the money actually went, and the one number that turned out to be the most interesting thing in the whole dataset.&lt;/p&gt;
&lt;h2 id=&quot;sources&quot;&gt;Sources&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Stanford CS336, &lt;em&gt;Language Modeling from Scratch&lt;/em&gt; - the course this series leans on most heavily. Lectures and assignments are public. &lt;a href=&quot;https://github.com/stanford-cs336&quot; rel=&quot;noopener&quot;&gt;https://github.com/stanford-cs336&lt;/a&gt; and &lt;a href=&quot;https://online.stanford.edu/courses/cs336-language-modeling-scratch&quot; rel=&quot;noopener&quot;&gt;https://online.stanford.edu/courses/cs336-language-modeling-scratch&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Andrej Karpathy, &lt;em&gt;nanoGPT&lt;/em&gt; and &lt;em&gt;minbpe&lt;/em&gt; - the reference implementations for building a transformer and a tokenizer from nothing. &lt;a href=&quot;https://github.com/karpathy/nanoGPT&quot; rel=&quot;noopener&quot;&gt;https://github.com/karpathy/nanoGPT&lt;/a&gt; and &lt;a href=&quot;https://github.com/karpathy/minbpe&quot; rel=&quot;noopener&quot;&gt;https://github.com/karpathy/minbpe&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Hugging Face, &lt;em&gt;The Ultra-Scale Playbook&lt;/em&gt; - training large models across many GPUs. &lt;a href=&quot;https://huggingface.co/spaces/nanotron/ultrascale-playbook&quot; rel=&quot;noopener&quot;&gt;https://huggingface.co/spaces/nanotron/ultrascale-playbook&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Jacob Austin et al., &lt;em&gt;How to Scale Your Model&lt;/em&gt; - the systems arithmetic behind training and serving. &lt;a href=&quot;https://jax-ml.github.io/scaling-book/&quot; rel=&quot;noopener&quot;&gt;https://jax-ml.github.io/scaling-book/&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Stas Bekman, &lt;em&gt;Machine Learning Engineering Open Book&lt;/em&gt; - the practical failure modes nobody writes papers about. &lt;a href=&quot;https://github.com/stas00/ml-engineering&quot; rel=&quot;noopener&quot;&gt;https://github.com/stas00/ml-engineering&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Anastasios N. Angelopoulos and Stephen Bates, &lt;em&gt;A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification&lt;/em&gt;, arXiv:2107.07511. The mathematical basis for the capstone project. &lt;a href=&quot;https://arxiv.org/abs/2107.07511&quot; rel=&quot;noopener&quot;&gt;https://arxiv.org/abs/2107.07511&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Richard Sutton, &lt;em&gt;The Bitter Lesson&lt;/em&gt;. &lt;a href=&quot;http://www.incompleteideas.net/IncIdeas/BitterLesson.html&quot; rel=&quot;noopener&quot;&gt;http://www.incompleteideas.net/IncIdeas/BitterLesson.html&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content>
  </entry>
</feed>
