How to Read a 190-Post Blog: A Map of Sebastian Raschka's Writing
Sebastian Raschka's blog has about 190 posts over thirteen years. A reading order that moves from method to map to practice to the current frontier.
Some blogs are too big to read. Sebastian Raschka’s blog has about 190 posts across thirteen years, from early notes on PCA and naive Bayes to the 2026 notes on attention variants and reasoning models. You can’t read it front to back, and sorting by date won’t help. The useful question is which order to read it in.
One article to read first
His Recommendations for Getting the Most Out of a Technical Book (November 2025) is short and sets the method. It gives five steps for each chapter:
- Read it once, offline, without code, for about twenty minutes. Don’t look anything up. The goal is the big picture.
- Read it again and type the code yourself instead of copying it. If your results differ from the book’s, check the repository, then package versions, random seeds and hardware, and ask the author last.
- Do the exercises. Try properly before looking at the solutions.
- Go back over your highlights and notes, and look up what is still unclear.
- Use an idea from the chapter in a small project of your own.
He adds that none of this is fixed. A chapter you already know can be skimmed, and one without code skips the code steps. The method is a starting point, not a rule.
It is the right first read because it tells you how to read the other 189: once for shape, once for detail, then build something.
The blog in eight parts
After that, the blog sorts into eight groups.
1. Building from scratch
Code first, with each mechanism built step by step.
- Understanding and Coding Self-Attention, Multi-Head, Causal and Cross-Attention (2023.2)
- Implementing a BPE Tokenizer From Scratch (2025.1)
- Building LLMs from the Ground Up: A 3-hour Coding Workshop (2024.9)
- Coding LLMs from the Ground Up: A Complete Course (2025.5)
- Building A GPT-Style LLM Classifier From Scratch (2024.9), a spam classifier
- Understanding and Coding the KV Cache in LLMs from Scratch (2025.6)
- Understanding and Implementing Qwen3 From Scratch (2025.9)
- Developing an LLM: Building, Training, Finetuning (2024.6), a one-hour talk on the three stages of LLM development
2. Fine-tuning and parameter-efficient methods
- The LoRA series: Parameter-Efficient Finetuning (2023.4), LoRA (2023.4), Finetuning Falcon (2023.6), DoRA from Scratch (2024.2)
- Using and Finetuning Pretrained Transformers (2024.4)
- Instruction Pretraining LLMs (2024.7) and Instruction Masking and LoRA experiments (2024.6)
- Optimizing LLMs From a Dataset Perspective (2023.9)
3. Post-training and reasoning models
- New LLM Pre-training and Post-training Paradigms (2024.8)
- How Good Are the Latest Open LLMs? And Is DPO Better Than PPO? (2024.5)
- Understanding Reasoning LLMs (2025.2), four ways to build a reasoning model
- The State of Reinforcement Learning for LLM Reasoning (2025.4), on GRPO
- Inference-Time Compute Scaling (2025.3) and Categories of Inference-Time Scaling (2026.1)
- Controlling Reasoning Effort in LLMs (2026.7)
- His book Build a Reasoning Model From Scratch (published 2026.6)
4. How architectures evolved
- The Big LLM Architecture Comparison (2025.7), from GPT-2 to DeepSeek V3
- From GPT-2 to gpt-oss (2025.8)
- A Visual Guide to Attention Variants in Modern LLMs (2026.3): MHA, GQA, MLA, sparse attention
- Beyond Standard LLMs (2025.11): linear attention, text diffusion and more
- The LLM Architecture Gallery (93 diagrams) with a comparison tool, and recent short notes on Kimi K3, Gemma 4 and Nemotron 3 Ultra
5. Evaluation, paper lists and trends
- Understanding the 4 Main Approaches to LLM Evaluation (2025.10): multiple-choice benchmarks, verifiers, leaderboards, LLM judges
- The yearly paper lists: 2024, 2025 January to June, 2025 July to December, 2026 January to May
- The State Of LLMs 2025 (2025.12), a yearly review with predictions for 2026
- State of AI 2026, a 4.5-hour interview with Lex Fridman and Nathan Lambert
6. Learning methods and workflow
- Recommendations for Getting the Most Out of a Technical Book (2025.11)
- My Workflow for Understanding LLM Architectures (2026.4)
- Keeping Up With AI Research and News (2023.3)
- Understanding LLMs: A Transformative Reading List (2023.2), a list of classic papers
7. Engineering practice
- PyTorch training optimization (2023): faster training, mixed precision, memory optimization and gradient accumulation
- DGX Spark and Mac Mini for Local PyTorch Development (2025.10)
- Using Local Coding Agents (2026.6)
8. Early classical machine learning (2013-2022)
PCA, LDA, naive Bayes and the model-evaluation series. Skip them unless you need to fill in traditional ML basics. One exception is Losses Learned: Optimizing Negative Log-Likelihood and Cross-Entropy in PyTorch (2022.4), which bears directly on how loss is computed for LLMs.
Suggested reading order
This is a suggestion, not a template.
Now:
- Recommendations for Getting the Most Out of a Technical Book
- Developing an LLM: Building, Training, Finetuning, the one-hour talk, as a map of the whole field
- Building A GPT-Style LLM Classifier From Scratch
- Losses Learned, to firm up the loss calculation
After the from-scratch material:
- The LoRA series (Parameter-Efficient Finetuning, LoRA, Finetuning Falcon, DoRA from Scratch), then New LLM Pre-training and Post-training Paradigms
- Understanding Reasoning LLMs, then the GRPO article
- The Big LLM Architecture Comparison
Long-term reference: the yearly paper lists (2024, 2025 January to June, 2025 July to December, 2026 January to May) and the LLM Architecture Gallery.
That order moves from method to map to practice to the current frontier. It works for any large body of technical writing. Pick the article that teaches you how to read the rest, build a frame, then go deeper in the order your work needs.