A beginner’s field guide to AI language models
Train Your Own LLM
A Friendly, Practical Guide to How AI Language Models Work, and How to Fine-Tune and Build One Yourself
Zafrullah Khan, Ed.D.
Coming soon
Understand the machine first, then train, fine-tune, and deploy honestly.
Welcome, reader.
This page collects the book's companion resources. Everything you need is printed in the book; what's here is a convenience.
About the book
Behind every fluent answer is one simple idea: predicting the next token. This friendly guide shows how that trick works, then walks you from curiosity to a model you can run. Part I needs no code. Later parts cover fine-tuning on affordable hardware, training a tiny model from scratch, evaluating honestly, and deploying with clear, step-by-step recipes.
Who it's for
Curious readers, managers, students: Part I (Chapters 1–3) needs no code. It covers what a model really is, an honest history, and how to decide whether to train one.
Builders: Parts II–IV are hands-on, in Python with tested scripts. They cover your workbench, data, tokenization, training, fine-tuning, LoRA/QLoRA, preferences, a tiny model from scratch, evaluation and deployment.
Part V adds case studies (Chapter 15) and capstone projects (Chapters 16–20). The book is 20 chapters in five parts, about 800 pages.
Formats
Three editions, all still to be published on Amazon.
| Edition | Details | Buy |
|---|---|---|
| Paperback, black & white | 7 × 10 in, about 800 pages | |
| Premium colour paperback | 7 × 10 in, full colour | |
| Kindle ebook | Colour ebook |
Contents
20 chapters in five parts, about 800 pages, 7 × 10 in. Chapters 1–7 are drafted and will be revised. Later chapters are listed from the outline.
Part I: Understanding the Machine
| Ch. | Title | Status | Links |
|---|---|---|---|
| 1 | What an LLM Really Is, and Whether You Should Train One | Draft | Chapter 1 code · Chapter 1 toolkit |
| 2 | A Short, Honest History of How We Got Here | Draft | Chapter 2 code · Chapter 2 toolkit |
| 3 | Your Model, Your Rules: Picking a Project Worth Training | Draft | Chapter 3 code · Chapter 3 toolkit |
Part II: Foundations You Can Run
| Ch. | Title | Status | Links |
|---|---|---|---|
| 4 | Setting Up Your Workbench | Draft | Chapter 4 code · Chapter 4 toolkit |
| 5 | Data: The Part That Matters Most | Draft | Chapter 5 code · Chapter 5 toolkit |
| 6 | Tokenization Up Close | Draft | Chapter 6 code · Chapter 6 toolkit |
| 7 | How Training Actually Works | Draft | Chapter 7 code · Chapter 7 toolkit |
Part III: Build
| Ch. | Title | Status | Links |
|---|---|---|---|
| 8 | Your First Fine-Tune | Chapter 8 toolkit | |
| 9 | Doing More with Less: LoRA and QLoRA | Chapter 9 toolkit | |
| 10 | Teaching Preferences: DPO and a First Look at RL | Chapter 10 toolkit | |
| 11 | Training a Tiny Model from Scratch | Chapter 11 toolkit |
Part IV: Measure, Ship, Grow
| Ch. | Title | Status | Links |
|---|---|---|---|
| 12 | Is It Any Good? Evaluating Your Model | Chapter 12 toolkit | |
| 13 | The Deployment Cookbook | Chapter 13 toolkit | |
| 14 | Where to Go Next | Chapter 14 toolkit |
Part V: Case Studies and Capstone Projects
Chapter 15 is case studies. Chapters 16–20 are capstone projects. Individual capstone titles are still open.
| Ch. | Title | Status | Links |
|---|---|---|---|
| 15 | Case Studies | Chapter 15 toolkit | |
| 16 | Capstone project | Chapter 16 toolkit | |
| 17 | Capstone project | Chapter 17 toolkit | |
| 18 | Capstone project | Chapter 18 toolkit | |
| 19 | Capstone project | Chapter 19 toolkit | |
| 20 | Capstone project | Chapter 20 toolkit |
Code downloads
Every listing is printed in the book. These files save you typing. Large files (model checkpoints, datasets you generate) are not included; the scripts recreate them.
| File | Size | Updated | Notes |
|---|---|---|---|
| train-your-own-llm-ch01-code.zip | 2.8 KB | 2026-10-08 | Chapter 1 README |
| train-your-own-llm-ch02-code.zip | 907 B | 2026-10-08 | Chapter 2 README |
| train-your-own-llm-ch03-code.zip | 1.8 KB | 2026-10-08 | Chapter 3 README |
| train-your-own-llm-ch04-code.zip | 4.6 KB | 2026-10-08 | Chapter 4 README |
| train-your-own-llm-ch05-code.zip | 9.0 KB | 2026-10-08 | Chapter 5 README |
| train-your-own-llm-ch06-code.zip | 6.1 KB | 2026-10-08 | Chapter 6 README |
| train-your-own-llm-ch07-code.zip | 6.5 KB | 2026-10-08 | Chapter 7 README |
Chapters 8–20 stay unpublished here until those chapters are final. Checksums are in downloads.json beside the files.
Toolkit by chapter
Chapters 8–14 links are based on the outline and will be confirmed against the final text. Link check, Oct 8, 2026: the 65 links in this list returned HTTP 200.
Chapter 1. What an LLM Really Is, and Whether You Should Train One
Chapter 2. A Short, Honest History of How We Got Here
Chapter 3. Your Model, Your Rules: Picking a Project Worth Training
Chapter 4. Setting Up Your Workbench
Chapter 5. Data: The Part That Matters Most
Chapter 6. Tokenization Up Close
Chapter 7. How Training Actually Works
Chapter 8. Your First Fine-Tune
Chapter 9. Doing More with Less: LoRA and QLoRA
Chapter 10. Teaching Preferences: DPO and a First Look at RL
Chapter 11. Training a Tiny Model from Scratch
Chapter 12. Is It Any Good? Evaluating Your Model
Chapter 13. The Deployment Cookbook
Chapter 14. Where to Go Next
Chapter 15. Case Studies
Toolkit links will be added with this chapter.
Chapter 16. Capstone project
Toolkit links will be added with this chapter.
Chapter 17. Capstone project
Toolkit links will be added with this chapter.
Chapter 18. Capstone project
Toolkit links will be added with this chapter.
Chapter 19. Capstone project
Toolkit links will be added with this chapter.
Chapter 20. Capstone project
Toolkit links will be added with this chapter.
Updates as tools change
No updates yet. When a library changes in a way that affects the book, you'll find the fix here.
Keep learning
Learn how language models work
The guessing machine: what a language model really does, and how to close the gap
From the author
I teach how today's LLMs work. NanoSI is what I'm building next.