Generative AI on Kubernetes
by Roland Huß & Daniele Zonca
- Non-fiction
- AI engineering
- Kubernetes
- MLOps
An LLM is just another workload until you count the GPUs. Then scheduling, metrics, and identity all have to change.
- Author
- Roland Huß & Daniele Zonca
- Type
- Non-fiction
- Genre
- AI engineering, Kubernetes, MLOps
- Pages
- 404
- Published
- 2026
Most of my career has been spent making ordinary services boring to run. This book is about what happens when the service is a large language model and "ordinary" stops applying.
What it's about
Roland Huß and Daniele Zonca walk through the full life of a generative AI workload on Kubernetes: deploying models behind optimized inference runtimes, scheduling GPUs, scaling across nodes, monitoring what actually matters for LLMs, choosing between fine-tuning and retrieval augmentation, evaluating models against benchmarks, and running agentic applications that call tools securely. It is written for the people who have to keep all of that up at 3am, not just demo it.
What stuck
The GPU is the scheduling problem. On a normal platform, CPU and memory are elastic enough that you mostly think about replicas. Here the accelerator is scarce, expensive, and specific. Detecting what hardware a node actually has, and placing work to match, becomes a first-class platform concern rather than a node-pool detail.
Your dashboards measure the wrong things. Request latency and error rate still matter, but an LLM endpoint lives or dies on time to first token and token throughput. A service can be "up" with healthy p95s and still feel broken to a user watching a blank chat box.
Fine-tune or retrieve is a design decision, not a trend. The book frames it as a trade-off between baking knowledge into weights and fetching it at request time, with different costs for freshness, evaluation, and operations. I appreciated that it treats this as an engineering choice with consequences, not a default.
Agents make identity everyone's problem. Once a model can call tools, the platform has to answer who is acting and with what authority. I spent a whole post on exactly this question, whose identity the tool sees, so it was good to see it treated as part of operations rather than an afterthought.
What I'd apply
- Put LLM-specific signals, time to first token and tokens per second, on the same dashboard as the classic golden signals from day one.
- Treat GPU capacity like any other scarce shared resource: explicit quotas, visible queues, and a clear owner.
- Make "fine-tune or retrieve?" a written design decision with its evaluation plan attached, the same way we would record any other architecture choice.