Saurabh Singh
Bookshelf

Book notes · 2 min read

Generative AI on Kubernetes

by Roland Huß & Daniele Zonca

  • Non-fiction
  • AI engineering
  • Kubernetes
  • MLOps
An LLM is just another workload until you count the GPUs. Then scheduling, metrics, and identity all have to change.
Author
Roland Huß & Daniele Zonca
Type
Non-fiction
Genre
AI engineering, Kubernetes, MLOps
Pages
404
Published
2026

Most of my career has been spent making ordinary services boring to run. This book is about what happens when the service is a large language model and "ordinary" stops applying.

What it's about

Roland Huß and Daniele Zonca walk through the full life of a generative AI workload on Kubernetes: deploying models behind optimized inference runtimes, scheduling GPUs, scaling across nodes, monitoring what actually matters for LLMs, choosing between fine-tuning and retrieval augmentation, evaluating models against benchmarks, and running agentic applications that call tools securely. It is written for the people who have to keep all of that up at 3am, not just demo it.

What stuck

The GPU is the scheduling problem. On a normal platform, CPU and memory are elastic enough that you mostly think about replicas. Here the accelerator is scarce, expensive, and specific. Detecting what hardware a node actually has, and placing work to match, becomes a first-class platform concern rather than a node-pool detail.

Your dashboards measure the wrong things. Request latency and error rate still matter, but an LLM endpoint lives or dies on time to first token and token throughput. A service can be "up" with healthy p95s and still feel broken to a user watching a blank chat box.

Fine-tune or retrieve is a design decision, not a trend. The book frames it as a trade-off between baking knowledge into weights and fetching it at request time, with different costs for freshness, evaluation, and operations. I appreciated that it treats this as an engineering choice with consequences, not a default.

Agents make identity everyone's problem. Once a model can call tools, the platform has to answer who is acting and with what authority. I spent a whole post on exactly this question, whose identity the tool sees, so it was good to see it treated as part of operations rather than an afterthought.

What I'd apply

  • Put LLM-specific signals, time to first token and tokens per second, on the same dashboard as the classic golden signals from day one.
  • Treat GPU capacity like any other scarce shared resource: explicit quotas, visible queues, and a clear owner.
  • Make "fine-tune or retrieve?" a written design decision with its evaluation plan attached, the same way we would record any other architecture choice.