Run production-grade GenAI workloads by containerizing, serving, and scaling LLMs, agents, and multi-model pipelines with Docker, MCP, and Kubernetes for cloud platforms
Key Features:
- Deploy and operate local and edge-friendly LLM inference using Docker Model Runner and an OpenAI-compatible API
- Orchestrate multi-model and multi-agent workloads with Docker Compose and Kubernetes patterns used by platform teams
- Purchase of the print or Kindle book includes a free PDF eBook
Book Description:
Modern AI systems don't fail at modeling; they fail in production. Moving from experiments to reliable, scalable systems requires more than notebooks and scripts. It requires infrastructure.
Operational AI with Docker shows you how to build, deploy, and operate AI systems that work beyond a single machine. You'll learn how to use Docker as a consistent runtime for machine learning workflows, package models as reproducible artifacts, and run them reliably across environments.
Starting with containerized machine learning, you'll progress to model serving, AI deployment, and scalable infrastructure using Kubernetes. You'll implement production-ready patterns for resource management, autoscaling, observability, and performance tuning, ensuring your AI workloads remain stable under real-world conditions.
The book goes beyond traditional MLOps by introducing agentic AI systems, including autonomous agents, multi-agent architectures, and secure execution environments. You'll also explore modern integration patterns using the Model Context Protocol (MCP), enabling AI systems to interact safely with tools, APIs, and data sources.
By the end of this book, you'll be able to design and operate production AI systems that are reproducible, scalable, and ready for real-world deployment using Docker and Kubernetes.
What You Will Learn:
- Containerize GenAI services using Docker images, registries, and Compose-based deployment stacks
- Package and distribute models as OCI artifacts for repeatable builds and controlled promotions across environments
- Choose GGUF quantization levels to balance cost, latency, and accuracy for cloud and hybrid runtimes
- Serve LLMs via Docker Model Runner with an OpenAI-compatible API suitable for internal platforms
- Integrate tools and data securely using MCP and Docker MCP Gateway with least-privilege access patterns
Who this book is for:
Cloud engineers, DevOps engineers, SREs, and platform engineers who need to deploy, operate, and scale GenAI workloads using Docker and Kubernetes on cloud, hybrid, or edge environments. You should be comfortable with the command line and basic service operations; prior Docker or Kubernetes exposure is helpful but not required.
Table of Contents
- Docker Desktop - The Runtime Foundation for AI/ML Workflows
- Understanding AI Models in Docker
- Model Service with Docker Model Runner
- Docker Offload for AI and ML Workflows
- Running ML Container Models on Kubernetes
- Protocol-Based AI Integration with MCP
- Building Autonomous AI Agents
- Multi-Model and Multi-Agent Architectures
- Advanced Agent Orchestration
外文書商品之書封,為出版社提供之樣本。實際出貨商品,以出版社所提供之現有版本為主。部份書籍,因出版社供應狀況特殊,匯率將依實際狀況做調整。
無庫存之商品,在您完成訂單程序之後,將以空運的方式為你下單調貨。為了縮短等待的時間,建議您將外文書與其他商品分開下單,以獲得最快的取貨速度,平均調貨時間為1~2個月。
為了保護您的權益,「三民網路書店」提供會員七日商品鑑賞期(收到商品為起始日)。
若要辦理退貨,請在商品鑑賞期內寄回,且商品必須是全新狀態與完整包裝(商品、附件、發票、隨貨贈品等)否則恕不接受退貨。