Build reliable Build reliable AI evaluation frameworks that measure quality, safety, grounding, and production readiness across modern LLM and SLM applications
Free with your book: DRM-free PDF version + access to Packt's next-gen Reader*
Key Features:
- Design evaluation frameworks for LLMs, SLMs, multimodal, reasoning, and agentic AI systems
- Measure quality, safety, grounding, robustness, and production readiness with practical metrics
- Apply unified evaluation methods to text, multimodal, and agentic AI systems
Book Description:
Modern AI systems are expected to do far more than generate fluent text. They should be able to retrieve information, reason through complex problems, understand images and documents, call external tools, execute workflows, and support critical business decisions. Evaluating these systems requires methods that go beyond traditional NLP benchmarks.
Taking a product-first approach, this book presents evaluation as a continuous operational capability spanning training, inference, and end-to-end system operation. You'll learn how to connect evaluation metrics directly to deployment gates, rollback criteria, monitoring systems, and production reliability objectives.
Using practical examples and real-world workflows, you'll explore evaluation strategies for text LLMs, vision-language models, multimodal conversational systems, mixture-of-experts architectures, reasoning models, agentic systems, retrieval pipelines, Text2SQL and Text2Cypher systems, embedding models, OCR workflows, and guardrail SLMs. You'll also learn how to manage non-determinism, design repeatable test suites, validate tool execution, and measure long-horizon agent behavior in production.
By the end of the book, you'll be able to design robust evaluation systems that help teams deploy reliable, safe, and economically viable LLM-powered applications with confidence.
*Email sign-up and proof of purchase required
What You Will Learn:
- Design repeatable evaluation pipelines for LLM systems
- Assess inference quality, latency, and operational cost
- Evaluate multimodal, agentic, and reasoning AI systems
- Build regression gates and deployment evaluation workflows
- Detect hallucinations and grounding failures in VLMs
- Assess routing stability in mixture-of-experts models
- Evaluate Text2SQL, OCR, and retrieval-based systems
- Translate evaluation signals into production decisions
Who this book is for:
ML engineers, GenAI engineers, AI architects, data scientists, platform engineers, and engineering managers responsible for deploying LLM-powered systems in production will benefit from this book. Applied AI researchers and technical decision-makers looking to measure reliability, safety, and operational readiness across modern AI systems will also find it valuable. Readers should have a working understanding of machine learning, Python, and modern LLM concepts.
Table of Contents
- Foundations of LLM Evaluation: Core Concepts and Primitives
- Building Reliable Text-Only LLMs Through Training-Time Evaluation
- Controlling Text-Only LLM Behavior at Inference Time
- Grounding and Reliability in Vision Language Models During Training
- Evaluating Visual Grounding and Reliability at Inference Time
- Evaluating Multimodal Conversational LLMs Across Training and Inference
- Evaluating Routing and Reliability in Mixture of Experts LLMs
- Evaluating Reliability and Control in Computer-Using Agent Systems
- Evaluating Information Extraction and Document-Understanding LLMs
- Evaluating Reasoning LLMs in Depth
- Evaluating Specialized LLM Systems
外文書商品之書封,為出版社提供之樣本。實際出貨商品,以出版社所提供之現有版本為主。部份書籍,因出版社供應狀況特殊,匯率將依實際狀況做調整。
無庫存之商品,在您完成訂單程序之後,將以空運的方式為你下單調貨。為了縮短等待的時間,建議您將外文書與其他商品分開下單,以獲得最快的取貨速度,平均調貨時間為1~2個月。
為了保護您的權益,「三民網路書店」提供會員七日商品鑑賞期(收到商品為起始日)。
若要辦理退貨,請在商品鑑賞期內寄回,且商品必須是全新狀態與完整包裝(商品、附件、發票、隨貨贈品等)否則恕不接受退貨。