Discover how AIOps is transforming the observability landscape for cloud-native and traditional systems. Learn how to build, monitor, and operate resilient services using AI-drive dynamic insights for smarter and more scalable operations
Key Features:
- Practical Integration of AI and Observability in Modern Engineering Workflows
- Real-World Use Cases Grounded in Industry Experience
- Tailored for Modern Engineering Roles and Organizations
Book Description:
Observability is mandatory for building and operating cloud-native distributed systems. Tools like OpenTelemetry have standardized how observability data is sourced, and AI now transforms how we extract value from the vast amounts of observability data generated by modern systems. This book guides you in implementing scalable observability, improving engineering efficiency with AI, and integrating observability throughout the Software Development Lifecycle (SDLC) via modern self-service internal developer platforms.
You'll start with observability basics and learn how AIOps enhances signal correlation, anomaly detection, and root-cause analysis. Using real-world examples, the book demonstrates how to implement AIOps, build proactive detection pipelines, and automate diagnostics and remediation. You'll explore best practices for expanding observability using OpenTelemetry, Prometheus, Grafana, Dynatrace, Datadog, and New Relic alongside machine learning models, ensuring your systems are accurate, efficient, and secure.
You'll also learn how to benchmark, measure, and secure your AIOps implementation, and gain a practical understanding of software compliance and how it applies to your systems. By the end of this book, you'll be ready to design and deliver AIOps-enabled observability solutions that make cloud-native systems more resilient, efficient, and secure.
What You Will Learn:
- Build observability pipelines for logs, metrics, traces and events
- Implement standards such as OpenTelemetry and Prometheus
- Correlate signals from multiple sources for better incident triage
- Apply AI/ML for anomaly detection and root cause analysis
- Design scalable architectures for intelligent monitoring
- Automate resiliency through self-healing and remediation agents
Who this book is for:
This book is for Software engineers and engineering leaders working on teams with operational responsibilities, such as platform engineering, site reliability engineering (SRE), DevOps, or application development, who want to integrate AIOps capabilities into their workflows will benefit from this book. If your team is responsible for building and running high-performing, resilient software systems, this book is for you.
Table of Contents
- Observability: The Art of Turning Data into Insights
- The Elephant in the Room: Artificial Intelligence
- From Observability to AIOps and the Use Cases it Solves Today
- ACME Financial Services: Implementing AIOps
- Democratizing Observability: A Primer to Self-Service Platforms
- The Observability Agent: Real-Life Use Cases
- ACME Financial Services: How to Move from AIOps to Agentic Platforms
- Evolving Operations: Proactive > Preventive > Self-Driven Architecture
- No Future Without Challenges
- ACME Financial Services: How Will the AI Future Shape Our Company?
外文書商品之書封,為出版社提供之樣本。實際出貨商品,以出版社所提供之現有版本為主。部份書籍,因出版社供應狀況特殊,匯率將依實際狀況做調整。
無庫存之商品,在您完成訂單程序之後,將以空運的方式為你下單調貨。為了縮短等待的時間,建議您將外文書與其他商品分開下單,以獲得最快的取貨速度,平均調貨時間為1~2個月。
為了保護您的權益,「三民網路書店」提供會員七日商品鑑賞期(收到商品為起始日)。
若要辦理退貨,請在商品鑑賞期內寄回,且商品必須是全新狀態與完整包裝(商品、附件、發票、隨貨贈品等)否則恕不接受退貨。