59,40 €
Versandkostenfrei per Post / DHL
Lieferzeit 2-3 Werktage ab Escheinungsdatum. Dieses Produkt erscheint am 04.10.2026
Moving from foundational concepts to advanced practice, Reliable Evals for LLMs and AI Agents introduces the four core levers of effective evals: sets, templates, metrics, and evaluators. It then extends these to the unique challenges of autonomous AI agents, where systems perceive, reason, act, and adapt in iterative loops that demand fundamentally different eval approaches. Along the way, it guides readers through benchmark selection, custom eval set design, statistical rigor in metrics, human and LLM-as-a-judge rating strategies, and the infrastructure needed to automate evals at scale.
For engineering leaders, applied researchers, data scientists, and product teams shipping LLM- and agent-powered experiences, this volume offers a blueprint for building eval flywheels that continuously improve AI quality. It shows how to progress from ad-hoc checks to production-grade eval systems, align model metrics with real user satisfaction, integrate offline evals with online A/B testing, and design accessible interfaces that democratize rigorous testing across an organization.
Moving from foundational concepts to advanced practice, Reliable Evals for LLMs and AI Agents introduces the four core levers of effective evals: sets, templates, metrics, and evaluators. It then extends these to the unique challenges of autonomous AI agents, where systems perceive, reason, act, and adapt in iterative loops that demand fundamentally different eval approaches. Along the way, it guides readers through benchmark selection, custom eval set design, statistical rigor in metrics, human and LLM-as-a-judge rating strategies, and the infrastructure needed to automate evals at scale.
For engineering leaders, applied researchers, data scientists, and product teams shipping LLM- and agent-powered experiences, this volume offers a blueprint for building eval flywheels that continuously improve AI quality. It shows how to progress from ad-hoc checks to production-grade eval systems, align model metrics with real user satisfaction, integrate offline evals with online A/B testing, and design accessible interfaces that democratize rigorous testing across an organization.
Previously, Alexei led Google's Gemini Core Modeling and Evals Data Science Research organization, driving evaluation research and training-data quality initiatives that shaped the performance, reliability, and safety of Gemini models.
Earlier in his career, he was a Data Science Manager at Twitter, overseeing evaluation of Home Timeline ranking and personalization, and a Principal Data Science Manager at Microsoft, where he guided cross-functional teams that shipped production-level Azure customer-experience solutions.
Alexei co-authored Machine Learning Governance for Managers (Springer, 2024), holds an MBA from Duke University's Fuqua School of Business, and earned a [...]. in Electrical Engineering and Computer Science from Tel Aviv University.
Liliya Lavitas is an accomplished Data Science leader with a deep background in statistical modeling and machine learning. Liliiya currently leads a Data Science team in Google Search incorporating AI features in Google Search experience. Previously, under leadership of Alexei, she has been leading a Data Science team at Google DeepMind, responsible for the rigorous evaluation and training of core Gemini capabilities, ensuring their performance and reliability. Prior to her role at Google, Liliya managed Data Science teams at Netflix and Twitter. Liliya earned her Ph.D. in Statistics from Boston University in 2017, with a thesis specializing in Time Series analysis.
Yueqing Wang is a statistician specializing in novel methodology for system evaluation. In recent years, she has developed new techniques and frameworks for Gen AI evaluation for Google Gemini and at Microsoft AI. Earlier, she worked on paid product ecosystem at YouTube, statistical evaluations of startups at Google Ventures, and Media Mix Modeling at Google Ads, among other things. She earned her PhD in Statistics from the University of California, Berkeley in 2012 with Professor Bin Yu.
| Erscheinungsjahr: | 2026 |
|---|---|
| Genre: | Informatik |
| Rubrik: | Naturwissenschaften & Technik |
| Medium: | Taschenbuch |
| Inhalt: |
x
193 S. 1 s/w Illustr. 31 farbige Illustr. 193 p. 32 illus. 31 illus. in color. |
| ISBN-13: | 9783032267481 |
| ISBN-10: | 303226748X |
| Sprache: | Englisch |
| Herstellernummer: | 978-3-032-26748-1 |
| Einband: | Kartoniert / Broschiert |
| Autor: |
Robsky, Alexei
Lavitas, Liliya Wang, Yueqing |
| Hersteller: |
Springer
Springer, Berlin |
| Verantwortliche Person für die EU: | Springer Verlag GmbH, Tiergartenstr. 17, D-69121 Heidelberg, juergen.hartmann@springer.com |
| Abbildungen: | X, 193 p. 32 illus., 31 illus. in color. |
| Maße: | 16 x 153 x 231 mm |
| Von/Mit: | Alexei Robsky (u. a.) |
| Erscheinungsdatum: | 04.10.2026 |
| Gewicht: | 0,318 kg |