Doesn't suit? No problem! You can return items for up to 30 days
You won't go wrong with a gift voucher. The gift recipient can choose anything from our offer.
Up to 30 days for returns
Master the complete lifecycle of modern language model alignment.
This book provides a practical and rigorous guide to RLHF (Reinforcement Learning from Human Feedback), DPO, Constitutional AI, RLAIF, and modern LLM evaluation.
Covering both theory and implementation, you'll learn how alignment systems are built, optimized, evaluated, and monitored in production environments.
Inside you'll find:
• Mathematical foundations of RLHF, PPO, reward modeling, and preference optimization
• Modern alignment methods including DPO, KTO, ORPO, Constitutional AI, and RLAIF
• Hands-on implementations using the Hugging Face ecosystem
• Multidimensional evaluation of toxicity, bias, factuality, helpfulness, and harmlessness
• Production monitoring, guardrails, observability, and drift detection
Written for ML engineers, data scientists, AI researchers, and technical leaders, this book bridges the gap between alignment research and real-world deployment.
A complete technical reference for building safer, more reliable, and production-ready language models.
Hi! I'm Libroamiko, your book advisor.
How can I help you?