Reinforcement Learning from Human Feedback cover art

Reinforcement Learning from Human Feedback

LLM alignment and post-training

Preview

Get 30 days of Standard free

£5.99/mo after trial. Cancel monthly.
Try for £0.00
More purchase options

Reinforcement Learning from Human Feedback

By: Nathan Lambert
Narrated by: Julie Brierley
Try for £0.00

£5.99 a month after 30 days. Cancel anytime.

Buy Now for £14.63

Buy Now for £14.63

Season's Savings | £0.99/mo for 3 months

£5.99/mo after 3 months - terms apply. Cancel monthly

"Reinforcement Learning from Human Feedback: LLM alignment and post-training" helps you understand how modern AI models can be adapted to better match the needs and expectations of their users. Rather than surveying the vast field of reinforcement learning, elite AI researcher Nathan Lambert concentrates exclusively on RLHF and its immediate importance to post-training generative AI models.

This compact book gets right to the point. Early chapters establish the training overview, explain instruction fine-tuning, and build reliable reward models. The middle chapters transition into the heart of alignment, exploring core policy gradient algorithms, Direct Preference Optimization (DPO), and inference-time scaling. Later chapters tackle the messy reality of data, guiding you through preference data collection, synthetic data generation, and the nuances of function calling.

As you go, you will see how these post-training methods work. You will explore common failure modes, such as qualitative over-optimization, reward hacking, and the unreliability of external evaluation comparisons. Difficult concepts like KL regularization, proximal policy optimization, and generative reward modeling are clarified with hands-on experiments.

The book’s seventeen short chapters lay out the core material, while supplements like vocabulary definitions, compute cost management, evaluation variance, and training performance tracking appear in handy appendixes.

About the listener:

For established engineers, AI scientists, and students trying to get a practical foothold in AI model alignment.

About the author:

Dr. Nathan Lambert is a leading AI researcher known for leading post-training at the Allen Institute for AI. With previous experience at HuggingFace, DeepMind, and Meta, he is a passionate advocate for open models. His work focuses on increasing access to, and the understanding of, AI technology—empowering listener to contribute to the advancement of AI outside closed corporate labs.

PLEASE NOTE: When you purchase this title, the accompanying PDF will be available in your Audible Library along with the audio.

©2026 Manning Publications (P)2026 Manning Publications
Computer Science History & Culture Machine Theory & Artificial Intelligence Data Science Management Technology Programming Machine Learning Software Software Development
adbl_web_anon_alc_button_suppression_t1
No reviews yet