NLP — Interactive Lecture
Reinforcement Learning from Human Feedback (RLHF) is a technique for fine-tuning large language models (LLMs) so their outputs better align with human preferences. Instead of relying solely on next-token prediction (which optimises likelihood), RLHF trains the model to maximise a reward signal learned from human judgements.
The RLHF pipeline consists of three sequential stages, each building on the previous one:
Stage 1 — Supervised Fine-Tuning (SFT): Start with a pre-trained LLM and fine-tune it on high-quality demonstration data (prompt, ideal-response pairs). This gives the model a baseline ability to follow instructions and produce coherent responses.
Stage 2 — Reward Model Training: Collect human preference data by showing annotators pairs of model responses and asking which is better. Train a reward model R(x, y) that predicts a scalar score for any (prompt, response) pair. The reward model learns to capture human preferences via the Bradley-Terry model:
where yw is the preferred (winning) response and yl is the rejected (losing) response.
Stage 3 — RL Optimisation (PPO): Use the reward model as the environment's reward signal and fine-tune the SFT model using Proximal Policy Optimisation (PPO). A KL divergence penalty prevents the policy from straying too far from the SFT baseline, avoiding reward hacking.
The reward model is trained on pairwise comparisons. Given a prompt x and two responses y1, y2, a human annotator indicates which is preferred. The Bradley-Terry model assumes:
The reward model is trained to minimise the negative log-likelihood of the observed preferences:
Without constraints, the policy would learn to exploit weaknesses in the reward model — producing outputs that receive high reward but are actually low quality (e.g., repetitive phrases that the reward model rates highly). The KL divergence penalty KL(π || πref) keeps the policy close to the SFT model:
Direct Preference Optimisation (DPO) is a newer approach that skips the reward model entirely. It reparameterises the RLHF objective into a classification loss directly on the preference pairs:
DPO is simpler to implement (no separate reward model, no RL loop), but RLHF with PPO can be more powerful when you have a strong reward model and want to optimise over diverse responses.