Reinforcement Learning from Human Feedback (RLHF)

NLP — Interactive Lecture

Theory & Background

What is RLHF?

Reinforcement Learning from Human Feedback (RLHF) is a technique for fine-tuning large language models (LLMs) so their outputs better align with human preferences. Instead of relying solely on next-token prediction (which optimises likelihood), RLHF trains the model to maximise a reward signal learned from human judgements.

Core Idea
Traditional LLM training optimises: max log P(token | context)
RLHF adds a preference-based objective: max E[R(response)] - β · KL(π || πref)
where R is the learned reward, π is the policy, πref is the reference (SFT) policy, and β controls the KL penalty.

The Three Stages of RLHF

The RLHF pipeline consists of three sequential stages, each building on the previous one:

Stage 1 — Supervised Fine-Tuning (SFT): Start with a pre-trained LLM and fine-tune it on high-quality demonstration data (prompt, ideal-response pairs). This gives the model a baseline ability to follow instructions and produce coherent responses.

Stage 2 — Reward Model Training: Collect human preference data by showing annotators pairs of model responses and asking which is better. Train a reward model R(x, y) that predicts a scalar score for any (prompt, response) pair. The reward model learns to capture human preferences via the Bradley-Terry model:

P(y_w > y_l | x) = sigmoid(R(x, y_w) - R(x, y_l))

where yw is the preferred (winning) response and yl is the rejected (losing) response.

Stage 3 — RL Optimisation (PPO): Use the reward model as the environment's reward signal and fine-tune the SFT model using Proximal Policy Optimisation (PPO). A KL divergence penalty prevents the policy from straying too far from the SFT baseline, avoiding reward hacking.

Why Not Just Use More Supervised Data?
Supervised learning requires knowing the correct output for every input — but for open-ended tasks like "write a helpful response," there are many valid answers. RLHF sidesteps this by learning from comparisons: it's much easier for humans to judge "Response A is better than Response B" than to write the perfect response from scratch.

The Reward Model (Bradley-Terry)

The reward model is trained on pairwise comparisons. Given a prompt x and two responses y1, y2, a human annotator indicates which is preferred. The Bradley-Terry model assumes:

P(y_1 > y_2) = exp(R(x,y_1)) / (exp(R(x,y_1)) + exp(R(x,y_2))) = sigmoid(R(x,y_1) - R(x,y_2))

The reward model is trained to minimise the negative log-likelihood of the observed preferences:

L = -E[ log sigmoid(R(x, y_w) - R(x, y_l)) ]

KL Divergence & Reward Hacking

Without constraints, the policy would learn to exploit weaknesses in the reward model — producing outputs that receive high reward but are actually low quality (e.g., repetitive phrases that the reward model rates highly). The KL divergence penalty KL(π || πref) keeps the policy close to the SFT model:

KL(π || π_ref) = E_π[ log(π(y|x) / π_ref(y|x)) ]
Reward Hacking
If β is too small, the model can "hack" the reward model by finding adversarial outputs that score high but aren't actually good. If β is too large, the model barely changes from the SFT baseline and doesn't improve. Finding the right β is critical.

DPO: A Simpler Alternative

Direct Preference Optimisation (DPO) is a newer approach that skips the reward model entirely. It reparameterises the RLHF objective into a classification loss directly on the preference pairs:

L_DPO = -E[ log sigmoid(β(log π(y_w|x)/π_ref(y_w|x) - log π(y_l|x)/π_ref(y_l|x))) ]

DPO is simpler to implement (no separate reward model, no RL loop), but RLHF with PPO can be more powerful when you have a strong reward model and want to optimise over diverse responses.

RLHF vs DPO at a Glance
RLHF: SFT → Reward Model → PPO  |  More flexible, harder to train
DPO: SFT → Direct preference loss  |  Simpler, single training stage
RLHF Pipeline Overview
The RLHF process takes a pre-trained LLM through three stages: supervised fine-tuning, reward model training, and RL optimisation. Click any stage on the right to explore it in detail.
StageOverview
InputPre-trained LLM
OutputAligned Model
Data NeededVaries by stage
MethodMulti-stage

RLHF Pipeline

Stage 1: SFT Stage 2: Reward Model Stage 3: PPO
Stage 1: SFT
Supervised Fine-Tuning on demonstration data
Pre-trained LLM → SFT Model
→
Stage 2: Reward Model
Train reward from human pairwise comparisons
Preferences → R(x,y)
→
Stage 3: PPO
RL fine-tuning with reward & KL penalty
SFT + R → Aligned Model
Pipeline Overview