Offline RL: KL, Advantage Weighting, and Guidance
Offline RL
Have an offline dataset collected by some behaviour policy . We want to learn a new policy with greater expected return.
Obvious idea: choose actions that maximise a learned critic:
However, has only been trained on / near actions in . Far from that, can be very wrong. Thus, the learned policy can snowball errors in and just be very bad.
Therefore we want: prefer actions with high value, but remain close to the data policy.
Entropy-regularised optimisation
This is a good start, but it ignores the behaviour policy. Consider one fixed state . For convenience, write .
We optimise over all possible action distributions:
is the entropy temperature. Maximise , but don’t collapse the distribution (like SAC).
Thus:
Can solve with Lagrange:
Differentiate with respect to :
Using the normalisation constraint gives:
Closed-form solution: a Boltzmann distribution. As , more like argmax. For continuous actions, the sum becomes an integral.
Cool link: entropy is KL regularisation to a uniform distribution
Suppose is a uniform reference distribution. Then:
Since is constant:
Thus, entropy regularisation is a special case of “stay close to the reference distribution,” where the reference is uniform.
KL-regularised policy improvement
Penalise the policy for leaving the behaviour policy via KL:
At the fixed state, this is:
Using Lagrange again:
Solve for and use the constraint again:
Or equivalently:
Basically, weight the prior with the (softmax).
Side thought: reverse KL tends to fit modes better and forward KL tends to fit means better. We kind of prefer reverse KL here.
Move to advantage
For fixed , doesn’t depend on :
The second factor is constant over all actions, so it disappears when normalised. Thus:
Advantage-weighted regression
That equation describes the improved policy distribution, but we still need to fit a parameterised model . One way is weighted maximum likelihood:
Basically gradient ascent to make the data more likely (supervised learning), but now with an extra loss-weighting term. This is the idea behind advantage-weighted regression and related methods (clone every action, but clone good actions more).
Is what Flow Q-learning describes as weighted behavioural cloning. Remember that increasing log probability is what they all sort of do.
Issue with weighted MLE: very sensitive to errors in value estimates. The exponential can make the weights highly concentrated.
Let’s try to rewrite the way of representing it.
Rewrite as a product distribution
More generally, suppose:
where is a non-negative, non-decreasing function. It could be or . Just another way to weight the prior favourably.
CFGRL paper idea
Suppose a target density is a product:
For smooth, positive factors, taking logs and gradients:
Looks like the scores that diffusion models learn (kind of). For policy improvement, formally:
Thus, a link: regularised policy improvement ↔ product of behaviour and optimality ↔ generative-model guidance.
Condition on optimality
Introduce a binary variable , where means the action is considered good. We can train two distributions: and . The first is behaviour and the second is behaviour conditioned on optimality.
By Bayes:
Since the denominator doesn’t depend on :
Thus, we can represent a policy as a product of two terms by conditioning on optimality. If:
we get back the KL-regularised improved policy:
But we can also pick and still learn. We basically use samples and supervised learning here.
CFG refresher
Train the model to operate both conditionally and unconditionally. During training, the condition is sometimes supplied and sometimes dropped.
The model therefore learns an unconditioned field and an optimality-conditioned field . Here “unconditioned” still includes the state .
Again, via Bayes:
Add some -scaled version of that vector: from the unconditional score towards the conditional score. Thus, for the fields we trained:
or equivalently:
The product-distribution interpretation is:
If the optimality factor is proportional to , then gives the earlier exponential reweighting. If it already uses , then gives that target.
This suggests a way to learn the unconditional and conditional fields, then combine them at inference with a chosen .
What CFGRL does in offline RL
- Learn a value estimate to get advantage.
- Turn advantage into an optimality condition, e.g. .
- Train a generative policy conditioned on and , but drop sometimes so it learns both.
- Apply CFG during sampling:
Thus, the value function is just for labels and we don’t have to backpropagate through it during sampling.
Difference from advantage-weighted regression
AWR controls the magnitude of the training loss. Looks like REINFORCE, but here we fit logged actions with positive exponential weights. Standard REINFORCE uses current-policy samples and return / advantage weights. CFGRL uses supervised learning where the condition is different.
It learns an “optimality direction” and amplifies it at inference.
RECAP and FQL
RECAP: shares this advantage-conditioning idea with CFGRL. They also do policy iteration: repeatedly deploy, recollect, redo advantage, update policy, etc. They use a task-dependent advantage threshold and usually just sample the positive condition, rather than turning up CFG.
I would say the benefit of not just learning the optimal actions is that it is like more general knowledge: even things not directly to do with what we want can help make the model better, so that you can then, for example, generate better samples. Also can tune your regularisation, so better than just the KL-regularised fit. (My intuition.)
FQL: also regularises the improved policy towards a behaviour model. Here the penalty is squared distance between the actions they produce from the same noise input, rather than KL. The behaviour model is an iterative flow policy, and the learned policy generates an action in one step.