
skull of a skeleton with burning cigarette
vincent van gogh
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
This is a writeup and presentation of the work done by Zhang et al on Optimistic Nash Policy Optimization (ONPO) a novel RLHF algorithm.