skull of a skeleton with burning cigarette
vincent van gogh

Improving LLM General Preference Alignment via Optimistic Online Mirror Descent


This is a writeup and presentation of the work done by Zhang et al on Optimistic Nash Policy Optimization (ONPO) a novel RLHF algorithm.

Link to the original paper.

Open the Slides in a separate tab

Full Report