ACL 2026 Main
Improving reasoning in diffusion language models
We improve reasoning in diffusion LLMs by adding process rewards to GRPO-based reinforcement learning (RL) during post-training. These rewards compare answer success rates before and after a denoising interval, giving the model credit for progress toward a correct answer.
No human-labeled reasoning steps or separate reward model required.
