β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

Published in arXiv, 2026