Signal2026-08-06
arXiv

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

Part of

Advanced Inverse Reward And Prompt Engineering