SPPO: Sequence-Level PPO for Long-Horizon Reasoning
Reformulates long-horizon reasoning as a sequence-level contextual bandit with a decoupled scalar value function to improve stability and reduce memory cost.
Work on reinforcement learning, post-training, and optimization. Use “Copy BibTeX” to copy a paper's citation in one click.