Tianyi Wang
I work on post-training algorithms for large language models, with a focus on reinforcement learning and reliable optimization.
My current research explores stable and efficient learning for long-horizon reasoning. I am also interested in optimizer design, implementation, and reproducible machine learning systems.
Updates
News
- SPPO was selected for an oral presentation at ACL 2026.
- SPPO was accepted to ACL 2026.
- APO was accepted to ICML 2026.
Research
All publications →Selected Works
SPPO: Sequence-Level PPO for Long-Horizon Reasoning
Reformulates long-horizon reasoning as a sequence-level contextual bandit with a decoupled scalar value function to improve stability and reduce memory cost.
DyJR: Preserving Diversity in Reinforcement Learning with Verifiable Rewards via Dynamic Jensen-Shannon Replay
Introduces Dynamic Jensen-Shannon Replay to preserve diversity in reinforcement learning through verifiable rewards mechanism.
Anchored Policy Optimization: Mitigating Exploration Collapse via Support-Constrained Rectification
Introduces support-constrained rectification for RLVR to mitigate exploration collapse and improve both Pass@1 and response diversity.
Background
Education & Experience
Education
BUPT
2023–2027

Experience
Youtu
Research Intern · 2026–Present
