Tianyi Wang

Undergraduate Researcher · BUPT

I work on post-training algorithms for large language models, with a focus on reinforcement learning and reliable optimization.

My current research explores stable and efficient learning for long-horizon reasoning. I am also interested in optimizer design, implementation, and reproducible machine learning systems.

Updates

News

  • SPPO was selected for an oral presentation at ACL 2026.
  • SPPO was accepted to ACL 2026.
  • APO was accepted to ICML 2026.
Research

Selected Works

All publications →
Preview of SPPO: Sequence-Level PPO for Long-Horizon Reasoning
ACL 2026 Oral Co-first Author

SPPO: Sequence-Level PPO for Long-Horizon Reasoning

Tianyi Wang*, Yixia Li*, Long Li, Yibiao Chen, Shaohan Huang, Yun Chen, Peng Li, Yang Liu, Guanhua Chen
Reformulates long-horizon reasoning as a sequence-level contextual bandit with a decoupled scalar value function to improve stability and reduce memory cost.
Paper Code
Preview of DyJR: Preserving Diversity in Reinforcement Learning with Verifiable Rewards via Dynamic Jensen-Shannon Replay
Under Review

DyJR: Preserving Diversity in Reinforcement Learning with Verifiable Rewards via Dynamic Jensen-Shannon Replay

Long Li, Zhijian Zhou, Tianyi Wang, Weidi Xu, Zuming Huang, Wei Chu, Zhe Wang, Shirui Pan, Chao Qu, Yuan Qi
Introduces Dynamic Jensen-Shannon Replay to preserve diversity in reinforcement learning through verifiable rewards mechanism.
Paper
Preview of Anchored Policy Optimization: Mitigating Exploration Collapse via Support-Constrained Rectification
ICML 2026 Co-first Author

Anchored Policy Optimization: Mitigating Exploration Collapse via Support-Constrained Rectification

Tianyi Wang*, Long Li*, Hongcan Guo, Yibiao Chen, Yixia Li, Yong Wang, Yun Chen, Guanhua Chen
Introduces support-constrained rectification for RLVR to mitigate exploration collapse and improve both Pass@1 and response diversity.
Paper Code
Background

Education & Experience

Education

BUPT

2023–2027

Experience

Youtu

Research Intern · 2026–Present