Shaohang Wei

魏少杭

Institute of Computational Linguistics (计算语言学研究所) · Peking University

5 Yiheyuan Rd, Haidian District, Beijing 100871, China

Shaohang Wei at Victoria Peak, Hong Kong
Taken at Victoria Peak, Hong Kong.

I am currently a Ph.D. student at the Institute of Computational Linguistics (计算语言学研究所), School of Computer Science, Peking University, advised by Prof. Houfeng Wang. My research focuses on 1) post-train: data-efficient and stable reinforcement learning across domains (RL, OPD, for agentic AI and reasoning); 2) data and eval: data synthesis research and reliable evaluation.

Before PKU, I received my B.Eng. degree in Artificial Intelligence from Beihang University in 2024.

I have industrial internship experience at Tencent (2026, 青云计划), Qwen Pilot, Alibaba · Tongyi Lab (2025), and PixVerse (2024), focusing on post-train, data synthesis and agentic ai & reasoning.

Industrial Experience

Tencent

· 青云计划

RL and OPD (multimodal agent): Working on post-training for visual agents: evidence diagnosis, trajectory reflection, targeted token-level supervision, and efficient tool use through ReVuE.

Qwen Pilot, Alibaba · Tongyi Lab

Algorithm Research: Worked on RLVR for mathematical reasoning and instruction following: token-level distribution shifts, cross-task learnability, and capability trade-offs. Infra & RL Framework: Built experience-guided rollouts in VeRL for sample efficiency and training stability.

PixVerse

Worked on data synthesis and evaluation for video generation: VLM-based video captioning, caption quality analysis, and side-by-side annotation and evaluation workflows.

News

Earlier updates
  • 3 papers were accepted by ICML 2026.
  • 1 paper was accepted by ACL 2026 Main.
  • 1 paper was accepted by ICLR 2026.
  • TIME was accepted as a NeurIPS 2025 Spotlight in the D&B Track.
  • 2 papers were accepted by ACL 2025 (Main and Findings).
  • 1 paper was accepted by Science China Information Sciences.
  • 1 paper was accepted by NAACL 2025 Main.

Open Source Projects & Academic Service

Open source. Co-contributor to BloodArena, an AI-powered Blood on the Clocktower environment for human–agent dialogue and LLM gameplay evaluation. Co-author of LeafyLingo, an interactive plant recognition system with a knowledge graph. Author of 24-Game-Reasoning, a minimal tutorial on LLM reasoning with Zero-RL, SFT, and SFT+RL.

Academic service. Reviewer for AAAI 2027, ICLR 2027, ICML 2026 (Gold Reviewer Award), and the EMNLP 2025 HCI-NLP Workshop.

Publications

16 publications and preprints · * Equal contribution

On-Policy Visual Evidence Distillation

Shaohang Wei, Feifan Song, Guangyue Peng, Wenhao Yu, Wei Li, Wen Luo, Yang Xu, Yufan Shen, Luke Mao, Yang Du, Asher Qin, Houfeng Wang

Preprint (under review)

TL;DR ReVuE diagnoses where visual agents lose evidence and turns those reflections into targeted on-policy distillation, improving visual reasoning while reducing redundant reasoning and tool use.

TimE: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios

Shaohang Wei, Wei Li, Feifan Song, Wen Luo, Tianyi Zhuang, Haochen Tan, Zhijiang Guo, Houfeng Wang

The 39th Annual Conference on Neural Information Processing Systems (NeurIPS 2025) · Spotlight (Datasets and Benchmarks Track)

TL;DR TimE evaluates temporal reasoning across Wikipedia, news, and dialogue, testing how models handle dense timelines, changing events, and time-dependent social interactions.

SWE-Ext: Scaling and Extending Augmented Data for Mid-Training of Repository-Level Coding Tasks

Wei Li, Xin Zhang, Shaohang Wei, Wen Luo, Feifan Song, Guangyue Peng, Yanjie Gao, Zhongxin Guo, Yangyu Huang, Houfeng Wang

The 42nd International Conference on Software Maintenance and Evolution (ICSME 2026)

TL;DR SWE-Ext turns GitHub pull requests into multilingual coding and completion data, improving repository-level mid-training and providing stronger initialization for downstream post-training.

Experience Augmented Policy Optimization for LLM Reasoning

Jinda Lu, Kexin Huang, Junkang Wu, Shuo Yang, Jinghan Li, Chiyu Ma, Shaohang Wei, Xiang Wang, Guoyin Wang, Jingren Zhou

The 43rd International Conference on Machine Learning (ICML 2026)

TL;DR EAPO reuses a trained policy's experience at critical token decisions, with importance correction to improve reasoning while reducing the inefficiency of learning from scratch.

Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning

Wenhao Yu, Shaohang Wei, Jiahong Liu, Yifan Li, Minda Hu, Aiwei Liu, Hao Zhang, Irwin King

The 43rd International Conference on Machine Learning (ICML 2026)

TL;DR RankTuner combines token probability and uncertainty to focus fine-tuning on under-learned tokens, improving reasoning and transfer without overemphasizing inherently uncertain targets.

One-Way Policy Optimization for Self-Evolving LLMs

Shuo Yang, Jinda Lu, Kexin Huang, Chiyu Ma, Shaohang Wei, Yuyang Liu, Guoyin Wang, Jingren Zhou, Li Yuan

The 43rd International Conference on Machine Learning (ICML 2026)

TL;DR OWPO separates verifier-guided update direction from reference-based update strength, preserving improvements beyond the reference policy and supporting continued self-improvement.

Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations

Wen Luo, Guangyue Peng, Wei Li, Shaohang Wei, Feifan Song, Liang Wang, Nan Yang, Xingxing Zhang, Jing Jin, Furu Wei, Houfeng Wang

The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026) · Main

TL;DR Truthfulness signals arise through distinct question- and answer-based pathways inside LLMs; disentangling them helps explain knowledge boundaries and improve hallucination detection.

Mitigating Overthinking through Reasoning Shaping

Feifan Song, Shaohang Wei, Bofei Gao, Yejie Wang, Wen Luo, Wei Li, Linli Yao, Weimin Xiong, Liang Chen, Tianyu Liu, Houfeng Wang

arXiv 2025

TL;DR GRSP regularizes reasoning at the step level, reducing excessive token use while retaining accuracy and improving reinforcement-learning stability.

MeNTi: Bridging Medical Calculator and LLM Agent with Nested Tool Calling

Yakun Zhu, Shaohang Wei, Xu Wang, Kui Xue, Shaoting Zhang, Xiaofan Zhang

The 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL 2025) · Main

TL;DR MeNTi uses nested tool calls to select medical calculators, fill inputs, and convert units, with CalcQA providing a physician-built benchmark for evaluating these abilities.

CiteCheck: Towards Accurate Citation Faithfulness Detection

Ziyao Xu, Shaohang Wei, Zhuoheng Han, Jing Jin, Zhe Yang, Xiaoguang Li, Haochen Tan, Zhijiang Guo, Houfeng Wang

arXiv 2025

TL;DR CiteCheck combines cost-conscious annotation with synthetic negative examples to build Chinese citation-faithfulness data and train smaller, effective evaluators for retrieval-augmented generation.