Contents

Verifier-Induced Support Reshaping
in On-Policy Optimization

Shaohang Wei1,‡Zikun Su2Feifan Song1Wen Luo1Wei Li1Guangyue Peng1Houfeng Wang1,†

1Peking University2BUPT

‡Project Lead†Corresponding Author

Overview of verifier-induced support reshaping: a finite rollout budget connects the current verifier to future training support.

What remains learnable next? A verifier changes which successful responses future on-policy training can still sample and reinforce.Enlarge ↗PDF ↗SVG ↗GIF ↓

Abstract

On-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives harder to sample. We study this verifier-induced support reshaping through mathematical reasoning and constrained instruction following across two model families. Math-RLVR raises average instruction-following success but leaves fewer prompts solvable under repeated sampling. IF-RLVR reduces math searchability and the reward variation available to later training. Distributional analyses and controlled interventions locate a key change at the response opening. Reference-policy constraints, routing priors, and on-policy distillation preserve cross-task support only partially. Tests on MathIF and ReasonIF further show that separate gains only partly translate into answers that are both correct and constraint-following.

TL;DR. A better score today does not guarantee that tomorrow’s successful responses remain reachable.

The question behind the score

With a fixed rollout budget, a successful response must remain likely enough to be sampled before an on-policy update can reinforce it. We call these reachable successes effective rewardable support.

We train Qwen3-8B-Base and Qwen2.5-Math-7B with math or instruction-following (IF) verifiers, then test both directions and both training orders. pass@1 measures average single-rollout success; best@k measures the share of prompts with at least one success in k samples.

Higher scores can hide narrower support

The two verifiers reshape cross-task behavior differently. Math-RLVR increases average IF success while reducing prompt coverage; IF-RLVR makes correct math responses harder to find even with more samples.

On IFEval with Qwen3-8B-Base, pass@1 rises by 6.5% while best@32 falls by 9.8% (absolute changes). The same split appears across both models and both IF benchmarks: more prompts become consistently correct, but more also become consistently wrong.

Math-RLVR increases IF pass@1 while best@32 decreases; prompts move toward all-correct and all-wrong outcomes.
Math-RLVR. Average success, repeated-sampling coverage, and prompt-level outcome shifts.PDF ↗SVG ↗

Training order shapes future learning

Switching verifiers exposes the consequence of narrower support. IF-first training leaves little reward variation for later math updates; in the reverse order, retaining math performance limits IF adaptation.

After the switch to Math-RLVR, the group of prompts with both successful and failed rollouts shrinks from 33.6% to 1.6% within 20 steps. Scores improve, but almost all groups become uniformly right or wrong, leaving little within-group learning signal.

After IF-first training, the mixed-success prompt group falls from 33.6% to 1.6% within 20 steps of Math-RLVR.
Sequential training. Better outcomes do not restore mixed support or step-by-step openings along this observed path.PDF ↗SVG ↗

Response openings change math searchability

The largest next-token distribution shift occurs at the first response token. RLVR reweights existing token candidates and can promote low-probability openings. Controlled interventions then test whether choosing a different opening changes which math solutions can be reached.

Without updating model parameters, forcing Base-side tokens or step-by-step (DRI) prefixes increases math best@32 from IF-RLVR checkpoints in both families. The reverse intervention is asymmetric across models, so the result supports an opening-route effect within the tested settings.

Forcing Base-side or step-by-step openings improves math best@32 from IF-RLVR checkpoints in both model families.
Controlled intervention. Free decoding versus forced first tokens and prefixes on AIME and MATH-500.PDF ↗SVG ↗

Preserving support takes more than a prior

The opening effect suggests two interventions: initialize a reasoning-friendly route, or provide dense teacher supervision through on-policy distillation (OPD). A one-time prior only delays the shift, and OPD outcomes depend on which teacher state is used.

Four SFT cold starts use 50 correct math responses before the same 100-step IF-RLVR protocol. Soft DRI delays the direct-answer shift longest, but all four conditions reach 100% direct-answer (DAI) openings by step 100. A stronger prior does not last longer in this run.

Four cold starts give different opening trajectories; all reach 100% DAI at step 100.
Cold starts. The prior changes when the shift occurs, not the observed endpoint route.PDF ↗SVG ↗

Separate gains only partly combine

Can one response be both correct and instruction-following? An independent endpoint test on MathIF and ReasonIF evaluates correctness (C), strict following (F), and their same-response conjunction (J), using 16 rollouts for each of 290 prompts per benchmark.

On MathIF, Math-RLVR raises correctness from 35.9% to 47.2%, but joint success moves only from 9.4% to 11.8%. IF-RLVR raises strict following from 25.0% to 43.1%, while joint success reaches 14.8%. ReasonIF shows the same gap between target-aligned gains and joint gains.

Correctness, strict instruction following, and joint success for Base, Math-RLVR, and IF-RLVR on MathIF and ReasonIF.
Joint-support stress test. Values from the paper’s Table 3; these independent step-720 endpoints are separate from the controlled trajectories above.PDF ↗SVG ↗

The implication. Evaluate more than the endpoint score: inspect reachable successes, the learning signal available after switching objectives, and whether desired behaviors coexist in the same response.

Scope: two starting model families and fixed sampling budgets, with one fixed-seed run per configuration. The findings characterize the tested training paths; they do not establish irreversible loss or a universal preservation strategy. Full settings and additional analyses are in the technical report.

Citation


@misc{wei2026verifier,
  title         = {Verifier-Induced Support Reshaping in On-Policy Optimization},
  author        = {Shaohang Wei and Zikun Su and Feifan Song and Wen Luo and Wei Li and Guangyue Peng and Houfeng Wang},
  year          = {2026},
  eprint        = {2608.00220},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  doi           = {10.48550/arXiv.2608.00220},
  url           = {https://arxiv.org/abs/2608.00220}
}

Paper figure

Scroll to inspect · Escape to close