On-Policy Distillation: A Unified Distributional View of Post-Training
📅 Published:
📘 TABLE OF CONTENTS
- 1. A Unified Distributional View of Post-Training
- 2. Supervised Fine-Tuning: Projection Toward a Demonstration Distribution
- 3. Reinforcement Learning: Reward-Induced Distribution Tilting
- 4. Direct Preference Optimization: Distribution Shaping Without Explicit RL
- 5. On-Policy Distillation: Teacher-Guided Shaping on Student States
- 6. Beyond Basic OPD: Making Teacher Supervision Useful
- 6.1 Teacher Compatibility: When a Stronger Model Is Not a Better Teacher
- 6.2 Adaptive Objectives: Learning Precisely Without Suppressing Alternatives
- 6.3 Outcome Alignment: Combining Teacher Guidance with Evidence of Success
- 6.4 Constructing Teachers: Privileged Information, Self-Distillation, and Weak-to-Strong Transfer
- 6.5 Capability Composition: Multiple Specialists and Multimodal Students
- 6.6 State and Compute Efficiency: Learning More from Fewer Interactions
- 6.7 Semantic and Temporal Credit: Supervising the Behavior That Matters
- 6.8 Connecting the Research Directions to the Unified Framework
- References
Modern post-training methods for large generative models are often described as fundamentally different paradigms: supervised fine-tuning (SFT) learns from demonstrations, reinforcement learning (RL) optimizes rewards, Direct Preference Optimization (DPO) learns from preference pairs, Reinforcement Learning from AI Feedback (RLAIF) uses model-generated feedback to guide RL, and On-Policy Distillation (OPD) transfers teacher behavior on student-generated trajectories. Despite their different implementations, these methods can be understood through a common probabilistic lens: they all reshape the conditional distribution induced by a pretrained base model. This article first develops that shared perspective, then examines what basic OPD leaves unresolved and how recent research changes its teachers, supervision, and training states.
RLAIF identifies the source of feedback within an RL pipeline; Constitutional AI is an early example.1 More broadly, AI-generated preference labels can also train offline methods such as DPO. The four approaches examined below are therefore SFT, RL, DPO, and OPD. They are useful categories for comparison, but they overlap: OPD can use policy gradients, and a training pipeline can combine several of them.
1. A Unified Distributional View of Post-Training
Consider a pretrained generative model $\pi_0(y\mid x)$, where $x\sim\rho(x)$ is an input condition and $y$ is an output. For a language model, $x$ is a prompt and
is a response, including its end-of-sequence (EOS) token. Its probability factorizes as
Here $\pi_\theta$ is the trainable policy, and $\theta_0$ denotes its initialization, so that $\pi_{\theta_0}=\pi_0$. At token position $t$, the state is the prompt and preceding tokens, $s_t=(x,y_{<t})$, and the action $a$ is the next token. All logarithms below are natural logarithms.
Post-training changes this conditional distribution to make desired outputs more likely. Probability-mass reallocation is a useful description of the resulting behavior, not a complete account of how the model’s internal representations change. It also does not imply that the model can only memorize or select outputs already present in its training data.
This viewpoint extends conceptually to image and video generation, but the probability space must be specified. Autoregressive visual models can use token sequences; diffusion models may use a final-sample distribution or a distribution over stochastic denoising paths. These are different objects. A deterministic flow may induce an output density through its random initial condition while its path measure is singular. The sequence likelihoods and KL identities below therefore assume discrete autoregressive responses; they cannot simply be reused as diffusion or flow training losses.
We will organize the discussion around four choices:
| Component | Meaning | Examples |
|---|---|---|
| $\mu(s)$ | Distribution of states receiving supervision | Demonstration prefixes, fresh student prefixes, replayed histories |
| $I(s)$ | Evidence used to evaluate behavior | Demonstrations, rewards, preference labels, teacher probabilities, verifier or environment feedback |
| $q^\star(\cdot\mid s)$ | Target distribution, when one can be defined | Demonstration conditionals, a reward-tilted reference policy, a teacher’s next-token distribution |
| $\mathcal U$ | Learning procedure | Maximum likelihood, divergence minimization, policy gradients, pairwise preference fitting |
Conceptually, a training procedure produces
This is bookkeeping for the learning process, not an optimization theorem. The states and targets can change with the policy. A finite preference dataset, for example, need not determine a unique $q^\star$ over all responses.
The first component determines where learning happens. Offline training uses a fixed data-induced state distribution; strictly on-policy training collects states from the policy being updated:
Here $d_\pi$ denotes a normalized visitation distribution over prefixes under a specified prompt distribution and token-weighting convention. The explicit trajectory losses below use sums over tokens; dividing each sequence by its own length instead would define a different weighting and generally a different objective.
Mixed sampling and replay occupy intermediate positions. On-policy refers to the relationship between the sampling policy and the update, not to training in real time after deployment. It addresses prefix mismatch only to the extent that training prompts, decoding rules, and interaction conditions represent deployment.
The remaining choices specify what evidence guides the model, what target that evidence supports, and how the update is computed. Keeping these distinctions separate lets us compare methods without assuming that they optimize the same objective.
2. Supervised Fine-Tuning: Projection Toward a Demonstration Distribution
Let a supervised dataset be sampled from
where $p_{\text{data}}$ denotes the demonstration distribution. In instruction tuning, these are prompt–response demonstrations, as in the supervised stage of InstructGPT.2 SFT minimizes the negative log-likelihood
For each input, cross-entropy decomposes into the entropy of the fixed demonstration distribution and a forward KL:
For discrete responses and finite cross-entropy, the entropy term is independent of $\theta$. Consequently, the minimizers, rather than the numerical values of the two objectives, are identical:
Thus SFT can be interpreted as projecting the model toward the demonstration distribution. The distributional picture is
At the distributional optimum, the model matches the demonstrated conditional distribution as closely as its model family allows. In a finite run, shared parameters couple different examples, so individual output probabilities need not change monotonically.
Each training token provides a direct supervised target. This dense signal makes SFT relatively straightforward to optimize. Compared with online rollout-based training, SFT often offers stable optimization, relatively low-variance gradients, and lower training cost once demonstrations are available. Those benefits still depend on data quality and the optimization setup.
However, teacher forcing trains on demonstrator prefixes, whereas inference uses the student’s own preceding predictions. Errors can therefore lead to unfamiliar states. This exposure-bias problem is a central motivation for learning on learner-visited states in imitation learning.3 SFT also depends on demonstration quality and coverage: its likelihood objective does not explicitly reward an equally valid alternative that the dataset omits. Generalization can still produce such alternatives, but the objective itself fits demonstrations rather than directly optimizing task success.
3. Reinforcement Learning: Reward-Induced Distribution Tilting
In reward optimization with PPO or GRPO,45 demonstrations do not directly specify the response target for the RL loss, although the pipeline may start from SFT or mix in a supervised objective. Instead, a reward function
evaluates the model’s own generated outputs. The basic objective is
A common language-model post-training objective adds a KL penalty relative to a fixed reference policy $\pi_{\mathrm{ref}}$, often an SFT checkpoint rather than the pretrained model $\pi_0$.2 For $\beta>0$,
Assume a fixed reward, a finite normalizing constant, and policies whose support is contained in the reference policy’s support. With unrestricted optimization over those response distributions, the optimum is
where
For a fixed prompt, this follows from
where the policies in the KL terms are conditioned on $x$. Thus the maximum is attained at $\pi=\pi^\star$ whenever this distribution is allowed. If the reference assigns zero mass to a response, finite-KL optimization cannot introduce mass there.
This characterizes the KL-regularized objective, not every RL algorithm or a guarantee that PPO/GRPO training reaches the optimum. PPO’s clipping against a recent rollout policy and a KL penalty against a fixed reference serve different purposes. The distributional interpretation here is: KL-regularized reward optimization exponentially reweights the reference distribution according to reward. Outputs with larger rewards receive larger multiplicative weights before normalization, increasing their relative odds against lower-reward alternatives. Unlike SFT, it does not have to correspond to one demonstrated answer. If many outputs have high reward, all may remain plausible.
On-policy RL reduces the teacher-forcing mismatch of SFT by collecting behavior from the current or a recent policy. PPO and GRPO commonly reuse each rollout batch for several updates, so clipping or related controls manage the resulting policy drift. This description concerns these on-policy approaches; RL also includes off-policy algorithms.
When feedback is provided only for the final outcome, RL has a much sparser signal than token-level supervision. Such settings face problems of sparse reward, credit assignment, exploration, high gradient variance, reward misspecification, reward hacking, and substantial rollout cost.
4. Direct Preference Optimization: Distribution Shaping Without Explicit RL
Preference learning assumes data of the form
where
A common latent-reward model assumes
where $\sigma(z)=1/(1+e^{-z})$ is the logistic sigmoid. This is the Bradley–Terry preference model, with its noise scale absorbed into $R$. DPO6 removes the explicit reward-model and policy-gradient stages by exploiting the relationship between the optimal KL-regularized RL policy and the latent reward. From
we obtain
Here $C(x)=\beta\log Z(x)$ cancels when comparing two responses to the same prompt. Replacing the optimal policy by the trainable policy and fitting observed labels yields the DPO objective
From a distributional perspective, DPO says: increase the reference-adjusted log-probability margin of the preferred output over the rejected output. This is a statement about relative odds: it does not guarantee that the absolute probability of every preferred response increases, or that every rejected response decreases, after a shared-parameter update.
Thus DPO is closer to RL than its supervised-looking loss may initially suggest. It implements a form of preference-induced distribution shaping corresponding to an implicit reward.
The key difference from on-policy RL is the source of the training comparisons: $(x,y^+,y^-) \sim D_{\mathrm{pref}}$ is usually sampled from a fixed, offline preference dataset. The original DPO setup is therefore an offline preference optimization method. Refreshing preference pairs with the current policy creates an online variant; the algebraic form of the loss alone does not determine whether collection is offline.
For a fixed preference dataset, DPO avoids a separately fitted reward model and an online rollout loop during optimization. It still needs response pairs and labels, whose collection may be expensive. Its gradients come from the pairwise classification loss rather than a policy-gradient estimator.
The connection to KL-regularized RL relies on the preference model and sufficient data coverage; an offline DPO run need not recover the same policy as on-policy reward optimization. Its main limitations follow from this offline setting. If the current policy moves into regions that are poorly represented in the preference dataset, DPO receives no direct correction there. Its behavior is therefore highly dependent on preference-data coverage and on the relationship between the policy being optimized and the policy that generated the preference pairs.
5. On-Policy Distillation: Teacher-Guided Shaping on Student States
Classical knowledge distillation trains a student to match a teacher’s soft predictions on supplied inputs.7 For autoregressive models, token-level distillation compares next-token distributions on chosen prefixes, while sequence-level distillation trains on teacher-produced responses.8 In a fixed-data setup, these prefixes are teacher-generated or externally supplied. Let
be a teacher policy and
the student. A generic token-level offline distillation objective is
Here $D$ is a chosen discrepancy; for forward KL it is $D(\pi_T,\pi_\theta)=D_{\mathrm{KL}}(\pi_T|\pi_\theta)$. Comparing these distributions directly assumes a common action space, usually shared tokenization and compatible special-token conventions. Full-vocabulary KL also requires access to teacher probabilities; a text-only API does not provide this information.
The idea of querying an expert on learner-visited states predates modern LLM distillation. DAgger iteratively collects such expert labels and aggregates them with past data.3 It provides a useful conceptual antecedent, but its expert assumptions and guarantees do not automatically apply to an LLM teacher on an arbitrary erroneous prefix.
On-Policy Distillation changes a crucial component: the student generates the prefixes on which the teacher is queried. At iteration $k$, let $\pi_{\theta_k}$ be the rollout policy and $s_t=(x,y_{<t})$ a visited prefix. A common direct distillation update minimizes
This is a local surrogate loss for one rollout batch. The batch is held fixed during differentiation, while gradients flow through the student’s next-token distribution and the teacher is held fixed. Reusing the batch after the first update can make its sampling policy stale; the method must refresh samples or explicitly manage that drift. Forward KL and other divergences are also possible: on-policy specifies where supervision is obtained, not which divergence must be used. The teacher can score the student’s sequence in a teacher-forced forward pass; it need not generate an entirely corrected response.
Earlier work already developed student-aware sequence distillation: ImitKD connected autoregressive distillation to imitation learning and learner-generated prefixes,9 while f-DISTILL studied sequence-level distillation through alternative f-divergences.10
Two influential LLM approaches clarify the distinction between local and sequence-level optimization. Generalized Knowledge Distillation (GKD) from Google DeepMind and collaborators studies student-generated training sequences and alternative divergence objectives. MiniLLM, from Tsinghua University and Microsoft Research, develops reverse-KL distillation using a policy-gradient formulation.1112 Both use student behavior, but their optimization details should not be treated as interchangeable. MiniLLM’s practical algorithm also uses teacher-mixed sampling, approximate importance weighting, and length normalization; it is not simply the unmodified gradient identity shown in Section 6.7.
For a fixed teacher and a shared tokenization, the sequence-level reverse KL obeys the chain rule
when responses include termination, both policies use the same stopping convention, the student is absolutely continuous with respect to the teacher, and the relevant expectations are finite. A fixed maximum length is also valid if both policies use the same truncation convention. However, differentiating this entire expectation also differentiates the distribution of visited prefixes. A local update that holds sampled prefixes fixed omits that dependence. Likewise, a sampled-token log-ratio update is not automatically the full sequence-level policy gradient; temporal return terms matter. For clarity, at a fixed prefix $s$ and with a fixed teacher,
A sampled-token coefficient can therefore estimate the local reverse-KL update at the sampling policy. Holding a sampled token fixed and differentiating its bare log-ratio is not this estimator: the coefficient must weight the student’s log-probability in a surrogate, with the coefficient detached. Section 6.7 explains why optimizing the full sequence distribution additionally requires credit for later states.
OPD combines student-generated training states and dense teacher supervision. It can therefore teach the student how to continue from its own imperfect reasoning, while extracting more local information than a single terminal reward. It can also consolidate specialist capabilities, as illustrated by NVIDIA’s multi-domain distillation pipeline.13
The comparison can now be summarized without conflating state collection with the loss:
| Approach | Typical training states | Main supervision | Objective or target |
|---|---|---|---|
| SFT | Demonstration prefixes | Observed target tokens | Maximum likelihood / forward KL to data |
| On-policy RL | Current or recent student rollouts | Task reward, often with a reference penalty | Expected reward; a Gibbs target for the fixed-reward KL-regularized idealization |
| Offline DPO | Fixed preference pairs and their prefixes | Pairwise preference labels | Reference-adjusted likelihood-ratio margin |
| OPD | Current or recent student prefixes | Teacher next-token probabilities | A chosen divergence or distillation policy-gradient objective |
These advantages do not establish that every teacher signal is useful. The teacher may be unreliable on student-generated prefixes; strict imitation may suppress valid alternatives; matching teacher probabilities may conflict with task success; and obtaining fresh trajectories and teacher scores can be expensive. Pure matching to a fixed teacher also provides no general mechanism for improving beyond that target, although finite-capacity approximation and evaluation differences mean that the teacher’s benchmark score is not a universal mathematical ceiling. These limitations motivate the research directions below.
6. Beyond Basic OPD: Making Teacher Supervision Useful
Basic OPD determines where learning happens: the student visits a state, and the teacher supplies a local target. What remains unresolved is whether that target is reliable, learnable, and aligned with the eventual task. A teacher can be strong in isolation yet unhelpful on a student’s mistaken prefix. A dense signal can be precise about token probabilities yet wrong about which reasoning path should be encouraged. An inexpensive gradient update can still require an expensive trajectory to obtain it.
The four-part framework in Section 1 provides a useful way to organize these problems. Current work changes the training-state distribution $\mu$, the supervision signal $I$, the effective teacher target $q^\star$, or the update rule $\mathcal U$. These choices interact: changing a teacher also changes which states it can supervise, while changing the loss changes which parts of its knowledge the student absorbs. The following subsections group papers by the underlying problem rather than by small variations in algorithm design.

6.1 Teacher Compatibility: When a Stronger Model Is Not a Better Teacher
On-policy sampling reduces the gap between the student’s training prefixes and its inference prefixes. It does not eliminate the gap between those prefixes and the teacher’s familiar states. Consider a student that has already made an invalid mathematical assumption. A capable teacher may solve the original question correctly, yet its next-token distribution after that assumption may continue the flawed argument rather than repair it. Teacher quality must therefore be evaluated conditionally on the states the student actually visits.

Rethinking On-Policy Distillation from Tsinghua’s THUNLP group and collaborators examines this issue through controlled teacher–student comparisons.14 Its experiments emphasize compatible thinking patterns and transferable capabilities, rather than model size or benchmark score alone. The proposed remedies include an off-policy cold start using teacher trajectories and selecting prompts better aligned with the teacher’s post-training experience. In distributional terms, the cold start helps the student reach states where useful teacher guidance becomes accessible. These are empirical findings within the paper’s settings, not universal necessary and sufficient conditions for distillation.
Apple’s Unmasking On-Policy Distillation approaches the same question through diagnostics.15 It estimates a per-node update direction that would increase the student’s probability of eventual success, then compares a candidate distillation gradient with that direction. This allows teacher and context choices to be assessed at the level of individual questions and tokens. The study finds that guidance can be more useful on incorrect rollouts than on already-correct ones, where additional imitation may introduce noise. Together, these papers shift teacher selection from a global ranking problem toward a conditional usefulness problem.

An unresolved challenge is to estimate that usefulness cheaply. A practical teacher-selection rule should account for complementary knowledge, compatibility with the student’s reasoning, and the likelihood that a local correction improves the final outcome. Teacher confidence alone cannot establish all three: a confident distribution may encode an unsuitable continuation, while an uncertain distribution may correctly represent several valid alternatives.
6.2 Adaptive Objectives: Learning Precisely Without Suppressing Alternatives
A fixed divergence applies the same notion of disagreement everywhere. Reasoning does not have that uniform structure. Some positions admit a nearly determined continuation; others begin a choice between several valid solution strategies. Under capacity constraints, reverse KL can favor concentration on a subset of teacher-supported behavior. That can help produce coherent responses, but excessive concentration can also reduce exploration and the benefit of sampling multiple attempts. This is an optimization tendency, not a claim that exact reverse-KL minimization must collapse a perfectly representable teacher distribution.
The DistiLLM line from KAIST and Microsoft studies the interaction between objectives and data sources.1617 DistiLLM uses skew-KL objectives and adaptive reuse of student-generated outputs to improve optimization and sampling efficiency. DistiLLM-2 develops a contrastive treatment of teacher- and student-generated responses, rather than applying an identical loss to both. The shared lesson is that the appropriate update depends on where a sample comes from and what it is intended to teach. Because these methods mix or reuse data sources, they belong to the broader development of student-aware distillation rather than a uniformly strict on-policy recipe.
Entropy-Aware On-Policy Distillation (EOPD) from IBM Research and collaborators adapts the objective to teacher uncertainty.18 When the teacher distribution has high entropy, it augments reverse KL with forward KL, preserving a broader set of plausible continuations. At more decisive positions, the objective retains the benefits of precise imitation. REOPOLD from Microsoft addresses instability in sampled-token feedback: many tokens contribute almost no information, while a few extreme negative log-ratios can dominate an update.19 Its clipping floor is motivated by a teacher–student mixture, while token selection in the refinement phase uses student entropy. An exploration-to-refinement schedule changes which feedback is retained over time. This differs from EOPD’s use of teacher entropy to adapt the divergence.
These approaches suggest a common design principle: distillation strength should depend on the meaning of the disagreement. A useful next step is to distinguish uncertainty, genuine errors, and harmless stylistic differences, rather than treating every large divergence as valuable supervision. Evaluation should also measure both single-attempt accuracy and performance under multiple samples; improving one while destroying diversity can conceal an important regression.
6.3 Outcome Alignment: Combining Teacher Guidance with Evidence of Success
Distribution matching and task optimization can disagree. A student may produce a correct solution in a style the teacher assigns low probability, or an incorrect solution whose local wording remains plausible to the teacher. In sampled-token OPD, a common feedback coefficient is
This compares the probabilities assigned by teacher and student to the sampled action; neither probability is necessarily a calibrated estimate of correctness. It is not, by itself, that action’s causal contribution to task success. Dense supervision therefore does not automatically solve credit assignment.
GKD already investigated combining on-policy distillation with sequence-level reward optimization.11 The same student rollout can support both signals: a teacher supplies local distributional information, while a reward evaluates the completed behavior. More recent work asks how to handle disagreement between them. On-policy Distillation with Verifiable Reward (OPDVR) from Tsinghua’s LeapLab and collaborators uses the verifier to gate token-level feedback.20 With a correctness sign $v(y)\in{-1,+1}$, its central operation can be expressed as
Correct trajectories retain nonnegative distillation feedback; incorrect trajectories retain nonpositive feedback. Updates that conflict with this outcome-level direction are removed. The teacher–student probability ratio determines the magnitude of the retained coefficient, while the verifier determines its permitted sign. In the corresponding sampled-token surrogate, this coefficient is detached from the gradient. This is the base OPDVR gate; the paper’s GRPD extension adds group-relative reward normalization.

A complementary approach improves the teacher’s information before forming the update. Multi-Rollout On-Policy Distillation via Peer Successes and Failures, from a Microsoft Research collaboration, conditions teacher supervision on other attempts at the same problem.21 Successful peers provide examples of viable reasoning, while failed peers reveal plausible mistakes. Its contrastive success–failure context makes supervision specific to the student’s observed alternatives. This is distinct from multi-teacher distillation, despite both sometimes being abbreviated MOPD.
The broader direction is to combine local guidance with outcome evidence. However, a terminal verifier does not identify every erroneous step: an incorrect trajectory can contain many useful decisions. Outcome gating is therefore a way to manage conflicting signals, not a complete solution to temporal credit assignment. Stronger methods should preserve correct intermediate work while targeting the decisions responsible for failure.
6.4 Constructing Teachers: Privileged Information, Self-Distillation, and Weak-to-Strong Transfer
A fixed external teacher creates both a cost problem and a capability problem. The strongest available student may have no stronger general-purpose model to imitate. Even when a teacher exists, reproducing its full behavior may be unnecessary: the student may only need a particular skill or an informative explanation of its own mistake. This motivates constructing a useful teacher from additional context or complementary capability, rather than requiring a larger model in every setting.
Privileged-context methods instantiate a teacher with additional information unavailable to the student at the decision being supervised:
where $c$ might be a demonstration, a verified solution, or feedback about the current attempt that was unavailable when the student generated it. For an interactive agent, earlier environment observations already belong to $s_t$; they need not disappear at deployment. Teacher targets are treated as fixed during each student update; methods differ in whether the teacher weights themselves are frozen, copied, or refreshed over training.
Self-Distilled Reasoner (OPSD), from Meta Superintelligence Labs, UCLA, and HKU, uses solution-informed predictions to supervise the model without that solution.22 Self-Distillation Fine-Tuning (SDFT), from MIT and ETH Zurich, uses demonstration-conditioned self-teachers to acquire skills while reducing forgetting.23 Self-Distillation Policy Optimization (SDPO), from ETH Zurich, MPI, MIT, and Stanford, uses rich environment feedback, such as execution errors, to form a feedback-conditioned teacher.24 These methods share a mechanism: information that is useful in context is converted into a persistent change in the unassisted policy. Their supervision is not information-free; the demonstrations, solutions, or feedback are essential resources.

Weak-to-Strong OPD (W2S-OPD), from Microsoft Research and collaborators, constructs a different kind of teacher.25 It extracts a capability direction from the logit difference between positive and negative weak models, then combines that direction with the student’s base model to create a proxy teacher. A small model before and after RL, for example, can reveal a skill improvement without requiring the strong student to imitate the small model’s entire distribution. The aim is selective transfer while maintaining compatibility with the student’s existing behavior.

These strategies introduce a new mismatch: the teacher may act as if it knows information the student must discover. Princeton’s Rethinking On-Policy Self-Distillation for Thinking Models finds that privileged-context supervision can degrade long-budget reasoning and suppress verification or backtracking in the studied thinking models.26 A teacher that already knows the answer may discourage the exploratory moves an uninformed student needs. Designing privileged context therefore requires more than maximizing teacher accuracy: it must preserve the process by which the student can reach the answer from its actual information state.

6.5 Capability Composition: Multiple Specialists and Multimodal Students
A single teacher need not dominate every domain. Coding, mathematics, instruction following, and multimodal reasoning can benefit from different experts. Sequentially training one student on these objectives can also improve one capability while degrading another. OPD offers a way to consolidate expertise on the student’s own responses, but doing so requires deciding which teacher is relevant and how much influence it should have.
NVIDIA’s Nemotron-Cascade 2 uses multi-domain OPD within a cascaded RL pipeline.13 Intermediate checkpoints that perform best in particular domains become teachers for subsequent consolidation. Distillation helps recover capabilities that regress during later specialization. This makes the teacher pool a reusable record of successful learning, rather than requiring a single larger model that is best at everything. It also illustrates how RL and OPD can play complementary roles: one discovers improved behaviors, while the other transfers and consolidates them.

On-Policy Omni Distillation (OPOD), a Tencent-involved collaboration, extends specialist composition across text, image, and audio tasks.27 It routes student responses to modality-appropriate teachers, controls their influence separately, and combines one-sided token guidance with teacher-derived verification. The motivation is that pooled multimodal training and uniform teacher pressure can cause interference. This verification is teacher-derived and should not be equated with an independent correctness oracle. The specialist models can be discarded after training, leaving one deployable student.
A different route appears in VOLD, from Tübingen, the MIT-IBM Watson AI Lab, and Inria/ENS collaborators.28 It transfers reasoning from a text-only LLM to a vision-language student through an SFT cold start followed by joint GRPO and OPD on text-only reasoning tasks. The resulting model is evaluated on visual reasoning. This separates improving a VLM’s reasoning policy from supplying new visual training data, and demonstrates why initial teacher–student alignment matters for transfer.

These works motivate conditional capability composition: choose guidance according to domain, modality, and the student’s remaining weaknesses. They do not establish that averaging expert distributions is sufficient, nor that understanding-oriented multimodal OPD directly solves image or video generation. Extending the idea to continuous generative trajectories requires defining compatible states, teacher signals, and update estimators for those processes.
6.6 State and Compute Efficiency: Learning More from Fewer Interactions
OPD’s dense gradients can obscure the cost of producing them. Each update may require fresh student generation, teacher scoring, and—in agent settings—live tool or environment interaction. Training efficiency should therefore be measured in teacher computation, generated tokens, environment calls, and wall-clock time, not only optimization steps.
The replay component of DistiLLM addresses this trade-off by reusing student-generated outputs.16 Replayed-Prefix On-Policy Distillation (ReOPD), from Microsoft Research and the University of Amsterdam, targets the environment cost of multi-turn agents.29 It replays previously collected teacher prefixes and lets the student generate at selected steps, receiving teacher supervision without executing new environment actions during student training. Collecting the reusable teacher traces still incurs environment cost. Its sampling schedule favors earlier prefixes to manage the mismatch between student behavior and teacher reliability. The selected continuation is student-generated, but the entire interaction history is not generated by the current student; this is an intentional hybrid rather than fully on-policy trajectory collection.
Data selection raises a related question. In OPD, a prompt primarily acts as a generator of training states: the teacher supplies the targets after those states are visited. Rethinking OPD II: One Training Example, from Tsinghua and collaborators, finds that a single prompt can recover much of full-data OPD’s improvement in its experimental settings, while a small diverse prompt set can approach full-data performance.30 Its analysis distinguishes broad coverage of teacher-relevant states from the increasingly slow rate at which the student absorbs their supervision. The reported coverage is based on the paper’s state representation and measurement procedure, not literal coverage of all possible token prefixes.

Together, these results suggest optimizing the usefulness of the training states, not merely increasing the number of prompts or insisting on fresh complete trajectories. Replay improves reuse, prefix selection changes where supervision is queried, and prompt selection changes which states are generated. The remaining challenge is to balance these gains against stale-policy bias, missing student failure states, and the cost of identifying useful examples. One-shot results should not be read as a universal claim that broad task coverage is unnecessary.
6.7 Semantic and Temporal Credit: Supervising the Behavior That Matters
Dense token feedback has two separate limitations. It may supervise the wrong unit of behavior, and it may ignore the future consequences of an action. A surface token is not always the semantic decision we care about; an immediately plausible token is not always a step toward a successful trajectory.
When EOS Tokens Disagree, from Microsoft, UNC Chapel Hill, and BYU, provides a concrete semantic example.31 Teacher and student models can prefer different tokens for the same stopping action. Penalizing the student’s particular EOS token may then discourage termination even when the two models agree on the total probability of stopping. The paper studies aggregating equivalent termination tokens into a semantic STOP action:
where $\mathcal E_{\mathrm{stop}}$ contains tokens that terminate the same kind of response in the chosen protocol. Merely allowing both tokens in the decoding stop set does not repair a misaligned training signal. This addresses a specific termination mismatch, not arbitrary cross-tokenizer alignment.

The temporal issue is already visible in MiniLLM, which accounts for future distillation returns rather than treating every token as an isolated decision.12 For the following exact identity, hold the teacher fixed and define
using the same current policy that samples $y$. This specializes the rollout-policy notation of Section 6.3 to the point at which the gradient is evaluated. Under standard regularity assumptions, maximizing the negative sequence-level reverse KL yields the policy-gradient form
for a fixed prompt, with its expectation omitted for readability. The direct derivative of the student log-density has zero expectation; the remaining score-function terms can be organized into this return-to-go expression. In a surrogate-loss implementation, the sampled return is detached. A baseline depending on the prefix but not the current action has zero expected score-function contribution and can reduce variance. Stale rollouts, clipped importance weights, per-response length normalization, or changing teacher targets require a separate analysis; the identity does not establish that those modified estimators remain unbiased. Replacing $G_t^{\mathrm{KD}}$ with only $r_t^{\mathrm{KD}}$ changes the update in general: an early action also changes the later prefixes on which the teacher will be queried.

Using a complete return does not make an imperfect teacher reliable. Rethinking OPD reports degradation of teacher guidance with trajectory depth, while Princeton’s thinking-model study shows how local feedback can suppress useful self-correction.1426 These findings point toward supervision that respects semantic decisions, reasoning branches, and eventual outcomes. Step-level targets, selective intervention, and better value or return estimates are promising directions, but should be evaluated for both bias and variance rather than assumed to improve on simpler token-level updates.
6.8 Connecting the Research Directions to the Unified Framework
The papers above can be compared through the same four questions that organized SFT, RL, DPO, and basic OPD. The table summarizes their main emphasis; several methods change more than one component.
| Research question | Main component changed | Representative work | Remaining difficulty |
|---|---|---|---|
| Which teacher can actually help this student? | Teacher selection and the effective target $q^\star$ | Rethinking OPD; Unmasking OPD | Measuring conditional usefulness without expensive trial training |
| Where should imitation be strict or permissive? | Divergence, weighting, and update rule $\mathcal U$ | DistiLLM-2; EOPD; REOPOLD | Preserving valid exploration while correcting genuine mistakes |
| Does the guidance improve task success? | Supervision signal $I$ and its use in updates | GKD with RL; OPDVR; peer-conditioned multi-rollout OPD | Reconciling local preferences with trajectory-level outcomes |
| How can a teacher be built without a stronger external model? | Teacher information and construction | OPSD; SDFT; SDPO; W2S-OPD | Avoiding information shortcuts and unreliable self-supervision |
| How can complementary expertise be combined? | Teacher routing and domain/modality-specific targets | Nemotron-Cascade 2; OPOD; VOLD | Preventing interference and preserving existing strengths |
| Which states are worth generating or revisiting? | Training distribution $\mu$ and data reuse | DistiLLM; ReOPD; Rethinking OPD II | Trading compute savings against coverage and policy freshness |
| What behavior should receive credit? | Supervision granularity and temporal update structure | MiniLLM; EOS-alignment analysis; thinking-model diagnostics | Aligning semantic actions and long-term consequences |
A useful experimental distinction follows from this framework. Lower teacher–student divergence demonstrates closer imitation under a particular state distribution; higher task reward demonstrates better behavior under a particular evaluation protocol. Neither alone establishes that the student retained other skills, maintained exploration, or learned efficiently. Evaluating advanced OPD therefore benefits from a small set of complementary measurements: task success, diversity under a fixed sampling budget, retention across domains, and total training cost. The central research question is how to turn teacher information into improvements the student can use under its own inference conditions.
References
Bai, Y., et al. Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073, 2022. Preprint. ↩
Ouyang, L., et al. Training Language Models to Follow Instructions with Human Feedback. NeurIPS, 2022. ↩ ↩2
Ross, S., Gordon, G., and Bagnell, D. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS, PMLR 15:627–635, 2011. ↩ ↩2
Schulman, J., et al. Proximal Policy Optimization Algorithms. OpenAI. arXiv:1707.06347, 2017. ↩
Shao, Z., et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. DeepSeek-AI. arXiv:2402.03300, 2024. ↩
Rafailov, R., et al. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. Stanford University. NeurIPS, 2023. ↩
Hinton, G., Vinyals, O., and Dean, J. Distilling the Knowledge in a Neural Network. arXiv:1503.02531, 2015. ↩
Kim, Y., and Rush, A. M. Sequence-Level Knowledge Distillation. EMNLP, 2016, pp. 1317–1327. ↩
Lin, A., Wohlwend, J., Chen, H., and Lei, T. Autoregressive Knowledge Distillation through Imitation Learning. EMNLP, 2020, pp. 6121–6133. ↩
Wen, Y., Li, Z., Du, W., and Mou, L. f-Divergence Minimization for Sequence-Level Knowledge Distillation. ACL, 2023, pp. 10817–10834. ↩
Agarwal, R., et al. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes. Google DeepMind and collaborators. ICLR, 2024. ↩ ↩2
Gu, Y., et al. MiniLLM: Knowledge Distillation of Large Language Models. ICLR, 2024. The later arXiv v6, revised January 31, 2026, is titled MiniLLM: On-Policy Distillation of Large Language Models. ↩ ↩2
Yang, Z., et al. Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation. NVIDIA. Technical report, 2026; arXiv:2603.19220. ↩ ↩2
Li, Y., et al. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe. Tsinghua University and collaborators. arXiv:2604.13016, 2026. Preprint. ↩ ↩2
Armandpour, M., et al. Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why. Apple. arXiv:2605.10889, 2026. Preprint; Apple research page published July 2026. ↩
Ko, J., Kim, S., Chen, T., and Yun, S.-Y. DistiLLM: Towards Streamlined Distillation for Large Language Models. KAIST and Microsoft. ICML, 2024. ↩ ↩2
Ko, J., et al. DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs. KAIST and Microsoft. ICML, 2025. ↩
Jin, W., et al. Entropy-Aware On-Policy Distillation of Language Models. IBM Research and collaborators. ICML, 2026. ↩
Ko, J., et al. Scaling Reasoning Efficiently via Relaxed On-Policy Distillation. Microsoft Research. arXiv:2603.11137, 2026. The Microsoft Research publication page lists the later title Relaxed On-Policy Distillation: Selective Credit Allocation for Scaling Reasoning Efficiently under NeurIPS 2026; the mechanism summarized here follows the arXiv preprint. ↩
Lin, W., et al. On-policy Distillation with Verifiable Reward. Tsinghua LeapLab, THUNLP, and collaborators. arXiv:2608.24696, 2026. Preprint; arXiv v4, September 29, 2026 (UTC). ↩
Yu, W., et al. Multi-Rollout On-Policy Distillation via Peer Successes and Failures. Microsoft Research and collaborators. May 2026. Preprint. ↩
Zhao, S., et al. Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models. UCLA, HKU, and Meta Superintelligence Labs. ICML, 2026. Author project page. ↩
Shenfeld, I., et al. Self-Distillation Enables Continual Learning. MIT and ETH Zurich. ICML, 2026. Project page. ↩
Hübotter, J., et al. Reinforcement Learning via Self-Distillation. ETH Zurich, Max Planck Institute for Intelligent Systems, MIT, and Stanford. ICML, 2026. Project page. ↩
Yu, F., et al. Weak-to-Strong On-Policy Distillation. Microsoft Research, University of Maryland, and MBZUAI. July 2026. Preprint. ↩
Kaur, S., et al. Rethinking On-Policy Self-Distillation for Thinking Models. Princeton Language and Intelligence, Princeton University. arXiv:2607.05184, 2026. Preprint. ↩ ↩2
Zhao, T., et al. OPOD: On-Policy Omni Distillation. Tencent-involved joint research. arXiv:2607.20918, 2026. Preprint; arXiv v3, August 4, 2026. ↩
Bousselham, W., Kuehne, H., and Schmid, C. VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation. Tübingen AI Center, University of Tübingen, MIT-IBM Watson AI Lab, and Inria/ENS/CNRS/PSL. CVPR, 2026. Preprint. ↩
Liao, B., et al. Multi-Turn On-Policy Distillation with Prefix Replay. Microsoft Research and University of Amsterdam. arXiv:2607.04763, 2026. Preprint. ↩
Fu, Z., et al. Rethinking On-Policy Distillation of Large Language Models II: One Training Example. Tsinghua University and collaborators. arXiv:2609.04172, 2026. Preprint. ↩
Yang, Y., et al. When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation. UNC Chapel Hill, BYU, and Microsoft. arXiv:2609.20511, 2026. Preprint. ↩
