Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation

Wenxuan Wang1,2,*, Zekai Liu3,*, Weinan Zhang1, Yu Cheng4,†, Yang Yang5,†
1Harbin Institute of Technology 2Shanghai AI Laboratory 3Shandong University
4Nanyang Technological University 5Shanghai Jiao Tong University

* Equal contribution. † Corresponding author.

Direct generation · four-benchmark average

60.52 → 65.09

+4.57 points after internalizing agent experience.

Updated generator · renewed skill evolution

82.33 → 84.16

+1.83 points from a second round of ASE.

Abstract

An agentic harness can improve text-to-image generation through memory, skills, verification, and iterative prompt refinement. These gains remain external to the diffusion model and require the harness at inference time. Diffusion On-Policy Context Distillation (D-OPCD) treats the agent-improved prompt as privileged context and distills its guidance into the generator’s weights. The updated generator retains part of the harness’s benefit when conditioned on the original query alone. With Auto Skill Evolver (ASE) turning task experience into reusable skills, direct generation improves from 60.52 to 65.09 on average across four benchmarks. After internalization, the harness resets its learned skills and evolves again around the updated generator, raising the average from 82.33 to 84.16.

Diffusion On-Policy Context Distillation

From prompt guidance to model weights

The teacher sees the original query and the matched harness prompt. The student learns on its own denoising trajectories, conditioned on the original query alone.

D-OPCD internalizes prompt-mediated harness gains through on-policy velocity matching.
01

Collect agent experience

Replay the selected skill-equipped harness to collect original queries and matched generation prompts.

02

Distill privileged context

Match the EMA teacher’s predictions along the student’s own sampling trajectory, with the harness prompt hidden from the student.

03

Resume harness evolution

Adopt the updated generator, reset the learned skills, and run ASE again to learn new guidance.

Auto Skill Evolver

Episodes become insights. Insights become skills.

ASE learns reusable prompt guidance from completed tasks while the generator remains fixed.

EPISODE

Record task experience

Capture the execution trajectory and external evaluation feedback from each completed task.

INSIGHT

Refine across tasks

Extract guidance and mature it with supporting or contradicting evidence from later episodes.

SKILL

Consolidate reusable guidance

Combine mature insights into bounded instructions that guide the first generation prompt.

Experiments

Internalization and continued co-evolution

Results on GenEval, GenEval2, WISE Verified, and R2I-Bench, with T2I-CompBench++ for out-of-distribution evaluation.

Main paper results
Method GenEval GenEval2 WISE R2I-Bench Avg.
Base generator, direct 78.22 78.14 42.86 42.87 60.52
Base + skill-free harness 83.51 87.50 78.00 68.12 79.28
Base + evolved skills 87.63 88.83 81.43 69.41 81.83
D-OPCD generator, direct 81.28 82.24 50.00 46.83 65.09
D-OPCD + skill-free harness 86.08 88.41 81.14 73.68 82.33
D-OPCD + newly evolved skills 91.75 89.74 85.71 69.45 84.16

D-OPCD improves direct generation on all four benchmarks. The updated generator with a skill-free harness reaches 82.33, slightly above the original skill-equipped agent’s 81.83. Renewed ASE raises the updated agent’s average to 84.16, improving three benchmarks; R2I-Bench declines after this second skill-evolution round.

Scores are on a 0–100 scale; higher is better. Avg. is the unweighted mean across the four benchmarks. Bold indicates the best score in each column.

Citation

@misc{wang2026internalizingagentexperiencediffusion,
  title = {Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation},
  author = {Wenxuan Wang and Zekai Liu and Weinan Zhang and Yu Cheng and Yang Yang},
  year = {2026},
  eprint = {2610.07250},
  archivePrefix = {arXiv},
  primaryClass = {cs.AI},
  url = {https://arxiv.org/abs/2610.07250}
}