Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation
Agent experience, retained in the generator
The same original query, before and after D-OPCD.
a photo of a car and a computer mouse
Selected test queries from the paper. Direct generation uses the original query only; harness runs may use agent-selected prompts. The evolved-skills setting uses the original skills with the base generator and newly evolved skills with the updated generator.
Direct generation · four-benchmark average
60.52 → 65.09
+4.57 points after internalizing agent experience.
Updated generator · renewed skill evolution
82.33 → 84.16
+1.83 points from a second round of ASE.
Abstract
An agentic harness can improve text-to-image generation through memory, skills, verification, and iterative prompt refinement. These gains remain external to the diffusion model and require the harness at inference time. Diffusion On-Policy Context Distillation (D-OPCD) treats the agent-improved prompt as privileged context and distills its guidance into the generator’s weights. The updated generator retains part of the harness’s benefit when conditioned on the original query alone. With Auto Skill Evolver (ASE) turning task experience into reusable skills, direct generation improves from 60.52 to 65.09 on average across four benchmarks. After internalization, the harness resets its learned skills and evolves again around the updated generator, raising the average from 82.33 to 84.16.
Diffusion On-Policy Context Distillation
From prompt guidance to model weights
The teacher sees the original query and the matched harness prompt. The student learns on its own denoising trajectories, conditioned on the original query alone.
Collect agent experience
Replay the selected skill-equipped harness to collect original queries and matched generation prompts.
Distill privileged context
Match the EMA teacher’s predictions along the student’s own sampling trajectory, with the harness prompt hidden from the student.
Resume harness evolution
Adopt the updated generator, reset the learned skills, and run ASE again to learn new guidance.
Auto Skill Evolver
Episodes become insights. Insights become skills.
ASE learns reusable prompt guidance from completed tasks while the generator remains fixed.
Record task experience
Capture the execution trajectory and external evaluation feedback from each completed task.
Refine across tasks
Extract guidance and mature it with supporting or contradicting evidence from later episodes.
Consolidate reusable guidance
Combine mature insights into bounded instructions that guide the first generation prompt.
Experiments
Internalization and continued co-evolution
Results on GenEval, GenEval2, WISE Verified, and R2I-Bench, with T2I-CompBench++ for out-of-distribution evaluation.
| Method | GenEval | GenEval2 | WISE | R2I-Bench | Avg. |
|---|---|---|---|---|---|
| Base generator, direct | 78.22 | 78.14 | 42.86 | 42.87 | 60.52 |
| Base + skill-free harness | 83.51 | 87.50 | 78.00 | 68.12 | 79.28 |
| Base + evolved skills | 87.63 | 88.83 | 81.43 | 69.41 | 81.83 |
| D-OPCD generator, direct | 81.28 | 82.24 | 50.00 | 46.83 | 65.09 |
| D-OPCD + skill-free harness | 86.08 | 88.41 | 81.14 | 73.68 | 82.33 |
| D-OPCD + newly evolved skills | 91.75 | 89.74 | 85.71 | 69.45 | 84.16 |
D-OPCD improves direct generation on all four benchmarks. The updated generator with a skill-free harness reaches 82.33, slightly above the original skill-equipped agent’s 81.83. Renewed ASE raises the updated agent’s average to 84.16, improving three benchmarks; R2I-Bench declines after this second skill-evolution round.
| Method | GenEval | GenEval2 | WISE | R2I-Bench | Avg. |
|---|---|---|---|---|---|
| Base generator, direct | 78.22 | 78.14 | 42.86 | 42.87 | 60.52 |
| Vanilla SFT | 80.09 | 80.94 | 46.00 | 46.51 | 63.38 |
| Diffusion-DPO | 77.56 | 79.04 | 45.43 | 46.70 | 62.18 |
| D-OPSD | 80.72 | 76.48 | 48.86 | 44.96 | 62.75 |
| D-OPCD | 81.28 | 82.24 | 50.00 | 46.83 | 65.09 |
All methods start from the same Z-Image-Turbo weights and use matched task records from the evolved harness. D-OPCD achieves the highest direct-generation score on every benchmark, with a four-benchmark average of 65.09.
| Method | GenEval | GenEval2 | WISE | R2I-Bench | Avg. |
|---|---|---|---|---|---|
| Original query + matched harness prompt | 81.28 | 82.24 | 50.00 | 46.83 | 65.09 |
| Harness prompt only | 80.93 | 81.13 | 49.14 | 47.32 | 64.63 |
| Original query + shuffled harness prompt | 74.50 | 75.38 | 48.57 | 45.76 | 61.05 |
The student always receives the original query. Task-matched query–prompt conditioning gives the best average. Shuffling the harness prompts reduces scores on all four benchmarks.
| Method | Color | Shape | Texture | 2D-Spatial | 3D-Spatial | Numeracy | Non-Spatial |
|---|---|---|---|---|---|---|---|
| Base generator | 77.45 | 57.54 | 72.79 | 32.32 | 41.27 | 67.98 | 31.24 |
| D-OPCD–GenEval | 81.61 | 57.47 | 73.00 | 36.19 | 41.53 | 70.26 | 31.52 |
| D-OPCD–GenEval2 | 81.44 | 56.56 | 72.36 | 37.33 | 42.02 | 71.05 | 31.31 |
| D-OPCD–WISE | 77.17 | 57.21 | 72.67 | 32.34 | 39.74 | 67.09 | 31.34 |
| D-OPCD–R2I-Bench | 77.06 | 57.84 | 73.09 | 32.00 | 41.07 | 68.42 | 31.33 |
Direct generation on seven held-out T2I-CompBench++ categories. Each D-OPCD checkpoint is named for its training benchmark. GenEval- and GenEval2-trained checkpoints improve color binding and 2D spatial relations; the WISE- and R2I-Bench-trained checkpoints remain close to the base overall.
Scores are on a 0–100 scale; higher is better. Avg. is the unweighted mean across the four benchmarks. Bold indicates the best score in each column.
Citation
@misc{wang2026internalizingagentexperiencediffusion,
title = {Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation},
author = {Wenxuan Wang and Zekai Liu and Weinan Zhang and Yu Cheng and Yang Yang},
year = {2026},
eprint = {2610.07250},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2610.07250}
}