ECoSim: Data Efficient Fine-Tuning for Controllable Traffic Simulation

ECCV 2026

Yu-Hsiang Chen1,2*, Wei-Jer Chang2*, Yi-Ting Chen1, Masayoshi Tomizuka2

1 National Yang Ming Chiao Tung University2 University of California, Berkeley

* Equal contribution

Teaser figure for ECoSim showing data-efficient multi-modal control of pretrained traffic models.

Three ways to specify behavior: sketch a path, transfer a reference behavior through a latent code, or describe an intent in text.

Abstract

Controllable traffic simulation is critical for testing autonomous driving systems, yet existing approaches often require retraining large generative models with extensive annotated data. We introduce a lightweight control adaptation framework that enables multi-modal controllability (sketch, latent behavior codes, and text) for pretrained state-of-the-art diffusion and autoregressive traffic models.

By modulating intermediate features through identity-initialized FiLM layers, our method efficiently adds new control modalities while preserving the base model's generative prior. Evaluated on Waymo Open Sim Agents Challenge, our approach demonstrates strong controllability with less than 1% of the paired control data.

Through context-aware condition transfer, our framework enables counterfactual scenario generation and long-tail synthesis while maintaining stable closed-loop driving realism and safety. Our framework unlocks new possibilities for controllable traffic simulation, enabling targeted scenario generation through lightweight adaptation of pretrained generative models.

Video

Method Overview

Model-agnostic control adaptation architecture diagram.

Train the control branch; freeze the traffic backbone. The modality encoder and FiLM layers modulate intermediate features. Near-identity initialization keeps the model close to its pretrained behavior at the start of adaptation. Paper, Fig. 2.

How are behavior latents learned? BehaviorVAE details
BehaviorVAE overview diagram.

BehaviorVAE encodes trajectories together with scene context into per-agent latent codes. It is trained separately on 10,000 WOMD scenarios; the paired-data percentages below refer to control-adapter training. Paper, Fig. 3.

Control across backbones

All control adapters below use 1% paired control data.

Table 1 · WOSAC controllable simulation Source
ControlTrainable paramsMeta ↑mADE (m) ↓mADE gain ↑
VBD-CL Diffusion
Unconditional12.3M0.71862.7223—
Sketch1.3M0.74320.964564.56%
Latent1.9M0.74400.922166.12%
Text1.9M0.73062.000526.51%
SMART-tiny-CLSFT Autoregressive
Unconditional7.0M0.77281.5044—
Sketch2.1M0.79840.256182.97%
Latent1.4M0.79120.323878.48%
Text2.0M0.77671.353710.02%

Reading the results. Meta measures distributional realism; mADE is the minimum average displacement error across sampled rollouts. Gain is the relative mADE reduction against each backbone's unconditional baseline. Bold marks the best accuracy and realism metrics within each backbone. Parameter counts are those reported for base-model training or control-adapter training, respectively.

Evaluation scope and data accounting

These are ground-truth-derived control evaluations on WOMD under WOSAC, not arbitrary-instruction success rates. The SMART sketch result uses Full-GT conditioning: all GT-valid agents receive controls. Under target-only conditioning, sketch mADE is 0.4602 m and Meta is 0.7898 (Table A2).

The 1% refers to paired data for adapter training (approximately 5,000 scenarios), excluding backbone pretraining and separate BehaviorVAE training. Transfer experiments below use the separate 10%-data SMART setting with decoding guidance described in Supplement E.1.

How much paired data is needed?

Latent control improves earliest; sketch and text benefit from more supervision.

Figure 5: SMART sample-efficiency curves for sketch, latent, and text. Top: mADE, lower is better. Bottom: WOSAC Meta realism, higher is better. Training data ranges from 0 to 10%, with 100%-data LoRA reference lines.

Fig. 5 · Sample efficiency. SMART backbone; 0% is the unconditional baseline. The reference lines jointly train the control branch and backbone LoRA adapters using 100% data. Percentages on the curves describe paired control-adapter data. Open PDF · Paper discussion.

Comparison with ProSim 1% vs. 100% paired control data
Table 2 · Same architecture, different adaptation Source
MethodPaired dataMeta ↑mADE (m) ↓mADE gain ↑
Original ProSim
Unconditional—0.7042.679—
Sketch100%0.7491.09958.96%
Text100%0.7092.33013.03%
ProSim with FiLM adaptation
Unconditional—0.7032.641—
Sketch1%0.7501.13057.23%
Text1%0.7171.82430.92%

FiLM achieves comparable sketch control and stronger text control with less paired supervision. Original ProSim uses the official checkpoint; FiLM uses a separately retrained unconditional ProSim backbone. Gains use each block's own baseline (Supplement E.3).

One scene, different futures

Transfer different reference intentions to the same target scene.

Paper Figure A1, scenario 1: reference behaviors for accelerating, turning right, decelerating, and going straight, with their corresponding conditioned trajectories in the same target scene.

Each column transfers a different reference behavior into the same scene. The query row shows the source behavior; the conditioned row shows the generated future. Paper, Fig. A1 (scenario 1).

Counterfactual results Control, diversity, and driving quality
Table 3 · Transferred controls Source
ControlControl ADE (m) ↓Success (%) ↑Coverage ↑Collision ↓Offroad ↓PDMScore ↑
Unconditional——201.020.05630.025278.40
Sketch0.372770.61287.170.08360.025279.30
Latent0.454273.89254.410.07740.025180.33
Text2.444669.84282.570.05200.022779.56

Control ADE measures alignment with the transferred behavior; success measures matching the requested maneuver; coverage measures spatial diversity. Collision and offroad are reported as rates, not percentages. PDMScore is adapted to the controlled agent and does not evaluate traffic-light compliance. This uses the separate 10%-data SMART setting with guidance (Supplement E.1).

Coverage and PDMScore improve across modalities, but sketch and latent control increase collision rates over the unconditional baseline. Stronger instruction following does not guarantee safer interactions.

Long-tail Scenario Generation

Match the context, then transfer a reference behavior through its latent code.

Context matching: filter candidates for feasibility, then rank their scene-context similarity to retrieve compatible agents.

Context matching finds compatible recipients for a reference behavior. The selected target receives latent control; surrounding agents remain unconditioned. Paper, Fig. 4.

Red: controlled target Orange: reference agent Blue: surrounding agents · Colored trail: reference trajectory · Control: latent

Unconditioned

Reference behavior

Conditioned

Table 4 · Why context matching matters Source
Context selectionControl ADE (m) ↓Collision ↓Offroad ↓PDMScore ↑
Random3.87230.31750.155339.06
Context match1.69940.15740.066165.27

Matching improves transfer accuracy and driving quality relative to random context selection. Collision and offroad are rates (0–1). These are latent-transfer results from the separate 10%-data SMART setting with guidance, not the 1% WOSAC experiment.

Limitations. Rare maneuvers can conflict with surrounding traffic: following a prompt may cause abrupt-braking or at-fault collisions. Behavior transfer here controls one agent at a time. Failure cases and discussion.

BibTeX

@inproceedings{ecosim,
  title={ECoSim: Data Efficient Fine-Tuning for Controllable Traffic Simulation},
  author={Chen, Yu-Hsiang and Chang, Wei-Jer and Chen, Yi-Ting and Tomizuka, Masayoshi},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2026}
}

For more details, please check our paper.