Hypothesis
Explicit affordance maps should make the generated motion easier to control than text-only conditioning, especially around interactions between agents.
First observation
The conditioning signal improves coarse scene layout before it improves fine motion. This suggests evaluating control fidelity at multiple temporal scales instead of using a single visual similarity score.
Next run
The next experiment will separate lane geometry, agent occupancy, and motion intent so that each source of improvement can be measured independently.