Problem
Rare accident scenarios are difficult to collect at scale and difficult to vary systematically. A controllable generator could provide a way to test perception and planning systems against interactions that are underrepresented in ordinary driving data.
Method
The project conditions a video generation model on scene-level affordances such as agent occupancy, lane structure, and intended motion. The goal is to control the event while preserving enough visual variation to remain useful for evaluation.
Results so far
Early experiments suggest that affordance conditioning improves coarse event structure first. Fine-grained temporal realism remains the main open problem.