Visual Futures for Manipulation
Exploring whether generated future video can guide robot actions, and where visual plans fall short.
Overview
This exploratory project asks whether a pretrained video model can suggest useful future observations for a manipulation policy. The visuals below illustrate the proposed pipeline; they are not evidence of reliable zero-shot robot performance.
The pipeline generates candidate future frames from a current observation and text goal, then uses them as context for action prediction. Subsequent visual-plan experiments did not pass their registered benefit gate, so this remains a research question rather than a validated zero-shot manipulation system.
The proposed policy uses the current observation and a feature derived from predicted frames. Cross-attention provides one way to combine the present observation with imagined futures. The next step is to test when those predictions add useful information beyond a policy trained on observed data.
Project proposal
The original course project proposal records the motivation and planned architecture. It predates the later evaluation and should be read as a proposal.