G2A: From Generative Models to Robot Actions
Workshop at CoRL 2026
About
Large-scale generative models pretrained on video, image, and 3D data learn rich representations of objects, spatial structure, physical interactions, and how the visual world changes over time. These models are a promising source of prior knowledge for robot learning, especially because large-scale robot data remains expensive to collect and limited in diversity.
However, most pretraining data contains no robot action labels, while a policy must ultimately produce executable actions that respect the embodiment, dynamics, and sensing constraints of a physical system. This mismatch creates an action gap between generative pretraining and robot policy learning.
A video model can predict how a scene should change, and a 3D generative model can capture the geometry of an environment — but neither directly specifies which robot actions will reliably produce the desired change. G2A places the interface between generative models and robot policies at the center, asking how models trained to represent or generate changes in the world can help robots learn the actions that cause those changes.
How can knowledge learned from large-scale video, image, and 3D data collected without robot action labels be converted into executable robot policies?
Invited Speakers
Leading researchers across generative world models, robot policy learning, embodiment grounding, human-video learning, and evaluation.
Call for Papers
We invite submissions on any aspect of turning pretrained generative models into robot policies. We welcome new, preliminary, and in-progress work — including negative results, empirical comparisons, position papers, benchmarks, and evaluation methodologies.
Topics of interest
- Generative world models for robot learning. Generative models pretrained on video, images, and 3D data that learn representations of objects, spatial structure, physical interactions, and temporal changes, and their use as priors for robot policy learning.
- Vision-language-action models. Vision-language models and action-grounded variants, including VLAs and VLNs, that connect visual and semantic understanding to executable robot actions for navigation, manipulation, and other embodied tasks.
- World action models. Models that couple predictions of future world evolution with robot action generation, providing a direct interface between generative world modeling and executable robot behavior.
- World-to-action grounding. Methods for translating generated future states, trajectories, or observations into executable robot actions, including action inference, inverse dynamics, trajectory optimization, and policy learning.
- Predictive dynamics models. Models that learn how environments evolve under robot actions, including prediction of future states, observations, rewards, and interaction outcomes for model-based policy learning and planning.
- Closed-loop control from generative models. Methods that ground generative predictions into reliable behavior through feedback, replanning, error correction, uncertainty estimation, and recovery from prediction errors.
- Sim-to-real transfer. Methods for transferring generative priors, representations, policies, and predictive models from simulation to physical robots, including adaptation to visual, sensing, embodiment, and dynamics gaps.
- Cross-embodiment grounding. Methods for transferring action knowledge learned from videos, demonstrations, simulations, or other robots across different embodiments while respecting their kinematic, dynamic, and sensing constraints.
- Object- and interaction-centric representations. Representations of objects, geometry, contact, affordances, and interactions that expose actionable structure for connecting generative world models to robot behavior.
- Action supervision from diverse data. Methods for extracting, generating, or acquiring action supervision from large-scale video, image, and 3D data, human demonstrations, simulation, and robot interactions, including action inference, action discovery, human-to-robot retargeting, demonstration synthesis, and scalable data collection.
- Robot data efficiency and knowledge transfer. Methods for leveraging pretrained generative models and diverse data sources to reduce the amount of robot data required for policy learning, including data selection, quality assessment, adaptation, fine-tuning, and transfer of knowledge across tasks, environments, and embodiments.
- Evaluation of the action gap. Benchmarks and evaluation protocols for measuring how well generative models translate world knowledge into executable behavior, including action feasibility, controllability, generalization, robustness, data efficiency, long-horizon consistency, and inference cost.
Submission tracks
Research papers
Novel methods, empirical studies, algorithmic advances, benchmarks, datasets, or system demonstrations that bridge generative pretraining and robot action.
4–9 pages + referencesPosition & negative-result papers
Opinionated positions on the right abstractions or evaluation criteria, reproducibility studies, and rigorous negative results that the community should know about.
4–9 pages + referencesSubmission guidelines
Submissions should follow the CoRL 2026 LaTeX template and be between 4 and 9 pages, excluding references and appendices. All papers must be submitted as anonymized PDFs for double-blind review via OpenReview , with author names, affiliations, and acknowledgments removed and prior work cited in the third person.
Accepted papers will be featured in a poster session, with a small number selected for oral presentation. Proceedings are non-archival, allowing future conference or journal submissions; work under review at or accepted by other venues is welcome, though CoRL 2026 main-conference papers are not eligible.
Important dates
| Milestone | Date (AoE) |
|---|---|
| Submission deadline | October 8, 2026 |
| Acceptance notification | October 29, 2026 |
| Camera-ready due | November 5, 2026 |
| Workshop day | November 12, 2026 |
All deadlines are 11:59 PM Anywhere on Earth.
Awards
Thanks to the generous sponsorship of Lambda , the workshop will present:
- Best Paper Award: $3,000 in compute credits, selected from all accepted papers across both tracks and announced during the closing synthesis.
- Oral Paper Awards: $1,500 in compute credits each, for a small number of accepted papers selected for oral presentation at the workshop, chosen on review scores and breadth of interest to the audience.
- All other accepted papers: $400 in compute credits.
Spread the word
The call is distributed through this website, relevant mailing lists, social media, and the networks of the organizers and speakers. We especially encourage submissions from junior researchers, underrepresented groups, and institutions across different geographic regions.
Questions about the call? Contact action.gap.workshop@gmail.com .
Schedule
An interactive half-day event rather than a sequence of talks. Invited talks are paired with oral presentations of accepted papers and open discussion.
Tentative morning schedule. Times subject to change as the program is finalized.
Organizers
Sponsor
We gratefully acknowledge Lambda for sponsoring the paper awards and compute credits at this workshop.