G2A: From Generative Models to Robot Actions

Workshop at CoRL 2026

Date Thursday, November 12, 2026 CoRL 2026 workshop day
Time 08:30 – 12:30 CST Half-day · morning session
Location Austin, Texas, USA Room assignment to be announced
Call for Papers OpenReview
01

About

Large-scale generative models pretrained on video, image, and 3D data learn rich representations of objects, spatial structure, physical interactions, and how the visual world changes over time. These models are a promising source of prior knowledge for robot learning, especially because large-scale robot data remains expensive to collect and limited in diversity.

However, most pretraining data contains no robot action labels, while a policy must ultimately produce executable actions that respect the embodiment, dynamics, and sensing constraints of a physical system. This mismatch creates an action gap between generative pretraining and robot policy learning.

A video model can predict how a scene should change, and a 3D generative model can capture the geometry of an environment — but neither directly specifies which robot actions will reliably produce the desired change. G2A places the interface between generative models and robot policies at the center, asking how models trained to represent or generate changes in the world can help robots learn the actions that cause those changes.

Central question

How can knowledge learned from large-scale video, image, and 3D data collected without robot action labels be converted into executable robot policies?

02

Invited Speakers

Leading researchers across generative world models, robot policy learning, embodiment grounding, human-video learning, and evaluation.

YD

Yilun Du

Harvard University · Kempner Institute
Committed · in person
Generative models for embodied decision-making — energy-based models, diffusion, compositional world models, and planning.
FH

Furong Huang

University of Maryland
Committed · in person
Foundation models as decision-making systems — robust RL, reasoning under uncertainty, and trustworthy self-improvement.
AB

Amir Bar

Imperial College London · AMI Labs
Committed · in person
Learning to perceive, reason, and act from visual data — world models, video prediction, planning, and visuomotor control.
QR

Qwen Robotics Team

Alibaba, Inc.
Committed · in person
Industrial-scale view on extending multimodal foundation models toward embodied intelligence — world modeling and manipulation.
03

Call for Papers

We invite submissions on any aspect of turning pretrained generative models into robot policies. We welcome new, preliminary, and in-progress work — including negative results, empirical comparisons, position papers, benchmarks, and evaluation methodologies.

Topics of interest

Submission tracks

Track 1

Research papers

Novel methods, empirical studies, algorithmic advances, benchmarks, datasets, or system demonstrations that bridge generative pretraining and robot action.

4–9 pages + references
Track 2

Position & negative-result papers

Opinionated positions on the right abstractions or evaluation criteria, reproducibility studies, and rigorous negative results that the community should know about.

4–9 pages + references

Submission guidelines

Submissions should follow the CoRL 2026 LaTeX template and be between 4 and 9 pages, excluding references and appendices. All papers must be submitted as anonymized PDFs for double-blind review via OpenReview (submission portal to be announced), with author names, affiliations, and acknowledgments removed and prior work cited in the third person.

Accepted papers will be featured in a poster session, with a small number selected for oral presentation. Proceedings are non-archival, allowing future conference or journal submissions; work under review at or accepted by other venues is welcome, though CoRL 2026 main-conference papers are not eligible.

Important dates

MilestoneDate (AoE)
Submission deadlineOctober 8, 2026
Acceptance notificationOctober 29, 2026
Camera-ready dueNovember 5, 2026
Workshop dayNovember 12, 2026

All deadlines are 11:59 PM Anywhere on Earth. Dates are tentative until the OpenReview portal opens.

Awards

Thanks to the generous sponsorship of Lambda, the workshop will present:

Spread the word

The call is distributed through this website, relevant mailing lists, social media, and the networks of the organizers and speakers. We especially encourage submissions from junior researchers, underrepresented groups, and institutions across different geographic regions.

Questions about the call? Contact xu.xinyi5@northeastern.edu.

04

Schedule

An interactive half-day event rather than a sequence of talks. Invited talks act as focused inputs for the breakout problem-solving sessions that follow.

08:30 – 08:45
FramingOpening framing
08:45 – 09:10
TalkInvited talk 1
09:15 – 09:40
TalkInvited talk 2
09:40 – 10:00
InteractiveBreakout problem-solving · part 1
10:00 – 10:30
DiscussionBenchmark / evaluation discussion
10:30 – 11:00
BreakCoffee break & poster session
11:00 – 11:25
TalkInvited talk 3
11:30 – 11:55
TalkInvited talk 4
11:55 – 12:15
InteractiveBreakout problem-solving · part 2
12:15 – 12:30
SynthesisReport-back and synthesis

Tentative morning schedule. Times subject to change as the program is finalized.

05

Organizers

XX

Xinyi Xu

Northeastern University
Undergraduate
YQ

Yu Qi

Northeastern University
PhD Student
YY

Yifan Yin

Johns Hopkins University
PhD Student
KZ

Kevin Zhang

Peking University
PhD Student
XC

Xiaowei Chi

CUHK
PhD Student
SM

Siyuan Ma

CUHK · Westlake University
PhD Student
SW

Siyuan Wang

CUHK
Postdoc Researcher
LW

Lawson Wong

Northeastern University
Associate Professor
TS

Tianmin Shu

Johns Hopkins University
Assistant Professor
JM

Jiayuan Mao

UPenn · Amazon
Assistant Professor
JT

Jonathan Tremblay

NVIDIA, Inc.
Research Scientist
VB

Valts Blukis

NVIDIA, Inc.
Research Scientist
WL

Weiyang Liu

CUHK
Assistant Professor
JX

Jianwen Xie

Lambda, Inc.
Research Scientist