Visual Input
Robot Cameras (Images / Video)
Generative 3D World Data Engine
Generates higher-quality embodied data distributions equipped with rich 3D structure
Robot Cameras (Images / Video)
User Text Instruction
Proprioception / Rollout Traces
EMBODIED FUNDATION MODEL
(Post-trained for Action CoT)
Autoregressive CoT for aligned with action token generation
ACTION TOKENIZER
(Pre-trained Action Decoder)
The Action token space is growing dynamically
RLHF
Action
Chunk
Interactive World Dynamics Sandbox
ACTION ROLLOUT
Action chunks are converted into smooth arm motion locally