Visual Input
Robot Cameras (Images / Video)
Generative 3D World Data Engine
Generates higher-quality embodied data distributions equipped with rich 3D structure.
Robot Cameras (Images / Video)
User Text Instruction
Proprioception / Rollout Traces
EMBODIED FOUNDATION MODEL
(Post-trained for Action CoT)
Autoregressive CoT aligned with action token generation.
ACTION TOKENIZER
(Pre-trained Action Decoder)
The action token space is growing dynamically.
Interactive World Dynamics Sandbox
Action rollout
Action chunks are converted into smooth arm motion locally.