Pelican-Unify 1.0 brings visual understanding, reasoning, future imagination, and action generation into a single embodied foundation model. One vision-language model maps visual observations, instructions, and action histories into a common semantic representation, then produces reasoning about tasks, actions, and future outcomes. Its final hidden state conditions a Unified Future Generator, which jointly predicts future videos and actions through separate output heads within one denoising process.
Language, video, and action objectives jointly update the shared representation during training. With one checkpoint, the model scores 64.7 across eight VLM benchmarks, reaches 66.03 on WorldArena, and achieves 93.5 on RoboTwin. These evaluations show that a unified architecture can retain strong performance across visual understanding, world modeling, and robotic action generation.