Shilong Zou.

邹世龙

Beijing Innovation Center of Humanoid Robotics

I received my master's degree from the National University of Defense Technology (2026) and my bachelor's degree from Dalian Maritime University (2023). During my master's studies, I was advised by Prof. Kai Xu and Prof. Chenyang Zhu. I currently work as a world model researcher at the Beijing Innovation Center of Humanoid Robotics.

My research focuses on Embodied AI (including WFM, VLA, and RL), Computer Vision, and Deep Learning.

World ModelsEmbodied IntelligenceComputer Vision
Portrait of Shilong Zou

We are actively exploring new ideas to create intelligent robotic systems capable of learning dexterous, human-like behaviors and enhancing human life, while also committed to building a universal world model.

News

  • [09/2026] 🎉🎉 We release the technical report of Pelican-Sim 1.0, a general world model simulator for embodied intelligence.
  • [07/2026] 🎉🎉 I am awarded the title of Outstanding Graduate by the National University of Defense Technology.
  • [05/2026] 🎉🎉 Pelican-Unified rank #1 globally on the World Arena leaderboard with an EWM Score of 66.03.
  • [05/2026] 🎉🎉 We release the technical report of Pelican-Unified 1.0, a unified embodied intelligence model for understanding, reasoning, imagination, and action.
  • [05/2026] 🎉🎉 Our team PAI@IAII wins the second place in the world model category of the AgiBot World Challenge @ICRA 2026.
  • [04/2026] 🎉🎉 I have been invited to give a presentation at NICE Academic (on the theme "HELLO, WORLD MODEL").
  • [03/2026] 🎉🎉 One paper is accepted by JNUDT 2026!
  • [01/2026] 🎉🎉 One paper is accepted by TOG 2026!
  • [01/2026] 🎉🎉 One paper is accepted by TIP 2026!
  • [09/2025] 🎉🎉 The code of CycleDiff has been released now!
  • [08/2025] 🎉🎉 One paper is accepted by CoRL 2025!
  • [05/2025] 🎉🎉 One paper is accepted by TVCG 2025!
  • [05/2025] 🎉🎉 Our 🤖AEG-bot project page is accessible at link.
  • [01/2025] 🎉🎉 I receive "Outstanding Second Prize Scholarship" of National University of Defense Technology.
  • [10/2022] 🎉🎉 I receive the highest undergraduate award — "President's Scholarship" and "Top Ten College Students". View details at link.

Publications (*indicates equal technical contribution)

Pelican-Sim 1.0 overview: action conditioning, simulator applications, and generalization
Arxiv 2026
Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence
Shilong Zou, Shilin Zhang, Yingji Zhang, Yuhang Huang, Yi Zhang, Zeyuan Ding, Han Dong, Junwei Liao, Yong Dai, Jian Tang, Xiaozhu Ju
arXiv preprint, 2026
[Paper] [Code] [Website] GitHub stars
  • World Model Simulator
  • Action-conditioned video prediction
  • Data Engine
Abstract

Pelican-Sim 1.0 predicts future visual observations from robot actions and recent visual context, providing a general simulator for embodied learning and decision making. A shared 28-dimensional action space accommodates different robot embodiments, while rendered action videos connect control inputs to visual generation. Sparse mixture-of-experts layers model heterogeneous dynamics, and causal adaptation with few-step distillation enables efficient autoregressive rollouts.

Training uses approximately one million trajectories collected in simulation and the real world. Evaluations show improvements in action controllability and video quality across AgiBotWorld Beta, RoboMIND, and RoboTwin. On RoboTwin, the simulator supports synthetic training data, policy evaluation, action selection, and policy improvement. Adding 500 generated trajectories to 50 demonstrations per task increases policy success from 70% to 93%, and simulated policy evaluations correlate with measured policy success at 0.994. Tests under changes in trajectories, scenes, objects, embodiments, and viewpoints further explore its generalization.

Adapted from the paper abstractSource ↗
Arxiv 2026
Pelican-Unified 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action
Lead unified world model capabilities, action-conditioned video prediction.
Lead the experiments and evaluation on the World Arena leaderboard (EWM Score of 66.03, rank #1).
[Report]
  • Embodied Foundation Models
  • Vision-Language-Action
  • World Models
  • Unified Learning
Abstract

Pelican-Unify 1.0 brings visual understanding, reasoning, future imagination, and action generation into a single embodied foundation model. One vision-language model maps visual observations, instructions, and action histories into a common semantic representation, then produces reasoning about tasks, actions, and future outcomes. Its final hidden state conditions a Unified Future Generator, which jointly predicts future videos and actions through separate output heads within one denoising process.

Language, video, and action objectives jointly update the shared representation during training. With one checkpoint, the model scores 64.7 across eight VLM benchmarks, reaches 66.03 on WorldArena, and achieves 93.5 on RoboTwin. These evaluations show that a unified architecture can retain strong performance across visual understanding, world modeling, and robotic action generation.

Adapted from the paper abstractSource ↗
Under review
Learning Visuotactile Policy for Fast and Stable Placement

Under review
[Paper] [Code] [Website]
  • Visuotactile Learning
  • Robotic Placement
  • Sim-to-Real
  • Policy Learning
Abstract

Placement is an essential yet underexplored component of robotic manipulation. Its contact-rich nature poses significant challenges for perception and control, making it inherently difficult to balance efficiency and stability. We present Fast and Stable Placement (FSP), a visuotactile policy learning framework that leverages multimodal sensing for enhanced contact perception and achieves reliable placement through online adjustment.

FSP adopts a two-stage curriculum for robust placement skill learning. An oracle policy is first trained with privileged information; then, it is distilled into a deployable policy that operates on carefully designed domain-invariant modalities, enabling direct sim-to-real policy transfer. These comprehensive modalities are tokenized and fused with the attention mechanism to support end-to-end policy learning. Extensive evaluations in both simulation and on real robots demonstrate that the learned FSP policy achieves a superior trade-off between efficiency and stability.

Project abstractSource ↗
TIP 2026
CycleDiff: Cycle Diffusion Models for Unpaired Image-to-image Translation
Shilong Zou*, Yuhang Huang*, Renjiao Yi, Chenyang Zhu, Kai Xu
IEEE Transactions on Image Processing (TIP), 2026
[Paper] [Code] [Website] GitHub stars
  • Diffusion Models
  • Unpaired Translation
  • Image Generation
  • Cross-Modal Learning
Abstract

CycleDiff addresses image translation between domains without paired training examples by jointly learning translation and diffusion. Existing diffusion-based translators often train these processes separately or connect them only loosely because diffusion operates on noisy inputs while translation concerns clean images. This separation can limit the quality of the learned mapping.

The framework extracts image components that represent the clean signal and translates these components within the diffusion process, allowing end-to-end optimization. A time-dependent translation network models mappings that change across diffusion steps. Together, these designs improve image fidelity and preservation of structure. Experiments cover RGB-to-RGB translation as well as mappings between RGB images, edges, semantic representations, and depth. The method improves FID over competing approaches, including reported gains of 19.61 and 19.67 over the second-best method on Dog–Cat and Dog–Wild translation.

Adapted from the paper abstractSource ↗
JNUDT 2026
From Geometric Analysis to Semantic Reasoning: The Evolution of Robotic Grasping Perception Paradigms
Shilong Zou, Yuhang Huang, Renjiao Yi, Chenyang Zhu, Kai Xu
Journal of National University of Defense Technology (JNUDT), 2026
[Paper]
  • Robotic Grasping
  • Visual Perception
  • Semantic Reasoning
  • Survey
Abstract

Robotic grasping perception serves as a fundamental prerequisite for autonomous manipulation and embodied intelligence, acting as a core technology for robots to interact with the physical world. As application scenarios expand from structured industrial assembly lines to unstructured environments such as households and logistics, the challenges facing grasping tasks have become increasingly complex. Modern robots are required not only to perceive the geometric attributes of objects in scenes characterized by clutter, occlusion, and varying lighting but also to understand semantic information and task-specific contextual constraints. Despite the proliferation of research, existing literature often focuses on single technical branches or physical stability assessments, lacking a systematic tracing of the evolutionary logic of perception paradigms. Bridging this gap is essential for developing general-purpose robotic systems that possess high generalization performance and robust task comprehension.

Author-provided abstractSource ↗
TOG 2026
LLM-enhanced Scene Graph Learning for Household Rearrangement
Wenhao Li*, Shilong Zou*, Zhinan Yu*, Zheng Zhou, Wenxuan Li, Chenyang Zhu, Ruizhen Hu, Kai Xu
ACM Transactions on Graphics (TOG), 2026
[Paper] [Website]
  • Scene Graphs
  • LLM Reasoning
  • Household Robotics
  • Task Planning
Abstract

AEG-bot is an autonomous mobile humanoid system for rearranging household scenes in previously unseen real-world environments. It identifies misplaced objects and reasons about suitable destinations using contextual knowledge and user preferences. LLM-enhanced scene graph learning enriches an initial scene graph with additional object information and contextual relationships, producing an Affordance Enhanced Graph.

Receptacle nodes encode context-dependent affordances that describe which objects can be placed on them. This representation supports planning to detect misplaced items and choose appropriate destinations. The complete system combines active exploration, RGB-D reconstruction, scene graph extraction, task planning, and execution in one pipeline, enabling the robot to perceive an unfamiliar environment and carry out rearrangement autonomously.

Adapted from the project abstractSource ↗
TVCG 2025
Part-aware Shape Generation with Latent 3D Diffusion of Neural Voxel Fields
Yuhang Huang*, Shilong Zou, Xinwang Liu, Kai Xu
IEEE Transactions on Visualization and Computer Graphics (TVCG), 2025
[Paper] [Code]
  • 3D Generation
  • Latent Diffusion
  • Neural Fields
  • Part-Aware Modeling
Abstract

This work generates structured 3D shapes with a latent diffusion model operating on neural voxel fields. The model targets accurate part structure while retaining detailed geometry and appearance. Diffusion in a compressed 3D representation makes higher-resolution generation possible, allowing the generated fields to capture finer visual and geometric detail.

A part-aware decoder incorporates part codes when reconstructing the voxel fields, improving decomposition of the shape into parts and the quality of rendered views. Experiments across multiple object categories compare the model with existing generative methods and show improvements in part-aware shape generation.

Adapted from the preprint abstractSource ↗
Arxiv 2026
AdaPower: Specializing World Foundation Models for Predictive Manipulation
Yuhang Huang*, Shilong Zou*, Jiazhao Zhang, Xinwang Liu, Ruizhen Hu, Kai Xu
Arxiv, 2026
[Paper]
  • World Foundation Models
  • Test-Time Adaptation
  • Predictive Control
  • Robot Manipulation
Abstract

AdaPower adapts general world foundation models for precise robotic manipulation, addressing the gap between realistic video generation and the accuracy needed for control. Existing uses of these models as synthetic data generators can be expensive and make limited use of pretrained vision-language-action policies.

The framework combines Temporal-Spatial Test-Time Training for adaptation during inference with Memory Persistence to maintain consistency over longer horizons. A model predictive control loop uses the adapted world model to improve the decisions of pretrained VLA policies. On LIBERO benchmarks, task success improves by more than 41% without retraining the policy, while the framework retains computational efficiency and general-purpose capabilities.

Adapted from the paper abstractSource ↗
CoRL 2025
LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation
Yuhang Huang, Jiazhao Zhang, Shilong Zou, Xinwang Liu, Ruizhen Hu, Kai Xu
Conference on Robot Learning (CoRL), 2025
[Paper] [Code] [Website] GitHub stars
  • Latent World Models
  • Diffusion Policy
  • Visual Foundation Models
  • Predictive Manipulation
Abstract

LaDi-WM improves predictive manipulation by forecasting future states in a visual latent space rather than generating pixel-level images of robot–object interactions. This space combines geometric features from DINO with semantic features from CLIP. Diffusion predicts how these representations evolve, making future-state learning easier and more generalizable than direct image prediction.

A diffusion policy incorporates the predicted states while repeatedly refining its actions, improving consistency and accuracy. Evaluations in simulation and on physical robots report policy-performance gains of 27.9% on LIBERO-LONG and 20% in the real-world setting. Further real-world tests examine the generalization of both the world model and its associated policies.

Adapted from the paper abstractSource ↗
Visual Computer
Enhancing Cross-domain Few-annotation Object Detection via Memory Storage-to-Adaptation Mechanism
Yuhang Huang*, Shilong Zou*, Ji Li, Xuesong Xu, Xiaohong Chen, Xinwang Liu, Kai Xu
The Visual Computer, 2026
[Paper] [Code] [Website] GitHub stars
  • Object Detection
  • Domain Adaptation
  • Few-Annotation Learning
  • Memory Learning
Abstract

MS2A tackles object detection when a model must transfer across domains with very few labeled target-domain examples. Changes in environment and limited annotations make this transfer difficult. The approach learns prior information from a large collection of unlabeled target-domain data, giving adaptation a broader view of the target distribution.

A memory storage module collects information about foreground objects and background context. A memory adaptation module then incorporates that knowledge into feature learning to produce more discriminative representations. Tests on newly constructed and public datasets show leading performance, including improvements of up to 10.4% over existing methods on challenging industrial datasets.

Adapted from the published abstractSource ↗

Projects (*indicates equal technical contribution)

AEG-bot
🤖 AEG-bot: A mobile humanoid robot for Scene Graph guided Household Rearrangement
Wenhao Li*, Shilong Zou*, Zhinan Yu*, Zheng Zhou*, Wenxuan Li, Yihan Cao, Zaisheng Sun, Yue Liu, Yue Chen, Shiquan Liu, Long Zhou, Chenyang Zhu, Ruizhen Hu, Kai Xu

👉 Click here [Website]

Talks

Selected Talks

Instructor

Development and Applications of Open-source World Model Systems

Embodied AI Training Camp · Cohort 7

Taught in the seventh cohort's customization training course, presented the development and applications of open-source world model systems, and received an honorary certificate.

Hangzhou, China

Event report
Shilong Zou teaching alongside a humanoid robot in Hangzhou Shilong Zou explaining world models beside a humanoid robot and presentation screen World model development lecture at the Embodied AI Training Camp
Speaker

11th Siyuan Learning Motivation Training Camp

Shared the latest world model research from the Beijing Innovation Center of Humanoid Robotics at the camp's exchange session.

World model exchange session with Siyuan Training Camp participants Group photo from the Siyuan Training Camp visit
Roundtable participant

World Models for Embodied Intelligence: Future Paradigms and the Development Landscape

World Model Roundtable · InnoAngel Fund

Joined the roundtable hosted by InnoAngel Fund and shared experience in world model research.

Yanqing, Beijing · 世园璞燊酒店

Event report
Shilong Zou sharing experience at the InnoAngel Fund World Model Roundtable
Keynote speaker

Pelican-Unified 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action

Technology and Application Summit · Gaogong Humanoid Robotics

Delivered a keynote on Pelican-Unified 1.0 and its unified capabilities for embodied intelligence.

Shanghai · 新华联索菲特大酒店

Event report
Shilong Zou delivering the Pelican-Unified keynote in Shanghai Shilong Zou presenting the world model architecture at the Gaogong summit Shilong Zou speaking during the Gaogong Humanoid Robotics Summit discussion Shilong Zou attending the Gaogong Humanoid Robotics Summit in Shanghai
Keynote speaker

From Perception to Execution: Embodied Intelligence Video Generation and Action Prediction Based on World Models

HELLO WORLD (MODEL): Frontiers, Practice and Breakthroughs · NICE Academic

Delivered a keynote at the academic forum hosted by NICE Academic.

Dinghao H3, Haidian, Beijing

Event report
Shilong Zou presenting at the NICE Academic forum World model presentation at Dinghao H3 Audience attending the NICE Academic forum Shilong Zou speaking at the NICE Academic forum

Honors Awards

Received more than 20 scholarships and grants totaling more than 100,000 RMB

⬆ scrollable

Competition Awards

Received more than 70 awards at international, national, provincial, and university level or above

First prize
National First Prize in the China College Students Computer Design Competition (only two teams in China)
Haobin Wang*, Shilong Zou*, Mengyue Deng*
China College Students Computer Design Competition
First prize
World Arena benchmark (rank #1 with an EWM Score of 66.03)
Leader: Shilong Zou

Others (Selected)

  • [07/2022] National First Prize in the National Undergraduate Transportation Science and Technology Competition
  • [10/2021] National First Prize in the National College Business Elite Challenge
  • [12/2020] National Second Prize in the National College Students Mathematics Competition
  • [11/2023] Third Prize of “Huawei Cup” 20th China Graduate Student Mathematical Modeling Competition at National Level
⬆ scrollable

Education

M.S. in Computer Vision and Robotics
School of Computer Science
Comprehensive evaluation score: 3.51, rank: 2/187

National University of Defense Technology
2023 - 2026
Bachelor of Engineering
School of Information Science and Technology
GPA: 4.38/5.0, rank: 2/98


Dalian Maritime University
2019 - 2023

Internship

Beijing Innovation Center of Humanoid Robotics
Department of Foundation Models

Focus: research on embodied world foundation models and action-conditioned video prediction
2026.01 - present

Miscellaneous

Travel enthusiast 🏞️ and marathon runner 🏃.

  • [01/2025] 🏃 Completed the first half marathon with a score of 138 (average 4.40) - Songya Lake New Year Half Marathon
  • [03/2025] 🏃 Ranked 50th (average score of 4.05) in the Campus Run, National University of Defense Technology
  • [03/2025] 🏃 Completed the first A1 marathon with a score of 135 (average 4.30) - Zhangjiajie Wulingyuan Half Marathon
  • [04/2025] 🏃 Bronze Medalist🥉 in the Men's 800 Meters at the Sports Meet of Cadet Second Battalion, College of Computer Science, National University of Defense Technology