Homepage

Biography

I am a master's student in the School of Computer Science at Fudan University, advised by Prof. Zuxuan Wu. My research explores generative visual modeling for embodied intelligence, with the goal of building controllable and efficient visual foundations that help embodied agents understand and predict how scenes evolve over time. My first- and co-first-author work has been accepted to CVPR 2025 and ACM MM 2026, alongside publications at ICCV, SIGGRAPH Asia, and ICASSP.

Research Interests

My research lies at the intersection of embodied intelligence and generative visual modeling. I focus on generative world models, spatial and motion intelligence, and efficient visual generation. In particular, I study camera- and motion-conditioned video synthesis, multimodal spatial reasoning and camera planning, large-motion temporal prediction, and efficient diffusion models as visual foundations for embodied agents.

Publications

Experience

Apr 2025 — Mar 2026

Research Intern · Shanghai, China

BiliBili Inc.

Virtual Human Group · One-step pixel diffusion for video frame interpolation (SPEED) and vision-language-driven camera-controllable video generation (CT-1).

Jul 2024 — Mar 2025

Research Intern · Shanghai, China

Huawei

Noah's Ark Lab · Large-motion video frame interpolation (EDEN) and lightweight video motion editing (MotionFollower).

Education

Sep 2018 — Jun 2022

Wuhan University

Undergraduate study · School of Chemistry and Molecular Sciences