I will receive my Ph.D. degree in Computer Science and Technology from Tsinghua University in June 2027, advised by Researcher Junliang Xing and Professor Chun Yu. I obtained my master’s degree from National University of Defense Technology in 2017 under the supervision of Professor Yong Dou.
From 2017 to 2019, I worked as an Algorithm Researcher at Samsung Research China‑Beijing. From 2019 to 2023, I served as Algorithm Team Lead and board member at Seaway Inc., a spin‑off incubated by the Institute of Computing Technology, Chinese Academy of Sciences.
My primary research interests cover image/video generation and editing and video world models. I am also deeply interested in embodied AI.
I am always happy to explore potential research collaborations. Feel free to get in touch via email.

Yushe Cao, Shikun Feng, Fei Shen, Dianxi Shi, Jianqiang Xia, Chun Yu, Junliang Xing
Association for the Advancement of Artificial Intelligence(AAAI,UnderReview) 2027
Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-conditioned video inpainting and rely on separate modules for human parsing, pose estimation, and garment warping. This multi-stage design complicates deployment and, more critically, allows errors in explicit geometric priors to propagate irreversibly into…

Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing
Association for the Advancement of Artificial Intelligence(AAAI,UnderReview) 2027
Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip dependence incurs prohibitive latency and computational overhead in practical continuous deployment. Naively enforcing causality disrupts pretrained bidirectional priors and substantially degrades synthesis quality. We introduce LiveVVT, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation. Within a fixed-size window, LiveVVT…

Yushe Cao, Xuechao Zou, Dianxi Shi, Junliang Xing, Chun Yu, Yuanze Wang
IEEE Transactions on Multimedia 2026
Although diffusion-based methods have substantially improved the controllability of multimodal face synthesis, their semantic alignment remains suboptimal because most existing approaches rely on implicit latent-space objectives to model the relationship between denoising variables and multimodal conditions. Such implicit modeling is often insufficient to enforce precise correspondence between synthesized faces and conditional inputs, especially under long-tailed semantic mask distributions where rare…

Yushe Cao, Dianxi Shi, Xing Fu, Xuechao Zou, Haikuo Peng, Xueqi Li, Chun Yu, Junliang Xing
Association for the Advancement of Artificial Intelligence 2026
While significant progress has been achieved in multimodal facial generation using semantic masks and textual descrip tions, conventional feature fusion approaches often fail to enable effective cross-modal interactions, thereby leading to suboptimal generation outcomes. To address this chal lenge, we introduce MDiTFace—a customized diffusion transformer framework that employs a unified tokenization strategy to process semantic mask and text inputs, eliminat…

Your Name, James Wang, Some Other Name, John Doe
International Conference on Machine Learning (ICML) 2024 Spotlight
Photo by Pineapple Supply Co. on Unsplash. Please put a tldr (too-long-didnt-read, 1~2 sentences) of your publication here. It is not recommended to put the actual abstract here because it is usually too long to fit in. $\LaTeX$ is supported. $a=b+c$.