Yushe Cao
Ph.D.
Tsinghua University

I will receive my Ph.D. degree in Computer Science and Technology from Tsinghua University in June 2027, advised by Researcher Junliang Xing and Professor Chun Yu. I obtained my master’s degree from National University of Defense Technology in 2017 under the supervision of Professor Yong Dou.

From 2017 to 2019, I worked as an Algorithm Researcher at Samsung Research China‑Beijing. From 2019 to 2023, I served as Algorithm Team Lead and board member at Seaway Inc., a spin‑off incubated by the Institute of Computing Technology, Chinese Academy of Sciences.

My primary research interests cover image/video generation and editing and video world models. I am also deeply interested in embodied AI.

I am always happy to explore potential research collaborations. Feel free to get in touch via email.


Education
  • Tsinghua University
    Tsinghua University
    Ph.D. in Department of Computer Science and Technology
    Sep. 2024 - present
  • National University of Defense Technology
    National University of Defense Technology
    M.S. in College of Computer Science
    Sep. 2019 - Jul. 2021
Experience
  • Seaway Inc.
    Seaway Inc.
    Alg. Lead & Board Member
    OCT. 2019 - Sep. 2023
  • Samsung Research China-Beijing
    Samsung Research China-Beijing
    Algorithm Researcher
    Jul. 2017 - Sep. 2019
News
2024
AI Transforms Music Industry: First AI-Composed Symphony Debuts in New York
Oct 19
Virtual Reality Theme Park Opens, Redefining Entertainment Industry
Mar 22
First Human Settlement Established on Mars, Marking New Era of Space Exploration. Read more
Jan 30
2023
Scientists Discover New Species of Bioluminescent Fish in Mariana Trench
Nov 28
AI-Powered Robot Chef Wins International Culinary Competition Featured
Sep 05
2022
Lorem ipsum sit amet, consectetur adipiscing elit, sed do eiusmod tempor
Jan 11
Selected Publications
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on

Yushe Cao, Shikun Feng, Fei Shen, Dianxi Shi, Jianqiang Xia, Chun Yu, Junliang Xing

Association for the Advancement of Artificial Intelligence(AAAI,UnderReview) 2027

Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-conditioned video inpainting and rely on separate modules for human parsing, pose estimation, and garment warping. This multi-stage design complicates deployment and, more critically, allows errors in explicit geometric priors to propagate irreversibly into…

UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on

Yushe Cao, Shikun Feng, Fei Shen, Dianxi Shi, Jianqiang Xia, Chun Yu, Junliang Xing

Association for the Advancement of Artificial Intelligence(AAAI,UnderReview) 2027

Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-conditioned video inpainting and rely on separate modules for human parsing, pose estimation, and garment warping. This multi-stage design complicates deployment and, more critically, allows errors in explicit geometric priors to propagate irreversibly into…

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time
LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing

Association for the Advancement of Artificial Intelligence(AAAI,UnderReview) 2027

Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip dependence incurs prohibitive latency and computational overhead in practical continuous deployment. Naively enforcing causality disrupts pretrained bidirectional priors and substantially degrades synthesis quality. We introduce LiveVVT, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation. Within a fixed-size window, LiveVVT…

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing

Association for the Advancement of Artificial Intelligence(AAAI,UnderReview) 2027

Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip dependence incurs prohibitive latency and computational overhead in practical continuous deployment. Naively enforcing causality disrupts pretrained bidirectional priors and substantially degrades synthesis quality. We introduce LiveVVT, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation. Within a fixed-size window, LiveVVT…

Long-Tail Adaptive Flow Matching with Explicit Conditional Consistency Guidance for Precise Multimodal Face Synthesis
Long-Tail Adaptive Flow Matching with Explicit Conditional Consistency Guidance for Precise Multimodal Face Synthesis

Yushe Cao, Xuechao Zou, Dianxi Shi, Junliang Xing, Chun Yu, Yuanze Wang

IEEE Transactions on Multimedia 2026

Although diffusion-based methods have substantially improved the controllability of multimodal face synthesis, their semantic alignment remains suboptimal because most existing approaches rely on implicit latent-space objectives to model the relationship between denoising variables and multimodal conditions. Such implicit modeling is often insufficient to enforce precise correspondence between synthesized faces and conditional inputs, especially under long-tailed semantic mask distributions where rare…

Long-Tail Adaptive Flow Matching with Explicit Conditional Consistency Guidance for Precise Multimodal Face Synthesis

Yushe Cao, Xuechao Zou, Dianxi Shi, Junliang Xing, Chun Yu, Yuanze Wang

IEEE Transactions on Multimedia 2026

Although diffusion-based methods have substantially improved the controllability of multimodal face synthesis, their semantic alignment remains suboptimal because most existing approaches rely on implicit latent-space objectives to model the relationship between denoising variables and multimodal conditions. Such implicit modeling is often insufficient to enforce precise correspondence between synthesized faces and conditional inputs, especially under long-tailed semantic mask distributions where rare…

Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation

Yushe Cao, Dianxi Shi, Xing Fu, Xuechao Zou, Haikuo Peng, Xueqi Li, Chun Yu, Junliang Xing

Association for the Advancement of Artificial Intelligence 2026

While significant progress has been achieved in multimodal facial generation using semantic masks and textual descrip tions, conventional feature fusion approaches often fail to enable effective cross-modal interactions, thereby leading to suboptimal generation outcomes. To address this chal lenge, we introduce MDiTFace—a customized diffusion transformer framework that employs a unified tokenization strategy to process semantic mask and text inputs, eliminat…

Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation

Yushe Cao, Dianxi Shi, Xing Fu, Xuechao Zou, Haikuo Peng, Xueqi Li, Chun Yu, Junliang Xing

Association for the Advancement of Artificial Intelligence 2026

While significant progress has been achieved in multimodal facial generation using semantic masks and textual descrip tions, conventional feature fusion approaches often fail to enable effective cross-modal interactions, thereby leading to suboptimal generation outcomes. To address this chal lenge, we introduce MDiTFace—a customized diffusion transformer framework that employs a unified tokenization strategy to process semantic mask and text inputs, eliminat…

Convallis a cras semper auctor neque vitae rutrum quisque non tellus orci ac
Convallis a cras semper auctor neque vitae rutrum quisque non tellus orci ac

Your Name, James Wang, Some Other Name, John Doe

International Conference on Machine Learning (ICML) 2024 Spotlight

Photo by Pineapple Supply Co. on Unsplash. Please put a tldr (too-long-didnt-read, 1~2 sentences) of your publication here. It is not recommended to put the actual abstract here because it is usually too long to fit in. $\LaTeX$ is supported. $a=b+c$.

Convallis a cras semper auctor neque vitae rutrum quisque non tellus orci ac

Your Name, James Wang, Some Other Name, John Doe

International Conference on Machine Learning (ICML) 2024 Spotlight

Photo by Pineapple Supply Co. on Unsplash. Please put a tldr (too-long-didnt-read, 1~2 sentences) of your publication here. It is not recommended to put the actual abstract here because it is usually too long to fit in. $\LaTeX$ is supported. $a=b+c$.