Portrait placeholder for Yinuo Wang

Yinuo Wang

Hi, I’m Yinuo. I build digital humans that can talk, listen, respond, and remember.

My research focuses on generative digital avatars, with an emphasis on modeling expressive listening behavior and creating interactive 3D heads that respond naturally to people in real time. Ultimately, I hope to make virtual humans feel less like scripted animations and more like perceptive, consistent conversational partners.

My recent work explores diffusion-based interactive head generation, multimodal interaction, long-term conversational memory, and controllable generation for 3D heads. Before focusing on digital avatars, I worked on a range of machine learning problems, including multi-view clustering, image classification, and robust time-series forecasting. This background continues to shape how I think about representing and generating complex human behavior.

I am currently in the final stages of my PhD at Xi'an Jiaotong University and am seeking new opportunities. If you are interested in my work or potential collaboration, please feel free to get in touch.


News

Publications

MimicTalker paper thumbnail

MimicTalker: A Multimodal Interactive and Memory-Enhanced Framework for Real-Time Dyadic 3D Head Generation

Yinuo Wang, Yanbo Fan, Xuan Wang, Boyao Zhou, Yu Guo, Yujun Shen, Fei Wang
CVPR, 2026

A framework for real-time, context-aware, and long-term consistent dyadic 3D head generation.

HeadLighter paper thumbnail

HeadLighter: Disentangling Illumination in Generative 3D Gaussian Heads via Lightstage Captures

Yating Wang, Yuan Sun, Xuan Wang, Ran Yi, Boyao Zhou, Yipengjing Sun, Hongyu Liu, Yinuo Wang, Lizhuang Ma
arXiv, 2026

Disentangles appearance and illumination in generative 3D Gaussian heads for explicit relighting and viewpoint editing.

Diffusion-based Realistic Listening Head Generation paper thumbnail

Diffusion-based Realistic Listening Head Generation via Hybrid Motion Modeling

Yinuo Wang, Yanbo Fan, Xuan Wang, Guo Yu, Fei Wang
CVPR, 2025 (Highlight)

Generates high-fidelity listening-head videos with expressive motion using hybrid motion modeling and diffusion.