Xiaojuan (Jeanne) Wang

I’m a research scientist at Google, working to make Gemini better at creative tasks.

During my PhD at the University of Washington, I worked on generative models for visual storytelling, advised by Steve Seitz, Brian Curless, and Ira Kemelmacher. Before that, I completed my master’s at ETH Zurich in the CVG lab. I’ve also interned at Meta, Adobe Research, and Google Research.

Away from research, I play tennis.

Portrait of Xiaojuan Wang

Blog

Diagram of an AI app turning voice or typed input into an interactive workspace We Built an AI App Without a Chat Box
Tong He, Xiaojuan Wang, and ChatGPT
September 2026
read post

What a trip-planning experiment taught us about replacing the chat transcript with an interactive workspace.


Selected Publications

Snow White standing with the seven dwarfs
Generative Keyframing
Xiaojuan Wang
PhD Thesis, University of Washington, 2025
thesis

Exploring how generative models can turn keyframes into expressive video, from smooth transitions and deep zooms to choreographed animal dances.

How Animals Dance (When You're Not Looking)
Xiaojuan Wang, Aleksander Holynski, Brian Curless, Ira Kemelmacher, Steve Seitz
arXiv Preprint, 2025
project page / arXiv

Starting from a small set of generated keyframes, e.g., a marmot in various poses, our method generates an animal dance video that follows a specified choreography pattern, extracted from a reference dance video.

Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation
Xiaojuan Wang, Boyang Zhou, Brian Curless, Ira Kemelmacher, Aleksander Holynski, Steve Seitz
ICLR, 2025
project page / arXiv

Given a pair of key frames as input, our method generates a continuous intermediate video with coherent motion by adapting a pretrained image-to-video diffusion model.

Generative Powers of Ten
Xiaojuan Wang, Janne Kontkanen, Brian Curless, Steve Seitz, Ira Kemelmacher, Ben Mildenhall, Pratul P. Srinivasan, Dor Verbin, Aleksander Holynski
CVPR, 2024 (Highlight)
project page / arXiv

Given a series of prompts describing a scene at drastically varying scales, our method creates a seamless zooming video.

Jump Cut Smoothing for Talking Heads
Xiaojuan Wang, Taesung Park, Yang Zhou, Eli Shechtman, Richard Zhang
arXiv Preprint, 2023
project page / arXiv

Given a talking head video, we remove the the filler words, repetitive words and so on, and create a seamless transition for the jump cut.


Earlier Work

• Learning 3D semantic reconstruction on octrees (GCPR 2019); Xiaojuan Wang, Martin Oswald, Ian Cherabier, Marc Pollefeys
• Supervised quantization for similarity search (CVPR 2016); Xiaojuan Wang, Ting Zhang, Guo-Jun Qi, Jinhui Tang, Jingdong Wang
• Multi-scale learning for low-resolution person re-identification (ICCV 2015); Xiang Li, Wei-Shi Zheng, Xiaojuan Wang, Tao Xiang, Shaogang Gong
• Cross-scenario transfer person reidentification (TCSVT 2015); Xiaojuan Wang, Wei-Shi Zheng, Xiang Li, Jianguo Zhang

Miscellanea

My first name in Chinese is 小鹃,which means "a little cuckoo bird". You can call me "jeanne".
I travel a lot, and enjoy spending time in the musuem.

This website borrows the template from Jon Barron