I am a Member of Technical Staff (MTS) at Cartesia, where I build voice agents. Previously, I was a research scientist in Vladlen Koltun‘s org at Apple, working on video world models and 3D perception, and at Apple’s SLAM team for ARKit.
I earned my Ph.D. degree in Computer Science at UC Berkeley. Prior to Berkeley, I received my Bachelor’s degree from Yao Class at Tsinghua University. Earlier, I received a silver medal at the International Olympiad in Informatics (IOI 2011).
My expertise spans voice AI, video/image synthesis, 3D perception, and neural rendering, with a broader interest in artificial intelligence and computer vision. My PhD research focused on geometric deep learning from images with dissertation on “Learning to Detect Geometric Structures from Images for 3D Parsing”.
Ph.D. in Computer Science, 2020
University of California, Berkeley
B.Eng in Computer Science and Technology, 2015
Tsinghua University
”*” indicates equal contribution
I was a teaching assistant of the following courses at UC Berkeley: