I am engaged in research in Computer Vision (CV) and Machine Learning (ML), particularly in the area of vision-language perception and generation. My ultimate research goal is to enable intelligent systems to perceive, reason, and predict in complex visual environments. Specifically, my current research interests lie in:
Controllable Video Generation: Developing controllable text/image-to-video frameworks, with interests in test-time optimization and compositional generation.
Generative World Modeling: Scaling generative models to follow interactive instructions and integrate structured human knowledge for better prediction.
Video Understanding & Reasoning: Leveraging foundations in action recognition and temporal grounding to interpret long-range event dynamics.
I am open to any communication and collaboration. Please feel free to contact me if you’re interested!
News
🎉 One paper is accepted by NeurIPS 2026! See U in Atlanta 🐠
🎉 One paper is accepted by ECCV 2026! See U in Malmö 🇸🇪
Starting my internship at Adobe in San Jose!
🎉 One co-authored paper is accepted by ICLR 2026. Congrats to Wenliang!
🎉 One paper is accepted by WACV 2026! See U in Tucson 🌵