I am currently a Ph.D. student at the School of Mathematics and Statistics, University of Melbourne, supervised by Prof. Mingming Gong (宫明明). Meanwhile, I work as an AI Research Intern at HeyGen, Singapore, mentored by Mr. Yi Ren (任意) and Mr. Zhibin Hong (洪智滨). Before that, I received my M.Phil. degree from The Chinese University of Hong Kong, Shenzhen (CUHK-SZ) in 2023, supervised by Prof. Xiaoguang Han (韩晓光), and my B.Eng. degree from Tsinghua University in 2021, supervised by Prof. Yipeng Li (李一鹏).
My research interests include generative models, visual language models and multimodal language models. I have published 10+ papers at top international AI conferences such as CVPR, ECCV, and ICLR, with total google scholar citations 700+ ().
If you are interested in any form of academic cooperation, please feel free to email me at antonio.chan.cc@outlook.com.
🔥 News
- 2026.07: 🎉 One paper is accepted by SIGGRAPH Asia 2026.
- 2026.06: 🎉 One paper is accepted by ECCV 2026.
- 2026.02: 🎉 One paper is accepted by CVPR 2026.
- 2026.01: 🎉 One paper is accepted by ICLR 2026.
- 2025.06: 🎉 I join HeyGen as an AI research intern in Singapore!
- 2024.07: 🎉 One paper is accepted by ECCV 2024.
- 2023.02: 🎉 One paper is accepted by CVPR 2023.
📝 Publications
* indicates equal contribution.
Video Generation & Understanding

TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
Ce Chen, Yi Ren, Yuanming Li, Viktor Goriachko, Zhenhui Ye, Zujin Guo, Zhibin Hong, Mingming Gong
Project | arXiv | GitHub | Blog
- Proposes the Shot Transition Detection (STD) task and TransVLM, a VLM framework that injects optical flow as motion prior for detecting continuous temporal segments of video transitions.
- Deployed to production at HeyGen.

TAVR: Generate Your Talking Avatar from Video Reference
Zujin Guo, Zhenhui Ye, Yi Ren, Yuanming Li, Ce Chen, Zhibin Hong, Chen Change Loy
- A novel framework for cross-scene talking avatar generation from video references, integrating token selection and a three-stage training scheme.
- Deployed to production at HeyGen.
Avatar V: Scaling Video-Reference Avatar Video Generation
Benjamin Liang, Ce Chen, Desmond Lin, Ivan Somov, Jiajun Zhao, Jiewei Yuan, Jingfeng Zhang, Junhao Huang, Nik Nolte, Pedram Haqiqi, Penghan Wang, Rong Yan, Rui Zhang, Sam Prokopchuk, Sivan Wang, Viktor Goriachko, Yi Ren, Yuanming Li, Yutao Chen, Zhenhui Ye, Zhibin Hong, Zilong Nie, Zujin Guo
- A production-scale framework for video-reference-conditioned avatar generation. Introduces Sparse Reference Attention for linear-complexity conditioning on long references.
- Deployed to production at HeyGen, trained across thousands of GPUs with a data engine curating 100M+ training clips.
Image Generation & Understanding

Mobile-VTON: High-Fidelity On-Device Virtual Try-On
Zhenchen Wan*, Ce Chen*, Runqi Lin, Jiaxin Huang, Tianxi Chen, Yanwu Xu, Tongliang Liu, Mingming Gong
- A privacy-preserving virtual try-on framework enabling fully offline execution on commodity mobile devices. Introduces a modular TGT architecture with Feature-Guided Adversarial Distillation.
- Matches or outperforms server-based baselines at 1024×768 resolution on VITON-HD and DressCode.
3D/4D Generation & Understanding

Condition Matters in Full-Head 3D GANs
Heyuan Li, Huimin Zhang, Yuda Qiu, Zhengwentai Sun, Keru Zheng, Lingteng Qiu, Peihao Li, Qi Zuo, Ce Chen, Yujian Zheng, Yuming Gu, Zilong Dong, Xiaoguang Han
- Proposes view-invariant semantic feature conditioning to decouple 3D head generation from viewing direction, achieving higher fidelity and diversity in full-head synthesis.

CT4D: Consistent Text-to-4D Generation with Animatable Meshes
Ce Chen, Shaoli Huang, Xuelin Chen, Guangyi Chen, Xiaoguang Han, Kun Zhang, Mingming Gong
- Presents a novel text-to-4D framework operating on animatable meshes. Introduces the Generate-Refine-Animate (GRA) algorithm for text-aligned mesh generation and uniform driving functions for surface continuity.

SphereHead: Stable 3D Full-head Synthesis with Spherical Tri-plane Representation
Heyuan Li, Ce Chen, Tianhao Shi, Yuda Qiu, Sizhe An, Guanying Chen, Xiaoguang Han
- Proposes a spherical tri-plane representation in the spherical coordinate system that mitigates “mirroring” artifacts in full-head 3D GAN synthesis, along with a view-image consistency loss.

SCoDA: Domain Adaptive Shape Completion for Real Scans
Yushuang Wu, Zizheng Yan, Ce Chen, Lai Wei, Xiao Li, Guanbin Li, Yihao Li, Shuguang Cui, Xiaoguang Han
- Proposes the novel SCoDA task for domain adaptive shape completion on real scans, along with the ScanSalon dataset. Introduces cross-domain feature fusion and volume-consistent self-training.
Medical Image Analysis

DenseMP: Unsupervised Dense Pre-training for Few-shot Medical Image Segmentation
Zhaoxin Fan, Puquan Pan, Zeren Zhang, Ce Chen, Tianyang Wang, Siyang Zheng, Min Xu
- Introduces a two-stage unsupervised dense pre-training pipeline with segmentation-aware dense contrastive learning and superpixel guided pre-training, achieving state-of-the-art few-shot segmentation on Abd-CT and Abd-MRI.
Super Resolution

Structure-Preserving Super Resolution with Gradient Guidance
Cheng Ma, Yongming Rao, Yean Cheng, Ce Chen, Jiwen Lu, Jie Zhou
- Proposes a structure-preserving super resolution method that exploits gradient maps to guide image recovery. Introduces a gradient branch and gradient loss for second-order structural constraints.
- Cited 380+ times.
🎖 Honors and Awards
- 2026.03 🎓 Melbourne Research Scholarship
- 2026.03 🎓 Rowden White Scholarship
📖 Educations

🎓 Ph.D.
📚 School of Mathematics and Statistics, Faculty of Science
🏫 University of Melbourne (UniMelb)
📅 Mar.2026 – Now
👨🏫 Supervisor: Prof. Mingming Gong (宫明明)

🎓 M.Phil.
📚 Computer & Information Engineering, School of Science and Engineering
🏫 The Chinese University of Hong Kong, Shenzhen (CUHK-SZ)
📅 Aug.2021 – Jul.2023
👨🏫 Supervisor: Prof. Xiaoguang Han (韩晓光) | GPA: 3.98 / 4.0 (Rank 1/32)

🎓 B.Eng.
📚 Department of Automation
🏫 Tsinghua University (THU)
📅 Aug.2017 – Jul.2021
👨🏫 Supervisor: Prof. Yipeng Li (李一鹏) | GPA: 3.44 / 4.0
💻 Internships

💼 AI Research Intern
🏢 HeyGen
📍 Singapore
📅 Jun.2025 – Now
📝 Mentor: Mr. Yi Ren (任意). Research on accurate and detailed human video generation via diffusion models.

💼 Research Assistant
🏢 Machine Learning Department, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
📍 Abu Dhabi, UAE
📅 Feb.2024 – Jun.2025
📝 Supervisor: Prof. Mingming Gong (宫明明) & Prof. Kun Zhang (张坤). Research on text-to-4D generation, on-device text-to-image, and virtual try-on.

💼 Research Assistant
🏢 School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen (CUHK-SZ)
📍 Shenzhen, Guangdong, China
📅 Aug.2023 – Jan.2024
📝 Supervisor: Prof. Xiaoguang Han (韩晓光). Research on photorealistic 3D human head generation with spherical tri-plane representation.

💼 Computer Vision Algorithm Engineer Intern
🏢 WeChat Group, Tencent
📍 Beijing, China
📅 Jun.2021 – Aug.2021
📝 Developed table recognition system converting camera-captured or PDF table images to structured HTML, deployed as part of the company’s e-office system.

💼 Computer Vision Algorithm Engineer Intern
🏢 AI Lab, ByteDance
📍 Beijing, China
📅 Jun.2020 – Nov.2020
📝 Developed apartment layout diagram vectorization system for structured room parsing, achieving >90% recognition accuracy with >5× efficiency improvement.
📋 Academic Services
- 📖 Conference Reviewer: NeurIPS 2026, ECCV 2026, CVPR 2026, ICLR 2026, NeurIPS 2025
- 📖 Journal Reviewer: IEEE Transactions on Multimedia, IEEE Signal Processing Letters, IEEE Transactions on Image Processing