Pengjun Fang
Incoming MPhil Student, The Hong Kong University of Science and Technology
Advised by Prof. Qifeng Chen
I work on multimodal generation, video-to-audio synthesis, and controllable video generation.

Pengjun Fang
About
I am an incoming MPhil student in Computer Science at HKUST, advised by Prof. Qifeng Chen. I received my B.Eng. in Computer Science from HKUST, and spent an exchange semester at EPFL, Switzerland.
My research focuses on generative AI, with an emphasis on multimodal generation, video-to-audio synthesis, and controllable video generation. I am interested in building models that connect perception across modalities and give users precise, interpretable control over generated content.
News
- 2026.06Joining HKUST as an MPhil student, advised by Prof. Qifeng Chen.
- 2026.01AC-Foley accepted to ICLR 2026.
- 2025.08Finished research internship at Everlyn Labs Inc.
- 2025.07Text-Driven Portrait Image Animation accepted to ICCV 2025 Workshop.
- 2025.02Started exchange semester at EPFL, Switzerland.
Publications
* Equal contribution † Corresponding author
- AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer[ICLR 2026]Pengjun Fang*, Yingqing He, Yazhou Xing, Qifeng Chen†, Ser-Nam Lim†, Harry Yang†
- AnimationBench: Are Video Models Good at Character-Centric Animation?[arXiv 2025]Leyi Wu*, Pengjun Fang*, Kai Sun*, Yazhou Xing, Yinwei Wu, Songsong Wang, Ziqi Huang, Dan Zhou, Yingqing He, Ying-Cong Chen, Qifeng Chen
- Text-Driven Portrait Image Animation for Controllable Facial Dynamics[ICCV Workshop 2025]Jiaxin Xie*, Pengjun Fang*, Yingqing He, Qiang Wen, Zhefan Rao, Liya Ji, Yazhou Xing†, Qifeng Chen†
- MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions[arXiv 2024]Xiaowei Chi*, Aosong Cheng*, Yatian Wang*, Pengjun Fang*, Zeyue Tian*, Yingqing He, Xingqun Qi, Zhaoyang Liu, Rongyu Zhang, Qifeng Chen, Wenhan Luo, Qifeng Liu, Wei Xue, Shanghang Zhang, Yike Guo
Experience
- Everlyn Labs Inc. — Research InternFeb 2025 – Aug 2025
Developed AC-Foley, a reference-audio-guided video-to-audio synthesis system. Built WanFM, a First–Last–Frame-to-Video pipeline based on Wan 2.2 with bidirectional denoising for stronger temporal consistency.
Selected Projects
- VideoTuna— Open-source codebase for text-to-video generation
Unified codebase integrating multiple AI video generation models across text-to-video, image-to-video, and text-to-image. Provides end-to-end pipelines for pre-training, continual training, post-training alignment, and fine-tuning.
- WanFM— First–Last–Frame-to-Video generation
Built on Wan 2.2 Image-to-Video with last-frame constraints, bidirectional denoising, and prompt-adapted attention for controllable FLF2V generation.
Contact
Feel free to reach out by email at pfangaf@connect.ust.hk for research discussions or collaborations.