Research Output

Publications

Research publications by Zebang Cheng in multimodal intelligence, affective computing, and related areas.

All Publications

Google Scholar ↗

Publications are grouped by year. My name is shown in bold; an asterisk (*) denotes equal contribution.

2026

  1. Emotion-LLaMAv2 and MMEVerse: A New Framework and Benchmark for Multimodal Emotion Understanding arXiv 2026

    Xiaojiang Peng, Jingyi Chen, Zebang Cheng, Bao Peng, Fengyi Wu, Yifei Dong, Shuyuan Tu, Qiyu Hu, Huiting Huang, Yuxiang Lin, Jun-Yan He, Kai Wang, Zheng Lian, Zhi-Qi Cheng

    A new framework and benchmark for advancing multimodal emotion understanding across perception and reasoning tasks.

  2. MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models ICLR 2026

    Fan Zhang*, Zebang Cheng*, Chong Deng*, Haoxuan Li, Zheng Lian, Qian Chen, Huadai Liu, Wen Wang, Yi-Fan Zhang, Renrui Zhang, Ziyu Guo, Zhihong Zhu, Hao Wu, Haixin Wang, Yefeng Zheng, Xiaojiang Peng, Xian Wu, Kun Wang, Xiangang Li, Jieping Ye, Pheng-Ann Heng

    * Equal contribution.

    A unified benchmark for evaluating emotional understanding and causal reasoning in multimodal large language models across diverse audio-visual scenarios.

  3. MER 2026: From Discriminative Emotion Recognition to Generative Emotion Understanding arXiv 2026

    Zheng Lian, Xiaojiang Peng, Kele Xu, Ziyu Jia, Xinyi Che, Zebang Cheng, Fei Ma, Laizhong Cui, Yazhou Zhang, Xin Liu, Liang Yang, Jia Li, Fan Zhang, Liumeng Xue, Erik Cambria, Guoying Zhao, Björn W. Schuller, Jianhua Tao

    A challenge report and research perspective on the transition from discriminative emotion recognition to open-ended generative understanding.

  4. EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs arXiv 2026

    He Hu, Tengjin Weng, Zebang Cheng, Yu Wang, Jiachen Luo, Björn W. Schuller, Zheng Lian, Laizhong Cui

    A benchmark focused on how emotions evolve over time and whether multimodal models can explain and predict those transitions.

  5. OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing arXiv 2026

    Zebang Cheng, Shuimu Chen, Boxue Yang, Yuanshen Guan, Jingyi Chen, Zheng Lian, Xiaojiang Peng, Fei Ma, Laizhong Cui, Qi Tian

    An on-policy self-distillation approach that uses privileged rationales to improve affective understanding and reasoning.

  6. Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding arXiv 2026

    Shuimu Chen, Yuteng Chen, Yuanshen Guan, Zebang Cheng, Zeyu Zhang, Shengqian Qin, Bin Xia, Jiaran Li, Wenming Yang, Fei Ma

    An evidence-driven reflection framework for detecting and correcting reasoning errors in long-video understanding.

  7. OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation arXiv 2026

    Shuimu Chen, Jing Jin, Nan Su, Hongbo Xu, Zebang Cheng, Wenming Yang, Fei Ma, Guijin Wang

    A self-distillation framework that transfers language-model judgment capabilities to industrial anomaly detection.

2025

  1. AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models ICML 2025

    Zheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun, Licai Sun, Yong Ren, Zebang Cheng, Bin Liu, Rui Liu, Xiaojiang Peng, Jiangyan Yi, Jianhua Tao

    Oral presentation.

    A unified data, model, and evaluation pipeline for generative multimodal emotion understanding, supported by the MER-Caption dataset and MER-UniBench.

  2. Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models CVPR Workshops 2025

    Yuxiang Lin, Jingdong Sun, Zhi-Qi Cheng, Jue Wang, Haomin Liang, Zebang Cheng, Yifei Dong, Jun-Yan He, Xiaojiang Peng, Xian-Sheng Hua

    An exploration of emotional reasoning in multimodal large language models beyond conventional emotion classification.

  3. HA-VLN: A Benchmark for Human-Aware Navigation in Discrete-Continuous Environments with Dynamic Multi-Human Interactions, Real-World Validation, and an Open Leaderboard arXiv 2025

    Yifei Dong, Fengyi Wu, Qi He, Heng Li, Minghan Li, Zebang Cheng, Yuxuan Zhou, Jingdong Sun, Qi Dai, Zhi-Qi Cheng, Alexander G. Hauptmann

    A benchmark for embodied agents that must navigate safely and socially around dynamic human participants.

  4. DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal Retrieval AAAI 2025

    Fan Zhang, Changhu Wang, Zebang Cheng, Xiaojiang Peng, Dongjie Wang, Yijia Xiao, Chong Chen, Xian-Sheng Hua, Xiao Luo

    A semi-supervised framework for aligning 2D images and 3D shapes through decoupled discriminative learning.

  5. EEG-VL: Integrating Visual Features with Large Language Models for Automated Seizure Detection BIBM 2025

    Zi Liang, Peng Hu, Lian Li, Zebang Cheng, Yisu Dong, Haibo He, Qiang Lu, Nan Lin

    A vision-language approach for representing EEG signals and supporting automated seizure detection.

  6. MER 2025: When Affective Computing Meets Large Language Models ACM Multimedia 2025

    Zheng Lian, Rui Liu, Kele Xu, Bin Liu, Xuefei Liu, Yazhou Zhang, Xin Liu, Yong Li, Zebang Cheng, Haolin Zuo, Ziyang Ma, Xiaojiang Peng, Xie Chen, Ya Li, Erik Cambria, Guoying Zhao, Björn W. Schuller, Jianhua Tao

    The MER 2025 challenge report connects established affective-computing tasks with the capabilities of large language models.

  7. Emotion-Qwen: Training Hybrid Experts for Unified Emotion and General Vision-Language Understanding arXiv 2025

    Dawei Huang, Qing Li, Chuan Yan, Zebang Cheng, Yurong Huang, Xiang Li, Bin Li, Xiaohui Wang, Zheng Lian, Xiaojiang Peng

    A hybrid-expert model designed to preserve general visual-language ability while strengthening fine-grained emotional reasoning.

  8. EMER-Ranker: Learning to Rank Emotion Descriptions in the Absence of Ground Truth arXiv 2025

    Zheng Lian, Licai Sun, Haoyu Chen, Zebang Cheng, Fan Zhang, Ziyu Jia, Ziyang Ma, Fei Ma, Xiaojiang Peng, Jianhua Tao

    A ranking-based evaluation approach for free-form emotion descriptions when a single definitive ground truth is unavailable.

2024

  1. Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning NeurIPS 2024

    Zebang Cheng*, Zhi-Qi Cheng*, Jun-Yan He, Jingdong Sun, Kai Wang, Yuxiang Lin, Zheng Lian, Xiaojiang Peng, Alexander Hauptmann

    * Equal contribution.

    An early audio-visual large language model for recognizing emotions and producing evidence-grounded reasoning through instruction tuning.

  2. MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models SemEval 2024

    Zebang Cheng*, Fuqiang Niu*, Yuxiang Lin, Zhi-Qi Cheng, Bowen Zhang, Xiaojiang Peng

    * Equal contribution. 3rd place in the Multimodal Emotion-Cause Analysis Challenge.

    A multimodal-language-model system for identifying emotions and their causal utterances in multi-party conversations.

  3. SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition MRAC at ACM Multimedia 2024

    Zebang Cheng*, Shuyuan Tu*, Dawei Huang*, Minghan Li, Xiaojiang Peng, Zhi-Qi Cheng, Alexander G. Hauptmann

    * Equal contribution. 1st place in MER-NOISE and 3rd place in the Open-Vocabulary Multimodal Emotion Recognition Challenge.

    A Conv-Attention extension of Emotion-LLaMA for robust multimodal emotion recognition under noise and open-vocabulary settings.

  4. Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model ACM Multimedia 2024

    Fuqiang Niu, Zebang Cheng, Xianghua Fu, Xiaojiang Peng, Genan Dai, Yin Chen, Hu Huang, Bowen Zhang

    A challenge dataset and model for recognizing stance in multimodal, multi-turn conversations.

  5. Dataset Growth ECCV 2024

    Ziheng Qin, Zhaopan Xu, Yukun Zhou, Zangwei Zheng, Zebang Cheng, Hao Tang, Lei Shang, Baigui Sun, Xiaojiang Peng, Radu Timofte, Hongxun Yao, Kai Wang, Yang You

    A study of strategies for growing training datasets and improving model performance with scalable data construction.

2023

  1. Semi-Supervised Multimodal Emotion Recognition with Expression MAE ACM Multimedia 2023

    Zebang Cheng, Yuxiang Lin, Zhaoru Chen, Xiang Li, Shuyi Mao, Fan Zhang, Daijun Ding, Bowen Zhang, Xiaojiang Peng

    2nd place in the ACM Multimedia 2023 Multimodal Emotion Recognition Challenge.

    A semi-supervised framework that uses masked autoencoding to learn expressive multimodal representations from limited annotations.