Research Output
Publications
Research publications by Zebang Cheng in multimodal intelligence, affective computing, and related areas.
All Publications
Google Scholar ↗Publications are grouped by year. My name is shown in bold; an asterisk (*) denotes equal contribution.
2026
-
Emotion-LLaMAv2 and MMEVerse: A New Framework and Benchmark for Multimodal Emotion Understanding arXiv 2026
A new framework and benchmark for advancing multimodal emotion understanding across perception and reasoning tasks.
-
MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models ICLR 2026
* Equal contribution.
A unified benchmark for evaluating emotional understanding and causal reasoning in multimodal large language models across diverse audio-visual scenarios.
-
MER 2026: From Discriminative Emotion Recognition to Generative Emotion Understanding arXiv 2026
A challenge report and research perspective on the transition from discriminative emotion recognition to open-ended generative understanding.
-
EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs arXiv 2026
A benchmark focused on how emotions evolve over time and whether multimodal models can explain and predict those transitions.
-
OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing arXiv 2026
An on-policy self-distillation approach that uses privileged rationales to improve affective understanding and reasoning.
-
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding arXiv 2026
An evidence-driven reflection framework for detecting and correcting reasoning errors in long-video understanding.
-
OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation arXiv 2026
A self-distillation framework that transfers language-model judgment capabilities to industrial anomaly detection.
2025
-
AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models ICML 2025
Oral presentation.
A unified data, model, and evaluation pipeline for generative multimodal emotion understanding, supported by the MER-Caption dataset and MER-UniBench.
-
Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models CVPR Workshops 2025
An exploration of emotional reasoning in multimodal large language models beyond conventional emotion classification.
-
HA-VLN: A Benchmark for Human-Aware Navigation in Discrete-Continuous Environments with Dynamic Multi-Human Interactions, Real-World Validation, and an Open Leaderboard arXiv 2025
A benchmark for embodied agents that must navigate safely and socially around dynamic human participants.
-
DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal Retrieval AAAI 2025
A semi-supervised framework for aligning 2D images and 3D shapes through decoupled discriminative learning.
-
EEG-VL: Integrating Visual Features with Large Language Models for Automated Seizure Detection BIBM 2025
A vision-language approach for representing EEG signals and supporting automated seizure detection.
-
MER 2025: When Affective Computing Meets Large Language Models ACM Multimedia 2025
The MER 2025 challenge report connects established affective-computing tasks with the capabilities of large language models.
-
Emotion-Qwen: Training Hybrid Experts for Unified Emotion and General Vision-Language Understanding arXiv 2025
A hybrid-expert model designed to preserve general visual-language ability while strengthening fine-grained emotional reasoning.
-
EMER-Ranker: Learning to Rank Emotion Descriptions in the Absence of Ground Truth arXiv 2025
A ranking-based evaluation approach for free-form emotion descriptions when a single definitive ground truth is unavailable.
2024
-
Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning NeurIPS 2024
* Equal contribution.
An early audio-visual large language model for recognizing emotions and producing evidence-grounded reasoning through instruction tuning.
-
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models SemEval 2024
* Equal contribution. 3rd place in the Multimodal Emotion-Cause Analysis Challenge.
A multimodal-language-model system for identifying emotions and their causal utterances in multi-party conversations.
-
SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition MRAC at ACM Multimedia 2024
* Equal contribution. 1st place in MER-NOISE and 3rd place in the Open-Vocabulary Multimodal Emotion Recognition Challenge.
A Conv-Attention extension of Emotion-LLaMA for robust multimodal emotion recognition under noise and open-vocabulary settings.
-
Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model ACM Multimedia 2024
A challenge dataset and model for recognizing stance in multimodal, multi-turn conversations.
-
Dataset Growth ECCV 2024
A study of strategies for growing training datasets and improving model performance with scalable data construction.
2023
-
Semi-Supervised Multimodal Emotion Recognition with Expression MAE ACM Multimedia 2023
2nd place in the ACM Multimedia 2023 Multimodal Emotion Recognition Challenge.
A semi-supervised framework that uses masked autoencoding to learn expressive multimodal representations from limited annotations.