Research directions in multimodal large language models, affective computing, and multimodal agents.
Research Directions
I study how multimodal systems can perceive, reason about, and respond to human states in realistic environments.
Models
Multimodal Large Language Models for Emotional Intelligence
I develop multimodal large language models that jointly interpret speech, visual behavior, language, and context. The goal is to move affective computing beyond closed-set classification toward systems that can explain what a person may be feeling, identify supporting evidence, and communicate their reasoning in natural language.
Audio-visual instruction tuning and multimodal fusion
Fine-grained emotion recognition and explanation
Unified models that retain general vision-language ability
Emotion Understanding, Reasoning, and Benchmarking
Reliable emotional intelligence requires evaluation beyond a single accuracy score. My work builds datasets, benchmarks, and evaluation protocols that measure recognition, causal reasoning, temporal transitions, robustness, and free-form descriptions across realistic multimodal scenarios.
Unified benchmarks for emotional understanding and reasoning
I explore agents that perceive long-term multimodal context, reflect on evidence, and act in environments involving people. This direction connects multimodal reasoning with navigation, egocentric interaction, self-correction, and human-aware decision making.