Research Agenda

Research

Research directions in multimodal large language models, affective computing, and multimodal agents.

Research Directions

I study how multimodal systems can perceive, reason about, and respond to human states in realistic environments.

Models

Multimodal Large Language Models for Emotional Intelligence

I develop multimodal large language models that jointly interpret speech, visual behavior, language, and context. The goal is to move affective computing beyond closed-set classification toward systems that can explain what a person may be feeling, identify supporting evidence, and communicate their reasoning in natural language.

  • Audio-visual instruction tuning and multimodal fusion
  • Fine-grained emotion recognition and explanation
  • Unified models that retain general vision-language ability

Data & Evaluation

Emotion Understanding, Reasoning, and Benchmarking

Reliable emotional intelligence requires evaluation beyond a single accuracy score. My work builds datasets, benchmarks, and evaluation protocols that measure recognition, causal reasoning, temporal transitions, robustness, and free-form descriptions across realistic multimodal scenarios.

  • Unified benchmarks for emotional understanding and reasoning
  • Evidence-grounded and temporal emotion analysis
  • Evaluation of open-ended emotion descriptions

Agents

Multimodal Agents and Human-Aware Intelligence

I explore agents that perceive long-term multimodal context, reflect on evidence, and act in environments involving people. This direction connects multimodal reasoning with navigation, egocentric interaction, self-correction, and human-aware decision making.

  • Egocentric and long-video multimodal reasoning
  • Evidence-driven reflection and self-correction
  • Human-aware navigation and embodied interaction