Research

What I work on in trustworthy multimodal AI.

My research focuses on trustworthy multimodal AI, particularly large vision-language models, with an emphasis on privacy-aware reasoning, adversarial robustness, and long-horizon visual understanding.

Privacy-Aware Vision-Language Reasoning

I study how vision-language models can remain useful while reasoning about privacy-sensitive visual information in a more controlled and reliable way.

  • Privacy boundaries and information disclosure.
  • Privacy leakage in multimodal reasoning.
  • Privacy-utility trade-offs and privacy-aware alignment.

Adversarial Robustness & Security

I study both adversarial attacks and defenses for vision and vision-language models, with a focus on understanding model vulnerabilities and improving robustness.

  • Transferable attacks across models and vision encoders.
  • Adversarial purification and detection.
  • Robustness and generalization under adversarial perturbations.

Long-Horizon Video Understanding

I study how multimodal systems can understand and reason over extended egocentric video histories while maintaining useful and privacy-aware behavior.

  • Long-horizon and egocentric VideoQA.
  • Retrieval and reasoning over extended visual context.
  • Utility and privacy in long-form visual understanding.

Research Questions

Questions that guide the work

Each question maps to one of the three research areas.

What should a model not reveal?

How can LVLMs reason about scenes without leaking sensitive identity, location, relationship, or behavioral information?

How does multimodal AI fail under attack?

What attacks transfer across encoders and models, and what defenses preserve accuracy without destroying useful visual content?

How should systems reason over long video histories?

How can multimodal systems handle long-horizon and egocentric VideoQA while balancing utility, privacy, and reliability?

Research Areas

Project-oriented view

These are public-facing project areas. Specific unpublished titles, internal figures, experimental results, and submission details are intentionally left out until they are public.

Privacy-Aware Vision-Language Reasoning

Privacy boundaries, information disclosure, privacy leakage in multimodal reasoning, and privacy-utility trade-offs.

Adversarial Robustness & Security

Transferable attacks across models and vision encoders, adversarial purification, detection, and robust generalization.

Long-Horizon Video Understanding

Long-horizon and egocentric VideoQA, retrieval and reasoning over extended visual context, and utility/privacy trade-offs.

Methods & Tools

Technical foundation

The technical stack combines research engineering, ML experimentation, and software development experience.

Models & Methods

TransformersLVLMsDiffusion ModelsAdversarial MLMultimodal Learning

Frameworks

PyTorchTensorFlowKerasscikit-learnOpenCV

Engineering

PythonC++JavaReactSPARQLRDFoxGitLaTeX