I am Kairui Hu, currently a Ph.D. student at CCDS, NTU, fortunate to be supervised by Prof. Ziwei Liu. I am also an AI Scientist at Ropedia.

I was a Founding Team Member at Synvo AI, supervised by Prof. Chen Change Loy. Prior to this, I was a Core Contributor at LMMs-Lab, where I was on an exciting journey towards Large Multimodal Models and feeling the AGI. I have received kind and valuable guidance from Yuanhan Zhang, Bo Li, and Jingkang Yang. I received my B.Eng. in Computer Science from Nanyang Technological University in 2024 with First Class Honours (Highest Distinction).

My research focuses on agents, multimodal models, and egocentric AI.

📧   Feel free to reach me at HUKA0001@e.ntu.edu.sg.

News

  • 2026.06:   HippoCamp is accepted to ECCV 2026. Congrats to all coauthors!
  • 2026.04:   Video-MMMU is accepted to ACL 2026 Main Conference.
  • 2026.04:   Released FileGram — a privacy-first AI memory layer.
  • 2026.03:   HippoCamp is out — we collected 3 real people’s entire digital lives to build the first file-system memory dataset.
  • 2026.02:   OpenMMReasoner is accepted to CVPR 2026.
  • 2026.01:   New blog post at Synvo AIThe Digital Avalanche: Building the Memory Layer for Next-Gen Corporation AI Agents.
  • 2025.04:   Aero-1-Audio — our first generation of lightweight audio models, outperforming Whisper and Qwen-2-Audio.
  • 2025.01:   Released the Video-MMMU benchmark — selected as the only video benchmark in the official public releases of OpenAI GPT-5 and Google Gemini 3.0 & 3.1.
  • 2025.01:   LMMs-Eval accepted to NAACL 2025 Findings.
  • 2025.01:   Joined Synvo AI as a Founding Team Member.
  • 2024.08:   Joined MMLab@NTU as a Research Staff.
  • 2024.07:   Released LMMs-Eval — a comprehensive benchmark for evaluating Large Multimodal Models.
  • 2024.06:   Led the release of lmms-eval v0.2.0 to support video evaluations.
  • 2024.06:   Graduated from NTU with First Class Honours (Highest Distinction).

Selected Publications

Video-MMMU

Video-MMMU: Evaluating Knowledge Acquisition from Multidisciplinary Professional Videos

Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Xiang Yue, Bo Li, Yuanhan Zhang, Ziwei Liu

ACL 2026 (Main)

A video reasoning benchmark for LMMs, evaluating knowledge acquisition from multi-discipline professional videos. Featured in OpenAI GPT-5 and Google Gemini 3.0 official releases. Adopted by Google DeepMind, OpenAI, Alibaba, ByteDance, and many others.

OpenMMReasoner

OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe

Kaichen Zhang, Keming Wu, Zuhao Yang, Bo Li, Kairui Hu, Bin Wang, Ziwei Liu, Xingxuan Li, Lidong Bing

CVPR 2026

An open and general recipe for pushing the frontiers of multimodal reasoning, achieving strong performance across comprehensive multimodal reasoning benchmarks.

HippoCamp

HippoCamp: Benchmarking Contextual Agents on Personal Computers

Zhe Yang, Shulin Tian, Kairui Hu, Shuai Liu, Hoang-Nhat Nguyen, Yichi Zhang, Zujin Guo, Mengying Yu, Zinan Zhang, Jingkang Yang, Chen Change Loy, Ziwei Liu

ECCV 2026

A benchmark for contextual agents on personal computers, built from three real users’ complete digital lives — the first file-system memory dataset.

LMMs-Eval

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Kaichen Zhang, Bo Li, Peiyuan Zhang, Fanyi Pu, Joshua Adrian Cahyono, Kairui Hu, Shuai Liu, Yuanhan Zhang, Jingkang Yang, Chunyuan Li, Ziwei Liu

NAACL 2025 (Findings)

A One-for-All Multimodal Evaluation Toolkit across Text, Image, Video, and Audio tasks — a one-command framework that makes LMM evaluation easier, more convenient, and reproducible. Widely adopted across the GenAI community for model development and benchmarking.

Preprints

FileGram

FileGram: Grounding Agent Personalization in File-System Behavioral Traces

Shuai Liu, Shulin Tian, Kairui Hu, Yuhao Dong, Zhe Yang, Bo Li, Jingkang Yang, Chen Change Loy, Ziwei Liu

arXiv 2026

Grounding agent personalization in file-system behavioral traces — a privacy-first AI memory layer learned from how you actually use your files.

Software

Local Cocoa

Local Cocoa

A privacy-focused local AI assistant that turns your files into searchable semantic memory — running entirely on your device.

Experience

Research Staff, MMLab@NTU
Supervised by Prof. Ziwei Liu and Prof. Chen Change Loy  ·  Aug 2024 – Jul 2026

  • Multimodal language models, video understanding, benchmarking and evaluation.

Founding Engineer, Synvo AI, Singapore
Jan 2025 – Jul 2026

  • Architected and implemented agentic file system search and reasoning.
  • Local Cocoa: a local AI assistant running on your device that turns your files into actionable memory.

Core Member, LMMs-Lab, Singapore
Aug 2024 – Present

Honors and Awards

  • 2024: NTU Information Technology Management Association (ITMA) Gold Medal cum Book Prize
  • 2022, 2023: NTU President Research Scholar (with Merit), URECA Undergraduate Research Programme
  • 2020-2022: Dean’s List (Top 5%), College of Computing and Data Science, NTU
  • 2019-2024: NTU Science and Engineering Undergraduate Scholarship (SM2), Ministry of Education Singapore

Education

  • 2020.08 - 2024.06, B.Eng. in Computer Science, Nanyang Technological University, Singapore. First Class Honours (Highest Distinction).

Miscellaneous

  • 🎹   I love music — I play piano and guitar, and I’m a fan of Chinese folk music, especially Zhao Lei.
  • 🎵   I’m genuinely into music creation and songwriting.
  • 🏀   Sports: basketball, badminton, marathon, and swimming.
  • 🧗   Currently learning bouldering and surfing — still very much a rookie!