Mu Cai

I am a Member of Technical Staff at Thinking Machines Lab, working on multimodal pretraining science and visual agents.

Previously, I was a Research Scientist at Google DeepMind DeepMind logo working on Gemini multimodal and visual agents (image2code).

I got my Ph.D. in Computer Sciences from University of Wisconsin-Madison, advised by Prof. Yong Jae Lee.

Mu Cai profile photo
Selected Publications
Gemma 4
Gemma 4 Technical Report
Gemma Team:   ..., Mu Cai, ...
arXiv, 2026
[arXiv] [Page]

Matryoshka Multimodal Models
Matryoshka Multimodal Models
Mu Cai, Jianwei Yang, Jianfeng Gao, Yong Jae Lee
Proceedings of the International Conference on Learning Representations (ICLR), 2025
[arXiv] [code] [Project Page] [Demo]

LLaVA-PruMerge
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
Yuzhang Shang*, Mu Cai*, Bingxin Xu, Yong Jae Lee†, Yan Yan†
IEEE/CVF International Conference on Computer Vision (ICCV), 2025 (*equal contribution, †equal advising)
[arXiv] [code] [Project Page]

ViP-LLaVA
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, Yong Jae Lee
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
[arXiv] [code] [Demo] [Project Page] [Youtube]

TemporalBench
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Mu Cai, Reuben Tan, Jianrui Zhang, Bocheng Zou, Kai Zhang, Feng Yao, Fangrui Zhu, Jing Gu, Yiwu Zhong, Yuzhang Shang, Yao Dou, Jaden Park, Jianfeng Gao†, Yong Jae Lee†, Jianwei Yang†
arXiv, 2024 (†equal advising)
[arXiv] [Project Page] [Code] [Datasets] [Leaderboard]

Magma
Magma: A Foundation Model for Multimodal AI Agents
Jianwei Yang*, Reuben Tan*, Qianhui Wu*, Ruijie Zheng†, Baolin Peng†, Yongyuan Liang†, Yu Gu, Mu Cai, Seonghyeon Ye, Joel Jang, Yuquan Deng, Lar Liden, Jianfeng Gao
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025 (*first author, †second author)
[arXiv] [code] [Project Page]

LLaRA
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
Xiang Li, Cristina Mata, Jongwoo Park, Kumara Kahatapitiya, Yoo Sung Jang, Jinghuan Shang, Kanchana Ranasinghe, Ryan Burgert, Mu Cai, Yong Jae Lee, and Michael S. Ryoo
Proceedings of the International Conference on Learning Representations (ICLR), 2025
[arXiv] [code]

Humanity's Last Exam
Humanity's Last Exam
Center for AI Safety, Scale AI
Nature, 2026
[paper]

Talks