How do we build better intelligence?
We believe progress comes from two complementary fronts. Better generative models and agents give systems richer capabilities; better guidance and alignment make those capabilities controllable, reliable, and useful.
Our work spans the full stack—from theory and algorithms to real-world applications, with natural collaboration points for researchers working on LLMs, diffusion/flow models, Agents, Physical AI, AI4Science, and industrial problems.
01
Generative AI & Agents
We seek gains that do not come solely from scaling data or parameter counts. We study better modeling choices for efficient and coherent generation, including unified multimodal generative models and visual agents that can understand, decompose, and create complex artifacts.
02
Video Generation & World Models
We develop video generative models that are physically plausible, controllable, and capable of predicting how the world evolves. Our goal is to move beyond visually convincing clips toward models that capture geometry, motion, interaction, and world state.
03
Video Understanding
Video remains a major perceptual bottleneck for current models. We aim to help models understand events coherently over time through better representations, temporal reasoning, and alignment—rather than treating video as a loose collection of frames.
04
Inverse Problems & Computational Imaging
Many problems in science and engineering ask us to recover an unknown world from incomplete, noisy, or indirect measurements. We develop general theory and measurement-aware algorithms that use generative priors to solve these inverse problems reliably and efficiently.
05
Medical AI
We target concrete challenges in modern medicine, including medical image reconstruction, medical report generation, segmentation, and digital therapeutics. We aim to combine strong general-purpose models with domain knowledge to improve the quality, accessibility, and efficiency of care.