Spatial Intelligence Lab

Research

Six research pillars, from visual geometry to embodied AI — all building toward machines that understand and act in 3D space.

00

Overview

The Spatial Intelligence Lab develops models and algorithms that let machines perceive, reason about, and act in 3D space. Our core toolkit spans visual geometry, including correspondence, pose estimation, structure-from-motion, and 3D reconstruction, together with equivariant representation learning that keeps perception stable under viewpoint and rotation changes. Building on this foundation, we push toward 3D visual foundation models, multi-modal LLMs, and embodied AI, with applications ranging from industrial robotics and autonomous driving to wearable AR/VR systems and human-centric AI.

01

Computer Vision & Visual Geometry

Recovering 3D structure from 2D images via correspondence, pose, geometry, reconstruction, and equivariance.

Computer Vision & Visual Geometry

Our research asks how a machine can recover the geometry of the physical world from nothing more than a handful of 2D images. We build vision systems that fuse classical multi-view geometry with modern representation learning, treating correspondence, pose estimation, multi-view reasoning, 3D reconstruction, and equivariance as five faces of the same underlying problem. Our long-term goal is a geometrically grounded perception stack that any downstream system — a robot, an AR headset, an autonomous vehicle, or an embodied agent — can trust to localise itself, understand object pose, and reconstruct unseen geometry from sparse views.

Visual CorrespondenceCamera & Object PoseMulti-View Geometry3D Reconstruction

Read the research ideas

02

Multi-Modal Large Language Models

Building multi-modal LLMs that see, hear, reason, and act across images, audio, video, and 3D scenes.

Multi-Modal Large Language Models

We ask how large language models can be taught to perceive, reason, and decide across the same rich mix of modalities — images, audio and speech, video, and 3D scenes — that humans handle effortlessly. We work along four interlocking threads: building open multi-modal foundation models, eliciting reliable chain-of-thought reasoning that uses both pixel and acoustic evidence, designing post-training recipes that couple SFT with preference optimization (DPO, GRPO), and injecting 3D and geometry-aware representations into the multi-modal stack.

Multi-Modal Foundation ModelsChain-of-Thought ReasoningPost-Training (DPO, GRPO)3D-Aware MLLMs

Read the research ideas

03

Geometric Representation Learning

Symmetry, invariance, and equivariance as inductive biases for geometrically aware representations.

Geometric Representation Learning

Modern foundation models excel at semantics yet remain surprisingly fragile under the geometric transformations that govern the physical world — rotations, scale changes, viewpoint shifts, and 3D rigid motions. We design group-equivariant and self-supervised representation learning frameworks in which transformations of the input induce predictable, structured transformations of the feature space, studying invariance, equivariance, spherical and SO(3) harmonics, and symmetry-aware self-supervision as a single, unified toolkit.

Invariance & EquivarianceSO(3) / Spherical HarmonicsSelf-Supervised LearningGeometric Deep Learning

Read the research ideas

04

3D Visual Foundation Models

Feed-forward 3D foundation models that recover geometry, correspondence, and renderable scenes from images.

3D Visual Foundation Models

Can a neural network look at a handful of unposed images and immediately understand the 3D world they depict? We study a new generation of 3D visual foundation models — feed-forward transformers that jointly predict camera parameters, dense point maps, correspondences, and renderable representations such as 3D Gaussians directly from pixels. Building on the DUSt3R, MASt3R, VGGT, and π³ line of work, we investigate stronger geometric inductive biases, permutation- and viewpoint-equivariance, dynamic-scene extensions, and unified outputs that bind geometry, appearance, semantics, and uncertainty.

DUSt3R / VGGTPoint Maps & 3D GaussiansFeed-Forward ReconstructionNovel-View Synthesis

Read the research ideas

05

Robotics for Real-World Physical AI

Building embodied agents that perceive in 3D, reason with language, and act reliably in the physical world.

Robotics for Real-World Physical AI

How can an autonomous agent perceive, reason about, and act in the unstructured physical world with the same fluency that modern foundation models exhibit on the web? We unify 3D geometric perception, language-conditioned policy learning, and large-scale simulation into a single stack, with explicit attention to equivariance, multimodal grounding, and closing the sim-to-real gap — integrating perception, planning, control, learning, and vision-language-action so a single agent can map raw sensor streams to robust closed-loop behavior.

Embodied AIVision-Language-ActionSim-to-RealManipulation

Read the research ideas

06

Application Research, AI+X

Spatial AI as a core engine for AI+X across AR/VR, driving, healthcare, infrastructure, and animal behavior.

Application Research, AI+X

We treat spatial intelligence as a shared core engine and pursue AI+X collaborations with domain experts across many verticals: on-device spatial perception for AR/VR, planning-aware perception and occupancy world models for autonomous driving, trustworthy 3D segmentation for medical imaging, spatio-temporal and foundation-model forecasting for transportation and construction demand, and 3D vision for animal behavior — all sharing a backbone of multi-sensor fusion across camera, LiDAR, depth, and IMU, with robustness, calibration, and domain shift as first-class objectives.

AR / VRAutonomous DrivingMedical ImagingAnimal Behavior

Read the research ideas