Agentic embodied navigation

SuperNav An Agentic Navigation System
for Any Task in Any Scene

  • 1Zhejiang University
  • 2Shenzhen University
  • 3Causa Robotics

†Corresponding author

TL;DR

  • SuperNav is an agentic navigation system for diverse tasks and scenes, from finding objects and visiting multiple targets to fulfilling high-level requests.
  • A pretrained multimodal model directs navigation through tool use, task-progress tracking, and context management, without navigation-specific fine-tuning.
  • Experiments show higher success rates than the evaluated baselines on single-object, multi-object, and demand-driven navigation, alongside real-world deployment on a quadruped robot.
Overview Video

Method

A pretrained multimodal model decides where to go, and SuperNav turns each decision into motion. No navigation-specific fine-tuning.

Figure 1
SuperNav teaser: searching an unfamiliar home for a sofa, with observations, reasoning and navigation paths
Finding a sofa in an unfamiliar home. The agent rechecks candidates, recovers from a blocked doorway, and verifies completion.
Figure 2
SuperNav system overview
System overview.

Results

Single-object · Habitat-GS · 150 tasks
Single-object, Habitat-GS · 150 tasks
MethodSR ↑SPL ↑
NaVid24.670.1599
UniNaVid34.000.1833
StreamVLN13.330.1036
OmniNav (Action Former)27.330.2133
SuperNav78.000.4127

BibTeX

BibTeX
@misc{zhang2026supernavagenticnavigationtask,
  title={SuperNav: An Agentic Navigation System for Any Task in Any Scene},
  author={Jinkai Zhang and Jingyi Xu and Yuanhong Yu and Jiarui Guo and Ruizhen Hu and Hujun Bao and Xiaowei Zhou and Sida Peng},
  year={2026},
  eprint={2610.12126},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2610.12126},
}