Unified Panorama-Language-Action Model for Instruction-Guided Navigation and Dynamic Person Tracking

Pengfei Qi1, Haoran Lin2, Sizhuang Chen2, Kai Luo2, Sirui Zhang2, Xinqi Liu2, Fei Cheng3,4, Wenrui Chen2, Liming Yin4, and Kailun Yang1,2,†
1College of Semiconductors (College of Integrated Circuits), Hunan University, Changsha, China.
2School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, Changsha, China.
3School of Advanced Technology, Xi'an Jiaotong-Liverpool University, China.
4Suzhou VSDeep Intelligent Technology Co., Ltd., China.
†Corresponding author: Kailun Yang.

Panoramic perception. A new benchmark. Two tasks. One unified model.

5,000Tracking Trajectories
10,000Navigation Routes
919,978Waypoint Samples
Real-World Dataset
96Routes
9.13 kmTotal Distance
924Local Segments
36,267Waypoint Samples

OVERVIEW VIDEO Real-world demonstrations of person tracking and navigation across diverse environments.

Abstract

General-purpose embodied robots should support both navigation toward language-specified destinations and dynamic person tracking under arbitrary initial target azimuths. However, existing methods typically rely on forward-facing observations and address these tasks with separate policies, limiting omnidirectional perception and unified closed-loop control. We present UniTrackPLA, a unified panorama-language-action model for instruction-guided navigation and dynamic person tracking.

Its Panoramic-Aware Encoding (PAE) preserves the temporal and azimuthal structure of perspective views projected from each panorama, enabling perspective-pretrained visual encoders to process omnidirectional observations. A shared vision-language backbone grounds instructions in the panoramic context and predicts continuous robot-centric waypoint chunks for both tasks. World-Action Consistency (WAC) further predicts action-conditioned future visual states and verifies waypoint prefixes online, allowing reliable actions to be reused while triggering replanning upon inconsistency.

We also introduce OmniTrackNav-Bench, comprising 5,000 simulated tracking trajectories, 10,000 simulated VLN routes, and 96 verified real-world routes, providing 919,978 waypoint-supervision instances. UniTrackPLA improves overall tracking SR from 23.50% to 35.00% and Omni-VLN SR/SPL from 13.00%/12.77% to 19.75%/19.29%. Incorporating 76 real-world routes further improves held-out EP@0.2m from 42.92% to 92.08%. Closed-loop experiments on a Go2-W robot demonstrate unified panoramic tracking and navigation across indoor and outdoor environments.

+11.50 pp

Overall Tracking SR

23.50% → 35.00%

+6.75 pp

Omni-VLN SR

13.00% → 19.75%

+49.16 pp

Held-Out EP@0.2m

42.92% → 92.08%

Contributions

01

Unified Panoramic VLA

UniTrackPLA unifies dynamic person following and language-guided navigation through shared panoramic perception, vision-language representations, and continuous waypoint actions; PAE jointly models azimuth and temporal context.

02

World-Action Consistency

WAC predicts the future latent effects of waypoint prefixes, adaptively regulating the action horizon and deciding whether to continue execution or trigger replanning.

03

OmniTrackNav-Bench

A unified benchmark supports joint training and consistent evaluation across panoramic tracking and navigation, with extensive simulation and real-robot validation.

Framework

UniTrackPLA framework: panoramic perception, PAE, WAC, and waypoint prediction

Overview of the proposed UniTrackPLA pipeline: the upper policy branch encodes four perspective views with PAE and fuses them with language instructions to predict waypoint chunks, while the lower WAC branch predicts action-conditioned future features and evaluates future-state consistency to continue execution or trigger replanning; snowflakes denote frozen modules.

Experiment

Omni-Tracking Comparison

Person tracking results on STF, DRF, and overall benchmarks

Evaluation set: 200 episodes · 100 STF + 100 DRF

Omni-VLN Comparison

Navigation results on R2R-CE and RxR-CE seen and unseen splits

Evaluation set: 400 routes · 100 per split

REAL-WORLD NAVIGATION DATA PERFORMANCE

Performance with increasing numbers of real-world training routes

Evaluation set: 20 route-disjoint held-out routes

Ablation Study

PANORAMIC VIEW CONFIGURATIONS

Comparison of panoramic view configurations

Validation set: 60 episodes · 30 STF + 30 DRF

ABLATION STUDY OF PAE AND WAC

Ablation study of PAE and WAC components

Evaluation set: 200 episodes · 100 STF + 100 DRF

Real-World Demos

01 Person Tracking