Vision-Language-Navigation in dynamic, human-centric environments exposes a fundamental tension: linguistic reasoning is slow and deliberative, whereas safe, socially compliant planning must be instant and reactive. The resulting observation staleness is safety-critical — a maneuver chosen during inference can already be unsafe by the time it executes. We observe that, long before a VLM finishes its inference, its intermediate hidden states already encode action-relevant intent. We propose SPARK-VLN, a dual-system framework in which a slow VLM reasoner streams its knowledge to a fast flow-matching expert planner token-by-token. We also introduce a human-centric benchmark suite for dynamic social VLN that keeps pedestrians and the robot active throughout inference. Across these settings, SPARK-VLN improves navigation success and social compliance while sustaining inference efficiency.
Three modules carry the VLM reasoner's evolving knowledge into the flow-matching expert planner without waiting for inference to complete.
Unlike static protocols, our suite keeps pedestrians and the robot active throughout inference, systematically varying three axes to expose the observation staleness that conventional benchmarks hide.
@article{hu2026sparkvln,
title = {SPARK-VLN: Dynamic Social Vision-Language Navigation},
author = {Hu, Tianshuai and Zhong, Yangyi and Gong, Zeying and Kong, Lingdong
and Mei, XiaoDong and Zhao, Guoyang and Li, Rong and Liang, Junwei},
journal = {Under Review},
year = {2026}
}