A memory-enhanced embodied visual language navigation method and related devices
By extracting features and back-projecting from RGB images and depth maps collected by the embodied robot, a hierarchical three-dimensional memory structure is constructed, which solves the problem of insufficient global scene understanding in embodied visual language navigation, realizes long-term navigation and efficient search in complex large-scale scenes, and improves the success rate and efficiency of navigation tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
- Filing Date
- 2026-05-12
- Publication Date
- 2026-07-21
AI Technical Summary
Existing embodied visual language navigation technology lacks global scene understanding and long-term spatial memory in large-scale environments, leading to short-sighted decision-making and repetitive exploration, making it difficult to achieve fine near-field positioning and efficient far-field exploration, resulting in a low success rate for long-term autonomous navigation missions.
By acquiring RGB images and depth maps collected by the embodied robot, preprocessing and feature extraction are performed, and backprojection is used to obtain 3D patch features. These features are then aggregated into voxels and stored in a hierarchical 3D memory structure map. Combined with navigation instructions, a task-oriented query is generated to retrieve features from the near-field and far-field memory regions and generate navigation actions.
It achieves global modeling and long-term information retention of large-scale environments, taking into account both target localization and space exploration, improving the success rate and efficiency of navigation tasks, and preserving key geometric and semantic details.
Smart Images

Figure CN122176663B_ABST