A memory-enhanced embodied visual language navigation method and related devices

By extracting features and back-projecting from RGB images and depth maps collected by the embodied robot, a hierarchical three-dimensional memory structure is constructed, which solves the problem of insufficient global scene understanding in embodied visual language navigation, realizes long-term navigation and efficient search in complex large-scale scenes, and improves the success rate and efficiency of navigation tasks.

CN122176663BActive Publication Date: 2026-07-21HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
Filing Date
2026-05-12
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing embodied visual language navigation technology lacks global scene understanding and long-term spatial memory in large-scale environments, leading to short-sighted decision-making and repetitive exploration, making it difficult to achieve fine near-field positioning and efficient far-field exploration, resulting in a low success rate for long-term autonomous navigation missions.

Method used

By acquiring RGB images and depth maps collected by the embodied robot, preprocessing and feature extraction are performed, and backprojection is used to obtain 3D patch features. These features are then aggregated into voxels and stored in a hierarchical 3D memory structure map. Combined with navigation instructions, a task-oriented query is generated to retrieve features from the near-field and far-field memory regions and generate navigation actions.

Benefits of technology

It achieves global modeling and long-term information retention of large-scale environments, taking into account both target localization and space exploration, improving the success rate and efficiency of navigation tasks, and preserving key geometric and semantic details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122176663B_ABST
    Figure CN122176663B_ABST
Patent Text Reader

Abstract

The application discloses a memory-enhanced embodied visual language navigation method and related equipment, and the method comprises the following steps: extracting a two-dimensional patch feature set from a preprocessed RGB image; performing inverse projection on the two-dimensional patch feature to obtain a three-dimensional space position, and fusing the three-dimensional space position with the two-dimensional patch feature to obtain a three-dimensional patch feature; voxelizing and aggregating the three-dimensional patch feature to obtain an aggregated three-dimensional patch feature; storing the aggregated three-dimensional patch feature and its aggregated three-dimensional coordinates into a preset hierarchical three-dimensional memory structure graph to obtain a new hierarchical three-dimensional memory structure graph; retrieving three-dimensional patch features of different regions from the new hierarchical three-dimensional memory structure graph based on task-oriented query to obtain memory features; and generating a navigation action according to the memory features, a navigation instruction and the two-dimensional patch feature. The application can balance target positioning and space exploration, improve the scene understanding ability, exploration efficiency and task success rate of long-term autonomous navigation, and can be widely applied to the technical field of navigation.
Need to check novelty before this filing date? Find Prior Art