Visual language navigation method and system based on dynamic grid map
By constructing a dynamic grid map in visual language navigation, combining the CLIP model, cascading attention mechanism and Mamba module, the existing methods are solved in the problem of insufficient map representation and real-time reduction of long-distance navigation in continuous environments, and accurate visual language navigation is achieved.
Patent Information
- Application Number
- CN202510884175.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-30
AI Technical Summary
The existing visual language navigation methods are difficult to construct map representations that can accurately represent environmental geometric and semantic information, and can dynamically adapt to environmental changes in a continuous environment. In long-distance navigation, it is difficult to deal with timing dependence and dynamic environmental changes, resulting in a decline in real-time and disconnection between semantic understanding and environmental perception.
The visual language features of RGB panoramic images are extracted using the CLIP model, and combined with the depth features of the depth images to build a dynamic grid map. The cross-modal interaction feature analysis is performed through the cascade attention mechanism and the Mamba module, and action decisions are made in combination with the DD-PPO strategy to achieve real-time update and accurate navigation of the map.
It realizes rapid response and accurate visual language navigation in a dynamic continuous environment, improves the fineness of environmental representation and task adaptability, and solves the problem of real-time decline caused by long-distance dependence and the disconnection between semantic understanding and environmental perception.
Smart Images

Figure CN120385352A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of navigation, and particularly relates to a vision-language navigation method and system based on a dynamic grid map. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] With the rapid development of artificial intelligence and robotics technologies, endowing an intelligent agent with the ability to understand natural language instructions and autonomously complete navigation tasks in a complex three-dimensional environment, namely vision-and-language navigation (VLN), has become one of the research hotspots and core challenges in the field of embodied AI. VLN not only requires the intelligent agent to have strong visual perception and natural language understanding capabilities, but also needs to be able to perform effective cross-modal information fusion, environmental memory, path planning, and decision-making.
[0004] However, although existing vision-language navigation methods have made significant progress, in a continuous environment, there are still some technical problems in existing vision-language navigation methods for intelligent agents, such as: (1) Existing environmental maps mainly include topological maps and fixed semantic maps. Among them, although topological maps can provide global path guidance, they lack environmental details (such as object geometries and semantic categories), resulting in the inability to support fine-grained navigation decisions; while fixed semantic maps are based on fixed semantic labels, so it is difficult to generalize to complex scenes containing unseen objects. Therefore, existing methods cannot construct a map representation that can accurately represent environmental geometric and semantic information and dynamically adapt to environmental changes according to navigation instructions.
[0005] (2) In long-distance navigation tasks, existing methods are difficult to effectively handle temporal dependencies and dynamic environmental changes. For example, Transformer-based models (such as BEVBert) rely on self-attention mechanisms, and the computational complexity increases quadratically with the sequence length, resulting in a decrease in real-time performance in long-instruction scenarios and thus making it difficult to adapt to continuous dynamic environmental changes.
[0006] (3) Some existing methods also rely on large language models (such as MC-GPT, NavGPT) for planning, but lack deep fusion with visual features, resulting in a disconnection between semantic understanding and environmental perception. Summary of the Invention
[0007] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a visual language navigation method and system based on a dynamic grid map, which can quickly respond to dynamic continuous environmental changes on the basis of avoiding long-distance dependencies, so as to achieve accurate visual language navigation for an agent.
[0008] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions: The first aspect of the present invention provides a visual language navigation method based on a dynamic grid map.
[0009] A visual language navigation method based on a dynamic grid map includes: Obtain the RGB panoramic image and depth image at the current moment; Use the CLIP model to extract the visual language features of the RGB panoramic image, and combine the obtained visual language features and the depth features of the depth image with the absolute coordinates of the grid cells to construct a grid map; update the grid map according to the semantic relevance between the grid cells and the current navigation instruction; Based on the cascaded attention mechanism, process the map features of the updated grid map and the text features of the current navigation instruction to obtain cross-modal interaction features; use the Mamba module to analyze the obtained cross-modal interaction features to output predicted waypoints; Use DD-PPO as the local policy, and use the obtained predicted waypoints as input for action analysis to generate a probability distribution; select and execute navigation actions according to the obtained probability distribution.
[0010] Further, combining the visual language features and the depth features of the depth image with the absolute coordinates of the grid cells to construct a grid map includes: using the visual language features extracted from the RGB panoramic image by the CLIP model as grid features, and adjusting the size of the corresponding depth features to match the grid features; calculating the absolute coordinates according to the grid features and the depth features.
[0011] Further, combining the visual language features and the depth features of the depth image with the absolute coordinates of the grid cells to construct a grid map further includes: storing the grid features and the absolute coordinates in the grid cells to realize the fusion of semantic information and geometric information to construct a grid map.
[0012] Further, updating the grid map according to the semantic relevance between the grid cells and the current navigation instruction includes: performing cosine similarity matching between the current navigation instruction and the feature vector under the grid cell to obtain attention weights; performing weighted fusion on the map features of the grid map according to the obtained attention weights to update the grid map.
[0013] Further, the feature vector under the grid cell is: the feature vectors observed at the same position and different perspectives at a set time step.
[0014] Further, the Mamba module is used to analyze the cross-modal interaction features to output predicted waypoints, including: inputting the obtained cross-modal interaction features into the Mamba module, processing the long sequence dependencies through a selective state space mechanism to output time-series enhanced features; generating predicted waypoints through decoding by a fully connected layer according to the obtained time-series enhanced features.
[0015] Further, DD-PPO is adopted as the local policy, and action analysis is performed with the predicted waypoints as the input to generate a probability distribution, including: using the coordinates of the obtained predicted waypoints as sub-goals to input into the distributed proximal policy optimization module, and generating the probability distribution of discrete navigation actions based on the policy generation network in the distributed proximal policy optimization module.
[0016] The second aspect of the present invention provides a visual language navigation system based on a dynamic grid map.
[0017] A visual language navigation system based on a dynamic grid map includes: An image acquisition module, configured to: acquire an RGB panoramic image and a depth image at the current moment; A map construction module, configured to: extract visual language features of the RGB panoramic image using a CLIP model, and combine the obtained visual language features and depth features of the depth image with the absolute coordinates of the grid cells to construct a grid map; update the grid map according to the semantic relevance between the grid cells and the current navigation instruction; A waypoint prediction module, configured to: process the map features of the updated grid map and the text features of the current navigation instruction based on a cascaded attention mechanism to obtain cross-modal interaction features; analyze the obtained cross-modal interaction features using the Mamba module to output predicted waypoints; An action decision module, configured to: adopt DD-PPO as the local policy, perform action analysis with the obtained predicted waypoints as the input to generate a probability distribution; select and execute a navigation action according to the obtained probability distribution. The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the steps in a visual language navigation method based on a dynamic grid map as described in the first aspect of the present invention are implemented.
[0018] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the steps in a visual language navigation method based on a dynamic grid map as described in the first aspect of the present invention are implemented.
[0019] The above one or more technical solutions have the following beneficial effects: (1) The present invention uses the CLIP model to extract visual language features of RGB panoramic images, and combines the obtained visual language features and depth features of depth images with the absolute coordinates of grid cells to construct a grid map. Through this method of constructing a dynamic visual language grid map (DGVL-Map), by using the CLIP model to extract visual language features of open vocabulary and fusing the spatial geometric information of depth images, each grid cell can store semantic features and three-dimensional coordinates simultaneously, so as to solve the problems that the topological map lacks details and the fixed semantic map has insufficient generalization. In addition, the present invention realizes real-time dynamic update of the map by calculating the semantic correlation (attention weight) between grid cells and navigation instructions, can preferentially strengthen the regional features related to the task, enable the map representation to be dynamically adjusted according to the instruction requirements, and thus significantly improve the fineness and task adaptability of the environmental representation. Therefore, the present invention can construct a map representation that can accurately represent environmental geometric and semantic information and dynamically adapt to environmental changes according to navigation instructions.
[0020] (2) The present invention processes the grid map features and text features of navigation instructions based on a cascaded attention mechanism to obtain cross-modal interaction features; then uses the Mamba module to analyze the cross-modal interaction features to obtain predicted waypoints. By introducing a cascaded attention mechanism based on instructions and the Mamba model, and using the selective state space module of Mamba to achieve long sequence modeling with linear time complexity, the problem that the computational complexity increases with the square of the sequence length due to the self-attention mechanism in the existing method architectures can be avoided. This mechanism dynamically aggregates historical observations and instruction features through incremental state updates, can capture long-distance dependencies and ensure real-time navigation efficiency, and can effectively solve the problem of decreased real-time performance in long instruction scenarios, so as to better adapt to continuous dynamic environmental changes.
[0021] (3) The cascaded attention mechanism designed by the present invention uses the instruction features extracted by BERT as key-value pairs and the visual language features of the grid map as queries, and realizes cross-modal depth fusion through multi-head attention and the Mamba model. This mechanism quantifies the semantic correlation between instructions and visual features in a unified embedding space, guides the model to focus on the environmental areas related to the task, can solve the problem of the disconnection between large language model planning and visual perception, enables accurate alignment of semantic understanding and environmental features, and thus improves the accuracy of waypoint prediction and the robustness of navigation decisions.
[0022] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0024] Figure 1 It is a flowchart of a visual language navigation method based on a dynamic grid map in Embodiment 1 of the present invention.
[0025] Figure 2 It is an architecture diagram of a visual language navigation system based on a dynamic grid map in Embodiment 2 of the present invention.
[0026] Figure 3 It is a structural diagram of a cross-attention mamba network in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0028] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.
[0029] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0030] The overall idea proposed by the present invention: The present invention provides a visual language navigation method based on a dynamic grid map. First, a pre-trained CLIP model is used to extract visual language features of RGB panoramic images and fuse the corresponding depth information to construct a grid map containing rich semantic and geometric information. Subsequently, a method for real-time dynamic update of the grid map is designed. By calculating the semantic correlation (attention weight) between grid cells and navigation instruction features, the area related to the navigation task is preferentially updated to achieve effective weighted fusion of map features. In terms of waypoint prediction, in order to efficiently extract key features from the grid map containing a large amount of possibly irrelevant information, a cascade attention mechanism based on instructions is innovatively introduced, and combined with the powerful long-sequence processing ability of the Mamba model, the map features highly related to the instructions are aggregated, so as to output accurate waypoint predictions. Finally, the action decision adopts the DD-PPO strategy to generate navigation actions according to the predicted waypoints.
[0031] Example 1 This example discloses a vision-language navigation method based on a dynamic grid map.
[0032] As Figure 1 shown, a vision-language navigation method based on a dynamic grid map includes: Step S1, obtain the RGB panoramic image and depth image at the current moment; Step S2, use the CLIP model to extract the vision-language features of the RGB panoramic image, and combine the obtained vision-language features and the depth features of the depth image with the absolute coordinates of the grid cells to construct a grid map; update the grid map according to the semantic correlation between the grid cells and the current navigation instruction; Step S3, process the map features of the updated grid map and the text features of the current navigation instruction based on the cascaded attention mechanism to obtain cross-modal interaction features; use the Mamba module to analyze the obtained cross-modal interaction features to output the predicted waypoints; Step S4, use DD-PPO as the local policy, and perform action analysis with the obtained predicted waypoints as the input to generate a probability distribution; select and execute navigation actions according to the obtained probability distribution.
[0033] Based on the above method, the present invention can quickly respond to dynamic continuous environmental changes on the basis of avoiding long-distance dependencies to achieve accurate vision-language navigation for the agent. For the convenience of understanding the technical solution of the present invention, the following further explains and illustrates the specific implementation method of the technical solution of the present invention.
[0034] In step S1, the RGB panoramic image and depth image at the current moment are obtained.
[0035] The agent is equipped with an RGB camera in the navigation environment (such as the Habitat simulator) to capture the panoramic visual information of the current node, that is, the RGB panoramic image; at the same time, it is also equipped with a depth camera to synchronously capture the depth information of the current node with the RGB camera. Moreover, each single-view depth image collected by the depth camera corresponds one-to-one with the viewing angle of the RGB image.
[0036] Although significant progress has been made in VLN in continuous environments with existing methods, there are still some key issues, especially in directions such as the agent's perception and understanding of the environment and the improvement of multimodal fusion technology in continuous environments. The present invention is directed to the representation of the navigation environment and the design of navigation planning. Based on the following steps, by dynamically constructing a grid map that fuses visual features and instruction information, and combining a waypoint prediction mechanism based on the Mamba model, it aims to improve the agent's understanding of complex scenarios, long-term memory, and the efficiency and robustness of task execution in continuous environments. Specifically: In step S2, the CLIP model is used to extract the visual-language features of the RGB panoramic image, and the obtained visual-language features and the depth features of the depth image are combined with the absolute coordinates of the grid cells to construct a grid map; according to the semantic relevance between the grid cells and the current navigation instruction, the grid map is updated. Specifically, it can be achieved through the following methods: Step S2-1, Grid map construction.
[0037] The visual-language features extracted from the RGB panoramic image by the CLIP model are used as grid features, and the size of the corresponding depth features is adjusted to match the grid features; subsequently, the absolute coordinates are calculated based on the grid features and the depth features. Specifically, the RGB panoramic image is denoted as: ; where represents the RGB panoramic image, represents the panoramic image observed by the agent at time step ; represents the serial number of the component of the RGB panoramic image.
[0038] For the obtained RGB panoramic image, the pre-trained CLIP-ViT-B / 32 model is used to extract the visual-language features of the RGB panoramic image, that is, the grid features: ; where represents the grid features of the RGB panoramic image, represents the grid features of the th view or region at time step represents the set of real numbers, represents the height dimension, represents the width dimension, represents the channel dimension. The grid features of the h-th row and w-th column are denoted as .
[0039] The corresponding depth image is downsampled to the same scale as the RGB panoramic image and denoted as , to ensure that the depth value is aligned with the grid features extracted from RGB in the spatial dimension; where represents the corresponding depth value after resizing the depth image. Correspondingly, the depth value (depth feature) at the h-th row and w-th column after resizing is denoted as . For convenience, all the indices are denoted as , ranges from 1 to I, and . Therefore, the grid feature is re-denoted as , and the depth feature after resizing is re-denoted as . On this basis, the absolute coordinates of the grid feature can be calculated, that is: ; where represents the apex angle between the current position of the agent and the grid feature , represents the Euclidean distance between the agent and the grid feature ; and are the abscissa and ordinate of the agent in the two-dimensional space respectively, used to represent its position in the global coordinate system (i.e., the current coordinates of the agent).
[0040] After that, all grid features and their absolute coordinates are stored in the grid cells to achieve the fusion of semantic and geometric features for constructing a grid map, that is: ; where represents the grid memory bank at time step , which stores the visual features extracted by the agent during navigation and their corresponding absolute coordinates.
[0041] Based on the above steps, the present invention designs a dynamic grid visual language map, obtains the semantic understanding ability of open vocabulary through the CLIP model, and is no longer limited to predefined labels; at the same time, according to the visual language semantic information extracted by the CLIP model and the spatial geometric information extracted from the depth image, it can provide a finer-grained environmental representation for the agent. Furthermore, it gets rid of the dependence on predefined labels in the prior art and enhances the generalization ability in the position environment.
[0042] Step S2-2: Update the grid map according to the semantic relevance between the grid cells and the current navigation instruction.
[0043] Update the grid map based on the real-time dynamic update algorithm, that is: ; wherein, represents the number of accumulated features in the current grid cell, and the variable represents the global feature at the h-th row and w-th column of the grid map; represents at time step the feature vectors observed from the same position but different perspectives; represents the attention weight. To quantify the semantic correlation between the grid cell and the navigation instruction, the feature vector and the navigation instruction feature are matched in the unified embedding space through cosine similarity to obtain the attention weight , which is used to guide the weighted fusion and update of the map features, that is: ; This update mechanism can preferentially process the areas related to the navigation task. The visual instruction information is combined with the observation data from the new perspective to enhance the task alignment. By averaging the features in each grid cell, while maintaining the attention to the task-related areas, it can also effectively fuse the multi-perspective observation information of the same object.
[0044] Based on the above steps, the present invention combines the construction of the visual language map and the grid map on the basis of the visual language map, and can dynamically update the map content with emphasis according to the instruction information during the navigation process, thereby effectively avoiding the dependence limitation of the existing methods on the statically constructed map. In addition, by designing the attention weight based on the current instruction to dynamically update and fuse the map grid features, the task relevance can be more directly incorporated into the map representation.
[0045] In step S3, based on the cascaded attention mechanism, the map features of the updated grid map and the text features of the current navigation instruction are processed to obtain cross-modal interaction features; and the obtained cross-modal interaction features are analyzed using the Mamba module to output the predicted waypoints.
[0046] The constructed grid map is gradually updated as the agent gradually visits all environments. To understand the navigation progress, the present invention designs an instruction localization module for predicting waypoints by combining visual information at the current time step. Due to the complexity of the navigation environment, each grid cell stores a large amount of grid features, but most of these features may be irrelevant to the navigation task. Therefore, the agent needs key grid information highly relevant to the navigation instruction to understand the environment. In addition, the agent needs to quickly complete the planning prediction. Thus, the present invention proposes a cascaded attention mechanism based on instructions, using the Mamba model with lower time complexity to aggregate key features from the grid to output the predicted waypoints. Specifically, a text encoder and a visual encoder are used to process the text information of the navigation instruction and the encoded information of the image respectively, that is: A. Text encoder: Use the BERT model to process the navigation instruction to extract text features, and use to represent the encoded feature of the th word.
[0047] B. Visual encoder: Use a pre-trained convolutional neural network (ConvNet) to encode the observed RGB-D information to obtain RGB features and depth features 。
[0048] Use the text feature T in the text encoder as the key matrix and value matrix V, while the grid map feature is used as the query matrix q. The Mamba attention mechanism then calculates the map features most relevant to the current instruction . The map feature is still represented as a matrix with the shape of H × W × D, capturing the spatial and target context relevant to the task. The detailed process is as follows: First, the grid features and text features are processed through the multi-head attention mechanism MHA (Multi-Headed Attention). MHA calculates the feature correlations in different subspaces in parallel through multiple "attention heads", enabling the model to simultaneously focus on the semantic information at different positions in the input sequence and improving the cross-modal feature interaction efficiency; for the input query, key, and value, the attention weights are calculated through the dot product. Specifically, the grid map feature is used as the query , the instruction feature is used as the key and value V, and thus an intermediate representation is obtained, that is: ; where, Represents a combined operation of residual connection (Add) and layer normalization (Norm), that is, the residual connection adds the input query to the output of MHA to solve the problem of gradient degradation in the training of deep networks; layer normalization normalizes the feature dimensions to stabilize the training process. Represents the multi-head attention mechanism, which calculates the query and the key semantic similarity to generate the map feature weights most relevant to the instruction, and aggregate the information based on the weighted sum .
[0049] Then, it is processed by FFN for . FFN (Feed Forward Network) is a multi-layer neural network whose neurons only transmit signals unidirectionally between adjacent layers without feedback connections. In the Transformer architecture, FFN is usually connected after the multi-head attention mechanism (MHA) to perform further feature transformation on the attention output; its standard form contains two layers of linear mapping (fully connected layers) with a non-linear activation function inserted in the middle, and its formula can actually be expressed as: FFN(Z)=ReLU(xW1+b1)W2+b2; where, W1, W2 are weight matrices, b1, b2 are bias terms, and the ReLU activation function is used to introduce non-linearity. Specifically, it is processed by FFN for , that is: ; where, represents the feature sequence after processing , that is, the cross-modal interaction feature. Represents the feed-forward operation, that is: first, project the input into a high-dimensional space through the weight matrix W1 to enhance the feature expression ability; after activation by the ReLU activation function, project it back to the weight matrix W2 of the original dimension to achieve non-linear transformation and compression of the features. Through high-dimensional projection and dimensionality reduction, FFN can enhance key features (such as the visual language features of task-related objects), suppress noise, and thus improve the discriminability of the features.
[0050] The Mamba module is used to analyze the cross-modal interaction features to output the predicted waypoints, that is: input the obtained cross-modal interaction features into the standard Mamba (SSM) module, process the long sequence dependencies through the selective state space mechanism to output the time series enhanced features; according to the obtained time series enhanced features, generate the predicted waypoints through the fully connected layer decoding, specifically: The Mamba module utilizes its selective state space mechanism to scan sequences (i.e., cross-modal interaction features) to capture long-range dependencies and context information. At this time, the output of the Mamba (SSM) module is denoted as . The output feature sequence has the same dimension as the input cross-modal interaction feature . This process is expressed as: ; To interface the feature sequence encoded by the Mamba module with the subsequent cross-attention mechanism, it is necessary to generate the query ( ), key ( ), and value ( ) required for the cross-attention layer from the feature sequence . Inspired by the parameter generation of the self-attention mechanism in the Transformer architecture, the present invention uses independent linear projections to achieve this transformation. Specifically, the present invention defines three learnable linear transformation matrices for generating the query, key, and value respectively, namely: ; ; ; wherein, , , and represent the learnable linear transformation matrices for generating the query, key, and value respectively, , , and represent the target dimensions of the query, key, and value respectively.
[0051] Furthermore, the generation process of the attention parameters is as follows: ; ; ; wherein, , , and are optional bias terms. After these transformations, the query sequence , key sequence , and value sequence can be obtained.
[0052] When generating the query sequence , key sequence , and value sequence After that, it is passed as input to the subsequent cross-attention module to achieve the interaction and information aggregation between different modalities or different hierarchical features, that is: ; Among them, represents the feature after cross-attention fusion, represents the cross-attention mechanism fusion operation.
[0053] The combination of the Mamba module and cross-attention represents the cascaded processing of cross-modal depth feature fusion, that is: ; Among them, represents the Mamba module of the th layer, which is used to capture long sequence dependencies; the final output is the map feature matrix . After obtaining the map feature matrix most relevant to the current instruction, an area-based pooling mechanism is applied to reduce the dimension of the map feature matrix by focusing on the areas most similar to the current task. The pooling operation converts the map feature matrix from the matrix to the pooled feature vector , that is: ; Among them, represents the area pooling operation, which is used to compress the spatial dimension and retain key features. Finally, the pooled feature vector is concatenated with the RGB feature extracted by the visual encoder and the depth feature to form the input feature vector of the current time step, that is: ; The present invention uses the SSM in the Mamba model to model the historical trajectory context. Different from the traditional RNN architecture, Mamba can more efficiently capture long-distance temporal dependencies through a selective state space dynamic update mechanism with linear time complexity. The state space-based Mamba module significantly improves the computational efficiency and the breadth of feature expression while retaining the temporal modeling ability, and is especially suitable for continuous sub-goal generation tasks in large-scale semantic scenarios.
[0054] Specifically, the time step input sequence is input into the Mamba module for temporal modeling: ; Among them, and The hidden states of the current step and the previous step respectively. The Mamba module models temporal dependencies through incremental state space updates. At each time step t , the agent only inputs the current input features and the hidden state of the previous step into Mamba, and dynamically updates the state of the current step to . This mechanism avoids recomputing the historical sequence, with a time complexity of O (1), significantly improving the real-time navigation efficiency. Subsequently, the hidden state output by Mamba is mapped to the predicted two-dimensional coordinates of the sub-goal at the current time step through a fully connected layer, that is: ; where, represents the predicted two-dimensional coordinates of the sub-goal at the current time step, that is, the predicted waypoint; represents the weight matrix, represents the bias term of the fully connected layer.
[0055] To obtain the ground truth waypoint coordinates, the following operations are performed: First, the waypoint LAW (Language Alignment Waypoint) supervision method based on language alignment is used for the visual-language navigation task in a continuous environment to solve the problem of inconsistency with instructions in the existing goal-oriented supervision method when dealing with off-path scenarios. Specifically, the nearest waypoint is selected according to LAW, that is: the language-aligned path is different from the shortest path to the target. Especially when the agent deviates from the reference path, LAW supervision encourages the agent to move towards the nearest waypoint on the language-aligned path at each step, rather than directly towards the target.
[0056] Subsequently, the CMA (Cross-Modal Attention) algorithm is used to plan the shortest path from the current agent position to the nearest waypoint on the instruction-related path; the CMA algorithm model contains two recurrent networks, one for encoding the agent's state history and the other for predicting actions based on attention-based visual and instruction features. The CMA algorithm model takes natural language instructions and RGB-D images as inputs to provide perception and decision-making capabilities for planning the shortest path. Based on the trained CMA algorithm model, when the agent receives a natural language instruction, the model will calculate the shortest path to the nearest waypoint on the language-aligned path according to the current agent position (perceiving the position by obtaining environmental information through RGB-D images).
[0057] Finally, draw a circle with the agent as the center and a radius of 3 meters, and the intersection points of the path and the circle are regarded as the true path points.
[0058] Furthermore, the mean squared error loss (MSE) is used to optimize the predicted sub-goal position , that is: ; Among them, represents the waypoint prediction loss, which is used to measure the error between the predicted waypoint and the true waypoint; represents the time step of the true waypoint coordinates.
[0059] In step S4, DD-PPO is adopted as the local policy, and the obtained predicted route points are used as inputs for action analysis to generate a probability distribution; and navigation actions are selected and executed according to the obtained probability distribution.
[0060] Adopt DD-PPO as the local policy, and use the route point as the input to generate the probability distribution of four discrete actions. DD-PPO is a distributed algorithm developed on the basis of proximal policy optimization (PPO). As a policy gradient method, PPO aims to achieve data efficiency and reliable performance similar to trust region policy optimization (TRPO) using first-order optimization; DD-PPO then realizes distributed training on this basis, allowing multiple machines to participate in the training process simultaneously to improve training efficiency. Action selection is achieved through the policy network and optimized through reinforcement learning. Specifically, the calculation method of action probability is as follows: ; Among them, represents the action probability distribution; represents the time step of the hidden state, generated by mamba, containing the context information of historical observations and instructions; represents that the Softmax function normalizes the Mamba output to produce a valid probability distribution.
[0061] Sample or select actions from the action probability distribution as follows: ; Among them, represents the high-level discrete action (such as forward, left turn, right turn, stop) executed by the agent at time , represents the high-level discrete action; represents the conditional probability distribution based on the hidden state where the agent selects the high-level discrete action .
[0062] To optimize the local policy, the present invention also designs two loss functions, namely the multi-class cross-entropy loss function for measuring the difference between the true navigable actions and the predicted action probabilities, and the binary cross-entropy loss function for evaluating the difference between the true stop action and the predicted stop probability. The weighted sum of these two loss functions is used as the total loss function. Specifically: A. During the training process of the multi-class cross-entropy loss function, the model adjusts its parameters according to this loss value to make the predicted action probabilities closer to the probability distribution of the true actions, guiding the intelligent agent to make correct high-level navigation decisions, such as choosing the correct forward and turning directions, etc. The multi-class cross-entropy loss function is expressed as: ; Among them, represents the multi-class cross-entropy loss function, represents the true navigable action.
[0063] B. The binary cross-entropy loss function is used to help the model better judge when to stop moving to avoid missing the target or over-moving. The binary cross-entropy loss function is expressed as: ; Among them, represents the binary cross-entropy loss function, represents the true stop action, represents the predicted stop probability.
[0064] After obtaining the accurate probability distribution, the navigation action can be selected and executed according to the obtained probability distribution, that is, directly select the navigation action with the highest probability for execution.
[0065] Based on the visual language navigation method based on a dynamic grid map provided by the present invention, first, the construction of a dynamic visual language grid map (DGVL-Map) is carried out. This method can fuse the visual language features and depth information extracted by CLIP and perform real-time dynamic updates according to navigation instructions, enhancing the task relevance and environmental adaptability of the map representation. Secondly, a cascade attention mechanism based on instructions is designed, which can aggregate the key information in the grid and filter out the map features related to the instructions to improve the navigation efficiency. Finally, a waypoint prediction method combined with the Mamba model is proposed, which can predict sub-goals according to the input, effectively improving the ability to extract key information from complex map features and perform long-distance dependence modeling.
[0066] Embodiment 2 This embodiment discloses a visual language navigation system based on a dynamic grid map.
[0067] A visual language navigation system based on a dynamic grid map includes: An image acquisition module, configured to: obtain an RGB panoramic image and a depth image at the current moment; A map construction module, configured to: extract visual - language features of the RGB panoramic image using a CLIP model, and combine the obtained visual - language features and depth features of the depth image with the absolute coordinates of grid cells to construct a grid map; update the grid map according to the semantic correlation between the grid cells and the current navigation instruction; A waypoint prediction module, configured to: process the map features of the updated grid map and the text features of the current navigation instruction based on a cascaded attention mechanism to obtain cross - modal interaction features; analyze the obtained cross - modal interaction features using a Mamba module to output predicted waypoints; An action decision - making module, configured to: adopt DD - PPO as a local policy, take the obtained predicted waypoints as input for action analysis to generate a probability distribution; select and execute a navigation action according to the obtained probability distribution.
[0068] As Figure 2 shown, a visual - language navigation system based on a dynamic grid map, based on its overall architecture design: First, process the RGB panoramic image observed by the agent based on the CLIP model, and integrate the depth - based absolute coordinates to construct a dynamic visual - language grid map. Then, the dynamic visual - language grid map and the text - encoded instruction are input into the CAM module for processing to identify the map features most relevant to the instruction. Finally, these features are combined with the visual information of the agent and input into the Mamba model for waypoint prediction and action decision - making accordingly.
[0069] As Figure 3 shown, the Cross - Attention Mamba Network (CAM) consists of two parallel branches: one is the instruction features extracted by the text encoder, and the other is the map features. CAM uses the text features as keys (K) and values (V), and the map features as queries (Q) to calculate the map features most relevant to the current instruction.
[0070] Based on the above systematic architecture design, the present invention can quickly respond to dynamic continuous environmental changes on the basis of avoiding long - distance dependencies to achieve precise visual - language navigation for the agent.
[0071] Embodiment III The purpose of this embodiment is to provide a computer - readable storage medium.
[0072] A computer - readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a visual - language navigation method based on a dynamic grid map as described in Embodiment I of the present disclosure.
[0073] Embodiment 4 The purpose of this embodiment is to provide an electronic device.
[0074] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in a visual language navigation method based on a dynamic grid map as described in Embodiment 1 of the present disclosure.
[0075] The steps involved in the devices of the above Embodiments 2, 3, and 4 correspond to those of Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to execute any method in the present invention.
[0076] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0077] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, this is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.
Claims
1. A visual language navigation method based on a dynamic grid map, characterized in that, Including: Obtain the RGB panoramic image and depth image at the current moment; Use the CLIP model to extract the visual - language features of the RGB panoramic image, and combine the obtained visual - language features and the depth features of the depth image with the absolute coordinates of the grid cells to construct a grid map; Update the grid map according to the semantic correlation between the grid cells and the current navigation instruction; Based on the cascaded attention mechanism, process the map features of the updated grid map and the text features of the current navigation instruction to obtain cross - modal interaction features; use the Mamba module to analyze the obtained cross - modal interaction features to output predicted waypoints; Adopt DD - PPO as the local policy, take the obtained predicted waypoints as input for action analysis to generate a probability distribution; select and execute navigation actions according to the obtained probability distribution.
2. The visual language navigation method based on a dynamic grid map according to claim 1, wherein Combining the visual - language features and the depth features of the depth image with the absolute coordinates of the grid cells to construct a grid map includes: using the visual - language features extracted from the RGB panoramic image by the CLIP model as grid features, adjusting the size of the corresponding depth features to match the grid features; calculating the absolute coordinates according to the grid features and the depth features.
3. A visual language navigation method based on a dynamic grid map according to any one of claims 1-2, characterized in that, Combining the visual - language features and the depth features of the depth image with the absolute coordinates of the grid cells to construct a grid map further includes: storing the grid features and the absolute coordinates in the grid cells, realizing the fusion of semantic information and geometric information to construct a grid map.
4. The visual language navigation method based on a dynamic grid map according to claim 1, characterized in that, Updating the grid map according to the semantic correlation between the grid cells and the current navigation instruction includes: performing cosine similarity matching between the current navigation instruction and the feature vectors under the grid cells to obtain attention weights; performing weighted fusion on the map features of the grid map according to the obtained attention weights to update the grid map.
5. The visual language navigation method based on a dynamic grid map according to claim 4, wherein The feature vectors under the grid cells are: the feature vectors observed at the same position and different perspectives under a set time step.
6. The visual language navigation method based on a dynamic grid map according to claim 1, wherein Using the Mamba module to analyze the cross - modal interaction features to output predicted waypoints includes: inputting the obtained cross - modal interaction features into the Mamba module, processing long - sequence dependencies through a selective state - space mechanism to output time - series enhanced features; generating predicted waypoints by decoding through a fully - connected layer according to the obtained time - series enhanced features.
7. A visual language navigation method based on a dynamic grid map according to claim 1, characterized in that, Adopting DD - PPO as the local policy, taking the predicted waypoints as input for action analysis to generate a probability distribution includes: taking the coordinates of the obtained predicted waypoints as sub - goals and inputting them into the distributed proximal policy optimization module, and generating a probability distribution of discrete navigation actions based on the policy generation network within the distributed proximal policy optimization module.
8. A visual language navigation system based on a dynamic grid map, characterized in that, Including: An image acquisition module, configured to: obtain the RGB panoramic image and depth image at the current moment; A map construction module, configured to: use the CLIP model to extract the visual - language features of the RGB panoramic image, and combine the obtained visual - language features and the depth features of the depth image with the absolute coordinates of the grid cells to construct a grid map; Update the grid map according to the semantic correlation between the grid cells and the current navigation instruction; The waypoint prediction module is configured to: process the map features of the updated grid map and the text features of the current navigation instruction based on the cascaded attention mechanism to obtain cross-modal interaction features; use the Mamba module to analyze the obtained cross-modal interaction features to output predicted waypoints. The action decision-making module is configured to: adopt DD-PPO as the local policy, perform action analysis with the obtained predicted waypoints as the input to generate a probability distribution; select and execute navigation actions according to the obtained probability distribution.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a visual language navigation method based on a dynamic grid map as described in any one of claims 1-7.
10. An electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a visual language navigation method based on a dynamic grid map as described in any one of claims 1-7.
Citation Information
Patent Citations
Visual language navigation system and method based on dynamic enhancement instruction attack module
CN113804200A
Visual language navigation method combining image description and text generation image
CN117571014A
Visual language navigation method based on double semantic graphs and modal alignment
CN117889864A
Unmanned aerial vehicle visual language navigation method based on multi-modal perception Mama
CN119063736A
Visual language navigation method fusing grid map and topological graph
CN119164385A
Cited By
Communication chip resource dynamic scheduling method based on deep reinforcement learning
CN121255721A
Robot map asynchronous establishment method, electronic equipment and storage medium
CN122170853A
Robot map asynchronous establishment method, electronic device and storage medium
CN122170853B
A zero-fine-tuning visual language robot navigation method based on label-enhanced multi-modal reasoning
CN122775093A