A Variable Environment Visual Language Navigation Method and System Based on Historical Reflection

Through a historical reflection method, the large-model visual encoder and graph-perceived self-attention mechanism are used to dynamically adjust the navigation strategy, solving the adaptability and fault tolerance of visual language navigation under non-perfect instructions, and achieving efficient navigation in complex environments.

CN119826835BActive Publication Date: 2025-07-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510309117.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-25
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The existing visual language navigation technology lacks adaptability and fault tolerance in non-perfect instructions, making it difficult to effectively navigate in complex and dynamic real-world environments.

Method used

A method based on historical reflection is adopted, and a large-model visual encoder and query transformer are used to process visual observation data, combined with the gated network and the reflection network to analyze the navigation historical information, generate correction instructions, and path planning is carried out through the graph-perceived self-attention mechanism, and navigation strategies are dynamically adjusted.

Benefits of technology

It improves the robustness and adaptability of the navigation system in complex environments, enhances the fault tolerance of environmental changes, and ensures the stability and efficiency of navigation strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119826835B_ABST
    Figure CN119826835B_ABST
Patent Text Reader

Abstract

The present invention provides a variable environment visual language navigation method and system based on historical reflection, which relates to the technical fields of computer vision, natural language processing, and robot navigation. The method is to use a large model visual encoder and a query transformer to process variable environment visual observation data to obtain a scene graph encoding embedding, navigation history information used in the next action, and a corresponding encoded embedding instruction; use a gating network and a reflection network to analyze the scene graph encoding embedding, navigation history information used in the next action, and the corresponding encoded embedding instruction to obtain a corrected instruction; the reflection network is a large language model for generating corrected instructions; based on the corrected instruction, use a graph-aware self-attention mechanism to calculate and perform path navigation on the intelligent agent to obtain a visual language navigation result, thereby completing the visual language navigation in a variable environment. The present invention solves the problems of poor adaptability and low fault tolerance of visual language navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision, natural language processing and robot navigation technology, and in particular to a variable environment visual language navigation method and system based on historical reflection. Background Art

[0002] The Vision-and-Language Navigation (VLN) task aims to simulate human navigation behavior in a real environment. Its core goal is to enable an intelligent agent, that is, an entity that can autonomously perceive the environment and take actions to achieve a certain goal, to receive and understand natural language instructions, combine visual observations, and generate a reasonable sequence of action decisions to complete navigation tasks in an unknown environment. For example, the agent may receive an instruction: "Go downstairs, walk to the dining table, turn left to the kitchen, and stop in front of the refrigerator." Then, according to the content of the instruction, it needs to gradually complete the navigation path through visual perception and environmental interaction, and finally reach the target location. Due to the low cost, high controllability and scalability of the simulated environment to a variety of scenarios, many studies on vision-and-language navigation tasks rely on experiments based on simulated environments. However, migrating research results in simulated environments to real-world scenarios with higher complexity and uncertainty has always been one of the research focuses and technical difficulties in this field.

[0003] Existing visual language navigation technology frameworks are usually based on the "perfect instruction hypothesis", which assumes that natural language instructions can be perfectly aligned with the navigation path in the physical environment. This assumption assumes that the provider of the instructions has a full understanding of the navigation scenario and can generate clear and unambiguous instructions. However, in practical applications, this assumption is often not true because the instructions generated by the user are affected by a variety of subjective and objective factors and may not fully reflect the actual situation of the navigation environment. This deviation may cause the agent to face significant challenges when performing navigation tasks.

[0004] Existing research usually ignores the possible imperfect matching between natural language instructions and the physical environment, which leads to insufficient robustness of the agent in the real environment. How to design a more adaptive and fault-tolerant visual language navigation method in the case of imperfect instructions, combining the dynamic characteristics of the environment and the subjective bias of the instruction generator, is one of the core issues that need to be solved urgently. Summary of the invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides a method and system for variable environment visual language navigation based on historical reflection, which solves the problems of poor adaptability and low fault tolerance of visual language navigation.

[0006] To achieve the above invention objectives, the technical solution adopted by the present invention is as follows: A variable environment visual language navigation method based on historical reflection, including:

[0007] S1: Using a large model visual encoder and a query transformer, process the variable environment visual observation data to obtain a scene graph encoding embedding, navigation history information used in the next action, and corresponding encoded embedding instructions;

[0008] S2: Using a gating network and a reflection network, analyze the scene graph encoding embedding, the navigation history information used in the next action, and the corresponding encoded embedding instructions to obtain a correction instruction; wherein, the reflection network is a large language model for generating correction instructions;

[0009] S3: Based on the correction instruction, use the graph-aware self-attention mechanism for calculation, perform path navigation on the agent to obtain a visual language navigation result, and complete the visual language navigation of the variable environment.

[0010] Further, the S1 includes:

[0011] Using a frozen large model visual encoder, process the variable environment visual observation data to obtain image features; wherein, the frozen large model visual encoder is a large model visual encoder with its weights fixed in a lightweight manner;

[0012] Using a query transformer, perform self-attention mechanism fusion on the learnable query embedding and the instruction embedding to obtain instruction information;

[0013] Based on the cross-attention mechanism, fuse the image features and the instruction information to obtain a scene graph encoding embedding, navigation history information used in the next action, and corresponding encoded embedding instructions.

[0014] Further, the expression of the scene graph encoding embedding is:

[0015] ;

[0016] Wherein, represents the scene graph encoding embedding, represents the cross-attention mechanism, represents the image features, represents the instruction information.

[0017] Further, the S2 includes:

[0018] Concatenate the scene graph encoding embedding, the navigation history information used in the next action, and the corresponding encoded embedding instructions into sequence data, input it into the gating network, and obtain the output result of the gating network;

[0019] Input the output result of the gating network into a binary classification fully connected layer, and through calculation, obtain the error label value;

[0020] When the error label value is greater than the threshold, use the reflection network to identify the specific error type, and autoregressively generate a correction instruction through the large language model.

[0021] Furthermore, the expression of the error label value is:

[0022] ;

[0023] ;

[0024] where, represents the error label value, represents the standard deviation of, represents the binary classification fully connected layer, represents the output result of the gating network, represents the gating network, represents the scene graph encoding embedding, represents the encoding embedding instruction, represents the historical navigation information.

[0025] Furthermore, the expression of the error type is:

[0026] ;

[0027] ;

[0028] where, represents the error type, represents the maximum selection function, represents function, represents the first hyperparameter for error type recognition, represents the function value calculated according to the large language model, represents the second hyperparameter for error type recognition, represents the large language model, represents the scene graph encoding embedding, represents the encoding embedding instruction, represents the historical navigation information;

[0029] The expression of the correction instruction is:

[0030] ;

[0031] where, represents the correction instruction, Denote the large language model, Denote the function value calculated according to the large language model, Denote the error type.

[0032] Furthermore, the S3 includes:

[0033] Based on the action encoding sequence, unify the path points and visual features to obtain node embeddings;

[0034] Perform cross-modal cross-attention calculation on the correction instruction and the node embedding to obtain a cross-modal attention result;

[0035] Use variable environment visual observation data to judge the connectivity between nodes and construct a dynamic topology graph;

[0036] Combine the cross-modal attention result and the dynamic topology graph and perform a graph-aware self-attention mechanism to obtain a graph-aware result;

[0037] Use a feed-forward neural network to analyze the graph-aware result, and comprehensively consider the distance and difficulty of reaching the node to obtain an action score; where the action score is used to reflect the efficiency of the agent reaching this node;

[0038] Based on the action score, select the node with the highest score and navigate along the optimal path to this node to obtain a visual language navigation result and complete the visual language navigation in a variable environment.

[0039] Furthermore, the expression of the node with the highest score is:

[0040] ;

[0041] ;

[0042] ;

[0043] ;

[0044] ;

[0045] ;

[0046] ;

[0047] ;

[0048] Where, Denote the node with the highest score, Denote the maximum value selection function, Denote the action score, Denote the feed-forward neural network, Represents the graph perception result, Represents the graph perception self-attention mechanism, Represents the topological graph parameters, Represents the learnable transformation matrix for queries, Represents the learnable transformation matrix for keys, Represents the constant feature dimension, Represents the additional graph structure information, Represents the learnable transformation matrix for values, Represents the pairwise distance matrix between nodes obtained from the edge information of the graph, Represents the learnable weight matrix, Represents the learnable bias vector, Represents the cross-modal attention result, Represents the dynamic topological graph, Represents the cross-attention mechanism, Represents the node embedding, Represents the encoded correction instruction, Represents the navigable points after redundancy removal, Represents the traversed edges, Represents average pooling, Represents the action encoding sequence, Represents the direction embedding, Represents the step embedding.

[0049] A variable environment visual language navigation system based on historical reflection, comprising:

[0050] A visual language understanding module for processing variable environment visual observation data using a large model visual encoder and a query transformer to obtain a scene graph encoding embedding;

[0051] A visual language reflection module for using a gated network to determine whether to call a reflection network as part of a modular network. When a call is required, the reflection network analyzes the historical navigation information and the scene graph encoding embedding to obtain a correction instruction;

[0052] An action prediction module for calculating based on the correction instruction using a graph perception self-attention mechanism to perform path navigation on an agent to obtain a visual language navigation result and complete the visual language navigation of a variable environment.

[0053] The beneficial effects of the present invention are as follows: By introducing a historical reflection mechanism, the intelligent agent can utilize the historical information accumulated during the navigation process to deeply understand the environmental state and user intentions, and dynamically adjust the parsing of natural language instructions to bridge the gap between the instructions and the actual scenario, thereby improving the robustness and adaptability of the navigation system in real scenarios. (1) Dynamically detect and correct deviations in environmental changes and human instructions through the gating network and reflection network in the visual language reflection module, such as problems like increased obstacles, removed landmarks, modified textures or colors, and descriptor replacements. Combining with the historical reflection mechanism, the intelligent agent can dynamically adjust path planning and instruction parsing during navigation to bridge the gap between instructions and the environment, thus significantly enhancing the accuracy and adaptability of the navigation system in complex scenarios. (2) Use pre-trained large models (such as BLIP, BLIP2, InstructBLIP, LLaVA, etc.) and query transformers for deep alignment of vision and language, enabling the intelligent agent to generate high-quality navigation instructions in a dynamically changing environment. The historical reflection module, through in-depth analysis of historical information and combined with the dynamic update of the graph network, ensures that the performance of instruction parsing does not degrade due to environmental changes or increased path length, thereby enhancing the stability and robustness of instruction parsing. (3) Dynamically generate new navigation strategies by generating correction suggestions and combining multi-modal features, enabling the intelligent agent to flexibly handle diverse navigation environments. Through training on large-scale data collected in different environments and adopting efficient fine-tuning methods such as LoRA, the navigation model of the present invention demonstrates strong generalization ability in different scenarios. At the same time, combined with the graph network dynamic maintenance mechanism of the action prediction module, the intelligent agent can plan paths more efficiently and avoid redundant actions and repeated explorations, significantly improving the navigation efficiency. (4) Use the graph-aware self-attention mechanism combined with cross-modal attention results and dynamic topology graphs to analyze the scores of each node, making the selection of the optimal path more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] This specification will further illustrate by way of exemplary embodiments, which will be described in detail through the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same numbers represent the same structures, where:

[0055] Figure 1 is a schematic diagram of an application scenario of a variable environment visual language navigation system based on historical reflection according to some embodiments of this specification;

[0056] Figure 2 is an exemplary flowchart of a variable environment visual language navigation method based on historical reflection according to some embodiments of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0058] Embodiment 1

[0059] Figure 1 It is a schematic diagram of the modules of a variable environment visual language navigation system based on historical reflection shown in some embodiments of this specification.

[0060] In some embodiments, the variable environment visual language navigation system based on historical reflection may include a visual language understanding module, a visual language reflection module, and an action prediction module.

[0061] The visual language understanding module is used to process the variable environment visual observation data by using a large model visual encoder and a query transformer to obtain a scene graph encoding embedding.

[0062] The visual language understanding module mainly consists of a query transformer and a large model visual encoder. The query transformer encodes each perspective into a fixed-length visual token. For the candidate images obtained by the agent, the visual encoder is used to extract features. These visual features are then processed through a cross-attention mechanism with learnable query embeddings and self-attention learning with the text embedding of the instruction to obtain an instruction-based image query. After these queries are linearly projected, they are input into the large model for processing as image tokens. To handle imperfect instructions, the language input in this method not only includes the instruction but also contains new path planning information provided by the reflection module.

[0063] The visual language reflection module is used to determine whether to call the reflection network as part of the modular network by using a gating network. When it is necessary to call, the reflection network analyzes the historical navigation information and the scene graph encoding embedding to obtain a corrected instruction.

[0064] The visual language reflection module consists of a gating network and a historical reflection network.

[0065] The gating network is specifically designed to determine whether the current historical information contains errors, such as: obstacle addition, landmark removal, texture and color modification, descriptor replacement, etc. If any of these errors are detected, the agent will dynamically activate the historical reflection module to generate a navigation strategy for these error conditions. To ensure efficient navigation, instead of directly using a vision-language large model to identify potential vision-language mismatches, the gating network performs binary classification on visual and language information based on the Transformer architecture. Episodes marked as having a mismatch by the gating network will be passed to the vision-language reflection module. In this module, the model identifies the error type through a structured process, generates potential solutions, and formulates instructions.

[0066] The reflection network is used to generate a navigation strategy after the gating network confirms an error. In this case, the reflection network analyzes the current error based on the historical information, proposes a correction suggestion and corresponding verification criteria. This suggestion is then used to generate a new navigation instruction to replace the original instruction in the multimodal alignment stage. At the same time, in the decision-making stage, the new instruction is added to the prompt for the decision-making model to perform multi-step reasoning. After the decision-making model marks the end, if the suggestion newly proposed by the network is verified and it is determined that the comment associated with this suggestion is inconsistent with the actual navigation history, a new suggestion-comment pair will be generated and the above process will be repeated. If the two are consistent, the navigation process ends.

[0067] The action prediction module is used to perform calculations based on the corrected instruction, utilize the graph-aware self-attention mechanism to navigate the agent, obtain the vision-language navigation result, and complete the vision-language navigation in a variable environment.

[0068] A graph-based representation method is designed in the action prediction module to utilize the understanding and reasoning capabilities of the large language model, ensure the effectiveness of the semantic mapping from vision-language input to action, and enhance the model's understanding and representation capabilities of natural language. The topological graph is dynamically maintained and serves as a memory mechanism to track the navigation path. The next step is selected from the entire constructed topological graph. When there are imperfections in the confirmed instructions in the historical path, the visual markers obtained through vision-language understanding, as well as the new instructions and corresponding strategies derived from the vision-language reflection module, are input into the system prompt, enabling effective planning and backward tracking of unvisited nodes.

[0069] In some embodiments, the action prediction module can use supervised learning or reinforcement learning, and can be completed step by step: first train the action prediction to obtain the path planning ability, and then introduce the gating and reflection networks to achieve error correction and replanning. It is also possible to perform joint backpropagation of multi-module losses to achieve end-to-end optimization.

[0070] The graph network consists of visited nodes and un-explored nodes adjacent in the path. All candidate visual observations of each visited node are represented by that node after average pooling, while each un-explored node is represented by partial pooling of the corresponding visual observations of all its adjacent visited nodes. Each view is represented by the sum of the following three parts: 1. Visual features, representing the visual information of the node; 2. Direction embedding, representing the position of the node; 3. Step embedding, representing the traversal order of the agent in the current task; the step embedding of un-explored nodes is set to 0, and a "stay node" representing the stop action is added to the graph network.

[0071] At each time step, the navigation graph consists of a set of nodes and a set of edges. The node embeddings are passed into a multi-layer cross-modal Transformer for modeling the relationship between the instruction and the nodes. Specifically, the node embeddings are first cross-modally aligned with the instruction and then further processed by the graph-aware self-attention module. The graph-aware self-attention module combines the distance and visual similarity between nodes to enhance context understanding.

[0072] In some embodiments, a variable environment visual language navigation system based on historical reflection can be used to perform a variable environment visual language navigation method based on historical reflection, including: S1: Using a large model visual encoder and a query transformer to process variable environment visual observation data to obtain a scene graph encoding embedding; S2: Using a gating network to analyze historical navigation information, the encoded embedding instruction, and the scene graph encoding embedding to obtain a corrected instruction; S3: Based on the corrected instruction, performing path navigation on the agent to obtain a visual language navigation result, and completing the visual language navigation of the variable environment.

[0073] In some embodiments of this specification, the processor executes a variable environment visual language navigation method based on historical reflection using a variable environment visual language navigation system based on historical reflection. In this way, the difference between instructions and the actual scene can be bridged, and the robustness and adaptability of the navigation system in the real scene can be improved. (1) Dynamically detect and correct deviations in the environment changes and human instructions through the gating network and reflection network in the visual language reflection module, such as problems like increased obstacles, removed landmarks, modified textures or colors, and descriptor replacements. Combining the historical reflection mechanism, the agent can dynamically adjust path planning and instruction parsing during navigation, bridging the gap between instructions and the environment, thereby greatly enhancing the accuracy and adaptability of the navigation system in complex scenarios. (2) Use pre-trained large models (such as BLIP, BLIP2, InstructBLIP, LLaVA, etc.) and query transformers for deep alignment of vision and language, enabling the agent to generate high-quality navigation instructions in a dynamically changing environment. The historical reflection module, through in-depth analysis of historical information and combined with the dynamic update of the graph network, ensures that the performance of instruction parsing does not degrade due to environmental changes or increased path length, thereby enhancing the stability and robustness of instruction parsing. (3) Dynamically generate new navigation strategies by generating correction suggestions and combining multi-modal features, enabling the agent to flexibly handle diverse navigation environments. Through training on large-scale data collected in different environments and using efficient fine-tuning methods such as LoRA, the navigation model of the present invention exhibits strong generalization ability in different scenarios. At the same time, combined with the graph network dynamic maintenance mechanism of the action prediction module, the agent can plan paths more efficiently and avoid redundant actions and repeated explorations, greatly improving the navigation efficiency.

[0074] Embodiment 2

[0075] Figure 2 is an exemplary flowchart of a variable environment visual language navigation method based on historical reflection shown in some embodiments of this specification. As Figure 2 shown, the process includes the following steps. In some embodiments, the process can be executed by a processor.

[0076] S1: Use the large model visual encoder and query transformer to process the variable environment visual observation data to obtain the scene graph encoding embedding, the navigation history information used in the next action, and the corresponding encoded embedding instruction.

[0077] The large model visual encoder is an encoder for extracting image features. For example, the large model visual encoder can include BLIP, BLIP2, InstructBLIP, LLaVA, etc.

[0078] The scene graph encoding embedding is an embedding vector for aligning visual features and language features.

[0079] In some embodiments, the processor may implement S1 based on the following steps: using a frozen large model visual encoder to process variable environmental visual observation data to obtain image features; using a query transformer to perform self-attention mechanism fusion on learnable query embeddings and instruction embeddings to obtain instruction information; based on the cross-attention mechanism, fusing the image features and the instruction information to obtain a scene graph encoding embedding, navigation history information used in the next action, and corresponding encoded embedding instructions.

[0080] The frozen large model visual encoder is a large model visual encoder with fixed weights.

[0081] In some embodiments, the processor may use lightweight methods such as fully freezing, partially freezing, or LoRA to adapt the large model visual encoder to the navigation and reflection tasks to obtain a frozen large model visual encoder.

[0082] The variable environmental visual observation data is image data visually acquired by the agent.

[0083] The image features are features reflecting the relationship between the variable environmental visual observation data and the environment.

[0084] In some embodiments, the processor may use the core weights of the frozen large model visual encoder to train and fine-tune the query transformer to obtain a trained query transformer.

[0085] The query transformer may be a Q-Former (Querying Transformer) model.

[0086] In some embodiments, the processor may input the variable environmental visual observation data into the frozen large model visual encoder for processing to obtain image features.

[0087] The learnable query embeddings and instruction embeddings are descriptions of path points and information on action sequences.

[0088] The instruction features are features comprehensively describing the movement of the agent.

[0089] In some embodiments, the processor may use the query transformer to perform self-attention mechanism fusion on 32 fixed-length learnable query embeddings and instruction embeddings to obtain instruction features.

[0090] In some embodiments, the expression of the scene graph encoding embedding may be:

[0091] ;

[0092] Where represents the scene graph encoding embedding, represents the cross-attention mechanism, represents the image features, represents the instruction information.

[0093] S2: Use the gating network and the reflection network to analyze the encoded embedding of the scene graph, the navigation history information used in the next action, and the corresponding encoded embedding instruction to obtain a corrected instruction.

[0094] The gating network is a lightweight neural network used to correct errors in the encoded embedding instruction. For example, the gating network can include a lightweight Transformer.

[0095] The historical navigation information is the visual observations and instructions, the sequence of visited nodes, etc. in the past several steps.

[0096] The encoded embedding instruction is the instruction in the historical instructions that does not match the visual observation and the navigation graph.

[0097] The corrected instruction is the instruction parameter that matches the visual observation and the navigation graph.

[0098] In some embodiments, the processor can implement S2 based on the following steps: Concatenate the historical navigation information, the encoded embedding instruction, and the encoded embedding of the scene graph into sequence data, input it into the gating network to obtain the output result of the gating network; Input the output result of the gating network into a binary classification fully connected layer, and through calculation, obtain an error label value; When the error label value is greater than the threshold, use the reflection network to identify the specific error type, and autoregressively generate a corrected instruction through a large language model.

[0099] The sequence data is the concatenated data integrating the historical navigation information, the encoded embedding instruction, and the encoded embedding of the scene graph.

[0100] The error label value is a numerical value reflecting whether there is a mismatch problem in the output result of the gating network.

[0101] In some embodiments, the expression of the error label value can be:

[0102] ;

[0103] ;

[0104] where, represents the error label value, represents the standard deviation of represents the binary classification fully connected layer, represents the output result of the gating network, represents the gating network, represents the encoded embedding of the scene graph, Represents an encoding embedding instruction, Represents historical navigation information.

[0105] In some embodiments, when the error tag value is greater than the threshold, it is determined that there is a mismatch problem, and the reflection network needs to be called to identify the specific error type and generate a correction instruction. Otherwise, the original instruction is continued to be executed.

[0106] The reflection network can be a deep neural network.

[0107] In some embodiments, the gating network can be trained with binary or multi-class cross-entropy loss to determine whether there is a mismatch.

[0108] In some embodiments, the reflection network can generate correction instructions based on a large language model and can be trained through autoregressive language modeling or reinforcement learning (such as RLHF).

[0109] In some embodiments, the expression of the error type can be:

[0110] ;

[0111] ;

[0112] Wherein, Represents the error type, Represents the maximum value selection function, Represents Function, Represents the first hyperparameter for error type recognition, Represents the function value calculated according to the large language model, Represents the second hyperparameter for error type recognition, Represents the large language model, Represents the scene graph encoding embedding, Represents the encoding embedding instruction, Represents the historical navigation information.

[0113] The first hyperparameter for error type recognition and the second hyperparameter for error type recognition are learnable function parameters.

[0114] In some embodiments, the expression of the correction instruction can be:

[0115] ;

[0116] Wherein, Represents the correction instruction, Represents the large language model, Represents the function value calculated according to the large language model, Represents the error type.

[0117] In some embodiments, the pre - provided prompt form can be: "Your task is to cut and extract declarative text into operation descriptors and landmarks. Note that some phrases are combined with words indicating directions to better display spatial location information, and each phrase must contain at least one action. Do not re - phrase. The representation of this text is '%s', please only return your answer, separated by '|'. Each sentence can be disassembled more than once." In the reflection part, after the large - language model reflection function receives the call signal from the gating network, it combines the instruction information and image features , and makes targeted modifications. For example: "Given that the current path is blocked, we propose the following detour plan: Backtracking: Backtrack from any possible side passage in the corridor or room. Avoiding the sofa: Move slightly left or right to avoid the sofa blocking the way to the carpet. Revised steps: Backtracking from the corridor: Re - enter the bedroom, keeping to the left route along the initial path. Bypassing the sofa: Lateral movement: Turn right at the corner of the living room and move laterally to bypass the sofa. Alternative route: If the lateral movement is not passable, then move diagonally across the dining area, close to the left side of the terrace. Criticism: The agent should anticipate backtracking from the corridor first, then reaching the living room by laterally bypassing the sofa or diagonally crossing the dining area, and finally stopping near the front carpet."

[0118] S3: Based on the corrected instruction, use the graph - aware self - attention mechanism for calculation, perform path navigation on the agent, obtain the visual - language navigation result, and complete the visual - language navigation in a variable environment.

[0119] The visual - language navigation result is the result reflecting the final navigation path of the agent.

[0120] In some embodiments, the processor can implement S3 based on the following steps: Based on the action encoding sequence, unify the path points and visual features to obtain node embeddings; perform cross - modal cross - attention calculation on the corrected instruction and the node embeddings to obtain cross - modal attention results; use the visual observation data of the variable environment to judge the connectivity between nodes and construct a dynamic topological graph; combine the cross - modal attention results and the dynamic topological graph, perform the graph - aware self - attention mechanism to obtain graph - aware results; use a feed - forward neural network to analyze the graph - aware results, comprehensively consider the distance and difficulty of reaching the node to obtain an action score; based on the action score, select the node with the highest score and navigate along the optimal path to this node to obtain the visual - language navigation result and complete the visual - language navigation in a variable environment.

[0121] The action encoding sequence is the sequence reflecting the actions of the agent throughout the navigation process.

[0122] In some embodiments, the processor may obtain an action encoding sequence by removing redundancy and performing encoding processing based on the action instruction.

[0123] The node embedding is the embedding representation of an unvisited node.

[0124] In some embodiments, the processor may construct a node embedding for the unvisited node by combining average pooling, direction embedding, and step embedding.

[0125] The cross-modal attention result is the attention result of the corrected instruction and the node embedding after comprehensive encoding.

[0126] In some embodiments, the processor may perform encoding processing on the corrected instruction to obtain the encoded corrected instruction.

[0127] The dynamic topology graph is a topology graph used to judge the connectivity between nodes.

[0128] In some embodiments, the processor may utilize variable environmental visual observation data, analyze the relationships between various nodes, and construct a dynamic topology graph.

[0129] The graph perception result is the result reflecting the distance and visual similarity between nodes.

[0130] The action score is the data obtained by scoring each node in the graph perception result by comprehensively considering the distance and difficulty for the agent to reach the node; the higher the action score, the higher the efficiency for the agent to reach the node.

[0131] The highest-score node is the node with the highest current score.

[0132] In some embodiments, the processor may select the highest-score node, navigate to the node along the shortest or optimal path, select and execute a navigation action on the environmental nodes in combination with the graph network structure, and at the same time compare and verify the historical path with the correction suggestion. If there is still a mismatch, the reflection process is repeated until the agent successfully reaches the target node or it is confirmed that there is no further feasible path, and the visual language navigation result is obtained.

[0133] In some embodiments, the processor uses a two-layer feed-forward network to process the output node representation to generate an action score. The agent will select the node with the highest score as the target and navigate to the selected node along the shortest path in the graph network. At the same time, the method masks the scores of the visited nodes to encourage the agent to explore other nodes.

[0134] In some embodiments, the expression of the highest-score node may be:

[0135] ;

[0136] ;

[0137] ;

[0138] ;

[0139] ;

[0140] ;

[0141] ;

[0142] ;

[0143] wherein, represents the highest score node, represents the maximum value selection function, represents the action score, represents the feed-forward neural network, represents the graph perception result, represents the graph perception self-attention mechanism, represents the topological graph parameter, represents the learnable transformation matrix of the query, represents the learnable transformation matrix of the key, represents the constant feature dimension, represents the additional graph structure information, represents the learnable transformation matrix of the value, represents the pairwise distance matrix between nodes obtained from the edge information of the graph, represents the learnable weight matrix, represents the learnable bias vector, represents the cross-modal attention result, represents the dynamic topological graph, represents the cross-attention mechanism, represents the node embedding, represents the encoded correction instruction, represents the navigable point after redundancy removal, represents the traversed edge, represents the average pooling, represents the action encoding sequence, represents the direction embedding, represents the step embedding.

[0144] In some embodiments of this specification, by introducing a historical reflection mechanism, the agent can utilize the historical information accumulated during the navigation process to deeply understand the environmental state and user intent, and dynamically adjust the parsing of natural language instructions to bridge the gap between the instructions and the actual scenario, improving the robustness and adaptability of the navigation system in real scenarios. (1) Dynamically detect and correct deviations in environmental changes and human instructions through the gating network and reflection network in the visual language reflection module, such as issues like increased obstacles, removed landmarks, modified textures or colors, and descriptor replacements. Combining the historical reflection mechanism, the agent can dynamically adjust path planning and instruction parsing during navigation to bridge the gap between instructions and the environment, thus significantly enhancing the accuracy and adaptability of the navigation system in complex scenarios. (2) Use pre-trained large models (such as BLIP, BLIP2, InstructBLIP, LLaVA, etc.) and query transformers for deep alignment of vision and language, enabling the agent to generate high-quality navigation instructions in a dynamically changing environment. The historical reflection module, through in-depth analysis of historical information and combined with the dynamic update of the graph network, ensures that the performance of instruction parsing does not degrade due to environmental changes or increased path length, thereby enhancing the stability and robustness of instruction parsing. (3) Generate correction suggestions and, combined with multi-modal features, dynamically generate new navigation strategies, enabling the agent to flexibly respond to diverse navigation environments. By training on large-scale data collected in different environments and using efficient fine-tuning methods such as LoRA, the navigation model of the present invention demonstrates strong generalization ability in different scenarios. At the same time, combined with the graph network dynamic maintenance mechanism of the action prediction module, the agent can plan paths more efficiently, avoid redundant actions and repeated explorations, and significantly improve navigation efficiency. (4) Utilize the graph-aware self-attention mechanism combined with the cross-modal attention results and the dynamic topology graph to analyze the scores of each node, making the selection of the optimal path more accurate.

Claims

1. A variable environment visual language navigation method based on historical reflection, characterized in that, Including: S1: Using a large model visual encoder and a query transformer to process variable environment visual observation data, respectively obtaining a scene graph encoding embedding, navigation history information used in the next action, and corresponding encoded embedding instructions; S2: Using a gated network and a reflection network to perform historical reflection on the scene graph encoding embedding, the navigation history information used in the next action, and the corresponding encoded embedding instructions, obtaining a correction instruction; including: concatenating the scene graph encoding embedding, the navigation history information used in the next action, and the corresponding encoded embedding instructions into sequence data, inputting it into the gated network, and obtaining the output result of the gated network; Inputting the output result of the gated network into a binary classification fully connected layer, and obtaining an error label value through calculation; When the error label value is greater than the threshold, using the reflection network to identify the specific error type, and autoregressively generating a correction instruction through a large language model; wherein, the reflection network is a large language model for generating correction instructions; S3: Based on the correction instruction, using a graph-aware self-attention mechanism to calculate, performing path navigation on the intelligent agent, obtaining a visual language navigation result, and completing the visual language navigation in a variable environment; including: unifying path points and visual features based on an action encoding sequence, obtaining node embeddings; Performing cross-modal cross-attention calculation on the correction instruction and the node embeddings, obtaining a cross-modal attention result; Using variable environment visual observation data to judge the connectivity between nodes, and constructing a dynamic topology graph; Combining the cross-modal attention result and the dynamic topology graph, and using a graph-aware self-attention mechanism to calculate, obtaining a graph-aware result; Analyzing the graph-aware result using a feed-forward neural network, comprehensively considering the distance and difficulty of reaching the node, and obtaining an action score; wherein, the action score is used to reflect the efficiency of the intelligent agent reaching this node; Based on the action score, selecting the node with the highest score, and navigating along the optimal path to this node, obtaining a visual language navigation result, and completing the visual language navigation in a variable environment.

2. The variable environment visual language navigation method based on historical reflection according to claim 1, wherein The S1 includes: Using a frozen large model visual encoder to process variable environment visual observation data, obtaining image features; wherein, the frozen large model visual encoder is a large model visual encoder with its weights fixed in a lightweight manner; Using a query transformer to perform self-attention mechanism fusion on learnable query embeddings and instruction embeddings, obtaining instruction information; Based on the cross-attention mechanism, fusing the image features and the instruction information, obtaining a scene graph encoding embedding, navigation history information used in the next action, and corresponding encoded embedding instructions.

3. The variable environment visual language navigation method based on historical reflection according to claim 2, wherein The expression of the scene graph encoding embedding is: ; Among them, represents the scene graph encoding embedding, represents the cross-attention mechanism, represents the image feature, represents the instruction information.

4. The variable environment visual language navigation method based on historical reflection according to claim 1, characterized in that The expression of the error label value is: ; ; Among them, represents the error tag value, represents the standard deviation of, represents the binary classification fully connected layer, represents the output result of the gating network, represents the gating network, represents the scene graph encoding embedding, represents the encoding embedding instruction, represents the historical navigation information.

5. The variable environment visual language navigation method based on historical reflection according to claim 1, wherein The expression of the error type is: ; ; Among them, represents the error type, represents the maximum value selection function, represents function, represents the first hyperparameter for error type recognition, represents the function value calculated according to the large language model, represents the second hyperparameter for error type recognition, represents the large language model, represents the scene graph encoding embedding, represents the encoding embedding instruction, represents the historical navigation information; The expression of the correction instruction is: ; Among them, represents a correction instruction, represents a large language model, represents the function value calculated according to the large language model, represents the error type.

6. The method for variable environment visual language navigation based on historical reflection according to claim 1, characterized in that, The expression of the node with the highest score is: ; ; ; ; ; ; ; ; Among them, represents the highest score node, represents the maximum value selection function, represents the action score, represents the feed-forward neural network, represents the graph perception result, represents the graph perception self-attention mechanism, represents the topological graph parameters, represents the learnable transformation matrix for queries, represents the learnable transformation matrix for keys, represents the constant feature dimension, represents the additional graph structure information, represents the learnable transformation matrix for values, represents the pairwise distance matrix between nodes obtained from the edge information of the graph, represents the learnable weight matrix, represents the learnable bias vector, represents the cross-modal attention result, represents the dynamic topological graph, represents the cross-attention mechanism, represents the node embedding, represents the encoded correction instruction, represents the navigable points after redundancy removal, represents the traversed edge, represents the average pooling, represents the action encoding sequence, represents the direction embedding, represents the step embedding.

7. A variable environment visual language navigation system based on historical reflection, for performing the variable environment visual language navigation method based on historical reflection according to any one of claims 1 to 6, characterized in that, Including: A visual language understanding module, used to process variable environment visual observation data using a large model visual encoder and a query transformer, obtaining a scene graph encoding embedding; The visual language reflection module is used to determine whether to call the reflection network as part of the modular network by using a gating network. When it is necessary to call, the reflection network analyzes the historical navigation information and the encoded embedding of the scene graph to obtain a correction instruction; The action prediction module is used to perform calculations based on the correction instruction by using a graph-aware self-attention mechanism to navigate the path of the agent, obtain a visual language navigation result, and complete the visual language navigation in a variable environment.

Citation Information

Patent Citations

  • Indoor visual navigation method based on causal attention

    CN115512214A

  • Visual language navigation method based on historical context information enhancement

    CN118010026A