Visual language navigation method, system and medium based on motion feature alignment

By encoding the agent's action features and integrating them with the visual features, the problem of alignment between visual and language modality is solved, and the navigation success rate is improved.

CN119807671BActive Publication Date: 2025-05-09SHANDONG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510293010.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-09
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

In the current visual language navigation task, the alignment problem between visual and language modalities is difficult to effectively solve, resulting in a low navigation success rate.

Method used

By encoding the action characteristics of the agent during movement, combining object characteristics, room characteristics and action characteristics, the characteristics are fused, and the attention mechanism is used to calculate the weight of the command words to ensure the alignment of the actions of the visual and language modalities.

Benefits of technology

The success rate of agent navigation is significantly improved, and the matching degree between visual modes and command modes is enhanced through action feature alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807671B_ABST
    Figure CN119807671B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of visual navigation technology, and specifically relates to a visual language navigation method, system and medium based on action feature alignment, including the following steps: encoding action features based on the coordinate relationship between the current position of the agent and the candidate position; fusing the action features with the room features and object features seen and updating the topological map; weighting the words in the instruction based on the fused features to obtain dynamic instruction features. The method disclosed by the present invention fully considers the action alignment of the visual modality and the instruction modality, greatly improving the success rate of navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of visual navigation technology, and specifically relates to a visual language navigation method, system and medium based on action feature alignment. Background Art

[0002] At present, the research methods of visual language navigation tasks are mainly divided into end-to-end learning methods and modular methods. The end-to-end method directly uses a deep learning model to perform integrated modeling from raw input (such as pictures, text) to navigation actions. These methods usually use neural network architectures such as Transformer, LSTM, CNN, and enhance the matching degree between instructions and visual information through modal alignment technology. However, this method is difficult to explain the decision-making process of the model and requires a large amount of data for training.

[0003] Modular methods are different from end-to-end methods. Modular methods split VLN tasks into multiple subtasks, such as text parsing, environmental perception, path planning, etc. Common modular methods include reinforcement learning (RL)-based methods and planning-based methods. Although modular methods can enhance the interpretability of the model, the independence between modules may lead to discontinuity in information transmission. A core difficulty of VLN tasks is the alignment of visual and language modalities. Since text instructions are usually abstract, flexible, and vague, while visual information is high-dimensional, dynamic, and rich in details, how to establish an accurate mapping relationship between the two is the focus of current research. Summary of the invention

[0004] Based on the above problems, this application enriches the source of visual information by encoding the action characteristics of the intelligent body during movement, ensures that the action-related words in the instructions can align with the visual information, and greatly improves the success rate of intelligent body navigation.

[0005] To achieve the above object, the technical solution of the present invention is as follows:

[0006] A visual language navigation method based on action feature alignment, including the following contents:

[0007] S1. Constructed topological map : Navigate to the environment As an undirected graph, the nodes represent the locations where the agent can move, and the edges represent the paths between two nodes. ;

[0008] S2. Action-aware coding: Based on the positional relationship between the agent's current node and any candidate node, encode the changes in the deflection angle and pitch angle to the candidate node;

[0009] S3. Feature fusion: combining object features , Room Features And the encoded action features, obtain visual features with action information , and the topological map Make updates;

[0010] S4. Dynamic instruction weighting: Calculate the weight of each word in the instruction based on visual features and attention mechanism, and calculate the weighted instruction features .

[0011] Preferably, the agent is initially at a random position in the navigation environment and needs to navigate to the required position according to instructions and visual discovery. At time t, the object features seen by the agent are The room features are , Navigation instructions , the constructed topological map is ; It contains three types of nodes, namely the nodes that have been visited, the nodes where the agent is currently located, and the nodes that the agent can reach next.

[0012] Preferably, based on the positional relationship between the agent's current node and any candidate node, the change in the deflection angle to the candidate node is encoded, assuming , Is the path Two consecutive points on is a path of the agent from the current point to the i-th candidate point, , The coordinates are , , the current heading angle of the agent is , then the relative deflection angle The calculation method is as follows:

[0013] ;

[0014] ;

[0015] function To convert radians to degrees, represents the deflection angle between two nodes, is 360, which is used to ensure The range is between 0 and 360 degrees;

[0016] get after, Updated to , continue to calculate the path according to the above formula The deflection angles between the remaining adjacent points on the graph are finally obtained. , indicating that the agent is on the path The relative deflection angle between all adjacent points on the The number of median edges;

[0017] Using an embedding layer Convert to feature encoding:

[0018] ;

[0019] For path The relative deflection angle encoding.

[0020] Preferably, based on the positional relationship between the agent's current node and any candidate node, the change in the pitch angle of the agent moving to the candidate node is encoded:

[0021] Assumptions , The coordinates are , , the current elevation angle of the agent is , then the relative pitch angle The calculation method is as follows:

[0022] ;

[0023] ;

[0024] is the depression angle between adjacent nodes, Indicates the maximum value of the relative pitch angle;

[0025] get back, Updated to , continue to calculate the path according to the above formula The pitch angle changes between the remaining adjacent points on the graph are finally obtained. , indicating that the agent is on the path The relative pitch angle between all adjacent points on the graph when moving;

[0026] Using an embedding layer Convert to feature encoding:

[0027] ;

[0028] For path The relative deflection angle encoding.

[0029] Preferred, comprehensive and You can get the path Overall action characteristics , the formula is as follows:

[0030] ;

[0031] ;

[0032] represents the action encoding of all paths, n is the number of paths, Represents an encoder consisting of a two-layer transformer structure.

[0033] Preferably, in step S3, visual feature fusion: combining object features , Room Features And the encoded action features , obtain visual features with action information :

[0034] ;

[0035] Represents a splicing operation;

[0036] Topology feature update: For the node where the agent is currently located, use and As its feature, for the candidate node of the agent's next step, use , as well as As its characteristic, as well as They represent the object features and room features in the visible view, respectively.

[0037] Preferably, the word weight calculation in step S4 is based on visual features: Calculate the weight of each word in the instruction with the attention mechanism;

[0038] First, get the attention weight of each word at time t :

[0039] ;

[0040] d is The dimension is then max pooling is used to obtain the final weight , T is the transpose:

[0041] ;

[0042] Instruction feature weighting: Calculate the weighted instruction feature based on the weight of each word :

[0043] ;

[0044] Where L is the number of words in the instruction.

[0045] Preferably, the visual features found at time t , instruction characteristics And topological maps Put in a Dual-scale Graph Transformer network to predict the next action of the agent. The formula is as follows:

[0046] ;

[0047] in, represents the next action to be performed by the agent, Represents Dual-scale GraphTransformer, which can understand visual features simultaneously , weighted instruction features and topological maps .

[0048] A visual language navigation system based on action feature alignment is used to implement the visual language navigation method based on action feature alignment described in this application.

[0049] A storage medium, when running on a computer, executes the steps of the visual language navigation method based on action feature alignment described in the present application.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] The action features are encoded based on the coordinate relationship between the current position of the agent and the candidate position; the action features are fused with the room features and object features seen and the topological map is updated; the words in the command are weighted based on the fused features to obtain dynamic command features. The method disclosed in the present invention fully considers the action alignment of the visual modality and the command modality, greatly improving the success rate of navigation. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 This is the flow chart of this application. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0055] The present invention provides a visual language navigation method based on action feature alignment, such as Figure 1 As shown, this method improves the accuracy of the agent's navigation by encoding the agent's action features, enriching the visual information, and ensuring the alignment of the visual modality and the language modality.

[0056] A visual language navigation method based on action feature alignment includes the following steps:

[0057] The agent is initially at a random position in the navigation environment and needs to navigate to the required position based on the instructions and visual discovery. It is an undirected graph where nodes represent locations where the agent can move and edges represent paths between two nodes.

[0058] At time t, the features of the object seen by the agent are Room features: (The data set is an indoor environment), navigation instructions The constructed topological map is . It contains three types of nodes, namely the nodes that have been visited, the nodes where the agent is currently located, and the nodes that the agent can reach next.

[0059] (1) Action-aware coding:

[0060] Assume that the agent is currently at node At , one of the candidate nodes is ,from arrive Path for , the path is The path to the i-th candidate node. Here we need to calculate the agent on the path The yaw and pitch angle changes caused by the movement between all adjacent nodes.

[0061] (1.1) Deflection angle encoding. Based on the positional relationship between the agent’s current node and any candidate node, encode the change in the deflection angle from the agent to the candidate node. Assume , The coordinates are , , the current heading angle of the agent is , then the relative deflection angle The calculation method is as follows:

[0062] ;

[0063] ;

[0064] Here the function To convert radians to degrees, represents the deflection angle between two nodes, is 360, which is used to ensure The range is between 0 and 360 degrees.

[0065] get after, Updated to , continue to calculate the path according to the above formula The deflection angles between the remaining adjacent points on the graph are finally obtained. , indicating that the agent is on the path The relative deflection angle between all adjacent points on the Next, we use an embedding layer to Convert to feature encoding:

[0066] ;

[0067] For path The relative deflection angle encoding.

[0068] (1.2) Pitch angle encoding. Based on the positional relationship between the agent’s current node and any candidate node, encode the change in pitch angle when it moves to the candidate node. Assume , The coordinates are , , the current elevation angle of the agent is , then the relative pitch angle The calculation method is as follows:

[0069] ;

[0070] ;

[0071] is the depression angle between adjacent nodes, Indicates the maximum relative pitch angle, which is 30 degrees.

[0072] get back, Updated to , continue to calculate the path according to the above formula The pitch angle changes between the remaining adjacent points on the graph are finally obtained. , indicating that the agent is on the path The relative pitch angles of all adjacent points on the graph are then converted into Convert to feature encoding:

[0073] ;

[0074] For path The relative deflection angle encoding.

[0075] comprehensive and You can get the path Overall action characteristics :

[0076] ;

[0077] ;

[0078] here represents the action encoding of all paths, n is the number of paths, Represents an encoder consisting of a two-layer transformer structure.

[0079] (2) Feature fusion:

[0080] (2.1) Visual feature fusion: Combining , And the encoded action features , obtain visual features with action information :

[0081] ;

[0082] here Represents a concatenation operation.

[0083] (2.1) Topology map feature update: Update The features of the agent's current position and all candidate positions are represented in the new topological map. For the node where the agent is currently located, use and As its feature, for the candidate node of the agent's next step (assuming it is the i-th one), use , as well as As its characteristic, here as well as They represent the object features and room features in the visible view, respectively.

[0084] (3) Dynamic instruction weighting:

[0085] (3.1) Word weight calculation: based on visual features The attention mechanism is used to calculate the weight of each word in the instruction. First, the attention weight of each word at time t is obtained. :

[0086] ;

[0087] Here d is The dimension is then max pooling is used to obtain the final weight :

[0088] .

[0089] (3.2) Instruction feature weighting: Calculate the weighted instruction feature based on the weight of each word :

[0090] ;

[0091] Where L is the number of words in the instruction.

[0092] Finally, the visual features found at time t are , instruction characteristics And topological maps Insert a Dual-scale Graph Transformer network to predict the next action of the agent.

[0093] The formula is as follows:

[0094] ;

[0095] in, represents the next action to be performed by the agent, Represents Dual-scale GraphTransformer, which can understand visual features simultaneously , weighted instruction features and topological maps .

[0096] The present invention fully considers the action information of the intelligent agent during movement and encodes it, enriches the visual information, and ensures the alignment of the two modalities of vision and text.

[0097] As shown in Table 1 and Table 2, the present invention has achieved better results than other methods on multiple public data sets.

[0098] Table 1 Experimental results of various methods on the R4R Validation Unseen dataset

[0099] .

[0100] Table 2 Experimental results of various methods on the RxR-English Validation Unseen dataset

[0101] .

[0102] The method embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative work.

[0103] The system is constructed to run the method of the present application. Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0104] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A visual language navigation method based on action feature alignment, characterized in that: Includes the following: S1. Constructed topological map : Navigate to the environment As an undirected graph, the nodes represent the locations where the agent can move, and the edges represent the paths between two nodes. ; The agent is initially at a random position in the navigation environment and needs to navigate to the required position based on instructions and visual discovery. At time t, the features of the object seen by the agent are , Room Features , Navigation instructions The constructed topological map is ; There are three types of nodes, namely the nodes that have been visited, the nodes where the agent is currently located, and the nodes that the agent can reach next; S2. Action-aware coding: Based on the positional relationship between the agent's current node and any candidate node, encode the changes in the deflection angle and pitch angle to the candidate node; S3. Feature fusion: combining object features , Room Features And the encoded action features, obtain visual features with action information , and the topological map Make updates; S4. Dynamic instruction weighting: Calculate the weight of each word in the instruction based on visual features and attention mechanism, and calculate the weighted instruction features ; Word weight calculation: based on visual features Calculate the weight of each word in the instruction with the attention mechanism; First, get the attention weight of each word at time t : ; d yes The dimension is then max pooling is used to obtain the final weight , T is the transpose: ; Instruction feature weighting: Calculate the weighted instruction feature based on the weight of each word : ; Where L is the number of words in the instruction.

2. The visual language navigation method based on action feature alignment according to claim 1, characterized in that: Based on the positional relationship between the agent's current node and any candidate node, encode the change in its deflection angle to the candidate node, assuming , Is the path Two consecutive points on is a path of the agent from the current point to the i-th candidate point, , The coordinates are , , the current heading angle of the agent is , then the relative deflection angle The calculation method is as follows: ; ; function To convert radians to degrees, represents the deflection angle between two nodes, is 360, which is used to ensure The range is between 0 and 360 degrees; get after, Updated to , continue to calculate the path according to the above formula The deflection angles between the remaining adjacent points on the graph are finally obtained. , indicating that the agent is on the path The relative deflection angle between all adjacent points on the The number of median edges; Using an embedding layer Convert to feature encoding: ; For path The relative deflection angle encoding.

3. The visual language navigation method based on action feature alignment according to claim 2 is characterized in that: Based on the positional relationship between the agent's current node and any candidate node, encode the change in its pitch angle when it moves to the candidate node: Assumptions , The coordinates are , , the current elevation angle of the agent is , then the relative pitch angle The calculation method is as follows: ; ; is the depression angle between adjacent nodes, Indicates the maximum value of the relative pitch angle; get back, Updated to , continue to calculate the path according to the above formula The pitch angle changes between the remaining adjacent points on the graph are finally obtained. , indicating that the agent is on the path The relative pitch angles between all adjacent points on the graph when moving; Using an embedding layer Convert to feature encoding: ; For path The relative deflection angle encoding.

4. The visual language navigation method based on action feature alignment according to claim 3 is characterized in that: comprehensive and You can get the path Overall action characteristics , the formula is as follows: ; ; represents the action encoding of all paths, n is the number of paths, Represents an encoder consisting of a two-layer transformer structure.

5. The visual language navigation method based on action feature alignment according to claim 1, characterized in that: Visual feature fusion in step S3: combining object features , Room Features And the encoded action features , obtain visual features with action information : ; Represents a splicing operation; Topology feature update: For the node where the agent is currently located, use the object feature and room features As its feature, for the candidate node of the agent's next step, use , And the overall action characteristics As its characteristic, as well as They represent the object features and room features in the visible view, respectively.

6. The visual language navigation method based on action feature alignment according to claim 1, characterized in that: The visual features found at time t , instruction characteristics And topological maps Put in a Dual-scale GraphTransformer network to predict the next action of the agent. The formula is as follows: ; in, represents the next action to be performed by the agent, Represents Dual-scale GraphTransformer, which can understand visual features simultaneously , weighted instruction features and topological maps .

7. A visual language navigation system based on action feature alignment, characterized in that: Used to implement the visual language navigation method based on action feature alignment as described in any one of claims 1-6.

8. A storage medium, characterized in that: When the storage medium is run on a computer, the steps in the visual language navigation method based on action feature alignment according to any one of claims 1 to 6 are executed.

Citation Information

Patent Citations

  • Visual language navigation system and method for motion prompt based on modal alignment

    CN114973402A

  • Multi-level cross-media fusion visual language navigation method

    CN118758310A