A visual language navigation system and method based on action prompts of modal alignment
Through the action prompt system of modal alignment, modal alignment loss and continuous consistency loss are used, and image and text features are aligned with CLIP model, solving the problem of insufficient modal alignment in visual language navigation, and improving navigation performance and generalization capabilities.
Patent Information
- Application Number
- CN202210467461.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-04-29
AI Technical Summary
The existing visual language navigation methods are not robust in action decisions in dynamic scenarios, making it difficult to achieve accurate modal alignment, affecting navigation performance and generalization capabilities.
The action prompt system based on modal alignment is adopted, and the action prompt set generates module, a visual language navigation module and an optimization learning module for modal alignment action prompts are used to use modal alignment losses and continuous consistency losses to force the agent to explicitly learn cross-modal action knowledge, and combines the CLIP model to align image and text features.
It improves the accuracy and generalization of the agent's action decision-making ability in visual language navigation tasks, improves navigation performance, and has good interpretability.
Smart Images

Figure CN114973402B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual language navigation, and more specifically, to a visual language navigation system and method based on modal alignment action prompts. Background Art
[0002] Vision-language navigation is a challenging task that requires an embodied agent to navigate to a target location following natural language instructions. For successful navigation, the agent should make correct action decisions sequentially to move in dynamically changing scenes by understanding the intent of the given instructions and gradually basing the instructions on the surrounding observations.
[0003] Early visual-language navigation methods explored different data augmentation strategies, efficient learning paradigms, and useful model architectures to improve agent performance. Inspired by the significant progress made in large-scale cross-modal pre-training models in visual-language tasks, more and more works have attempted to introduce pre-training paradigms and models into visual-language navigation tasks. PREVALENT performs self-supervised pre-training on a large number of image-language-action triplets. Introducing loop functions in pre-trained models makes the agents time-aware. Although object-level alignment capabilities may be significantly improved during pre-training, these agents are still implicitly learning action-level modality alignment, which greatly limits the robustness of action decisions in different scenarios.
[0004] The prior art discloses a patent for a navigation method that integrates vision and language multimodality. The patent belongs to the fields of robot navigation, natural language processing, and computer vision. The patent first installs a binocular camera on the robot, and uses the robot to train a multimodal fusion neural network model. The patent selects any real scene, issues natural language navigation instructions to the robot, and converts them into corresponding semantic vectors. The RGB images obtained by the robot at each moment are converted into corresponding features. The semantic vector and RGB image features are fused to obtain the action features at the current moment. After the action features are corrected using prompts, the neural network model finally outputs the robot's action at the current moment, and the robot executes the action until the navigation task is completed. However, the patent rarely reports on how to implement a more robust visual language navigation model, improve accuracy and generalization, and have good explanatory power. Summary of the invention
[0005] The present invention provides a visual language navigation system based on modality-aligned action cues, which can force an intelligent agent to explicitly learn cross-modal action knowledge to improve action decisions during navigation.
[0006] Another object of the present invention is to provide a navigation method for the above system.
[0007] In order to achieve the above technical effects, the technical solution of the present invention is as follows:
[0008] A visual language navigation system based on modality-aligned action cues, comprising:
[0009] The action prompt set generation module inputs the instruction to the action prompt set generation module, and the intelligent agent retrieves the action prompt set related to the instruction from the action prompt library before starting navigation;
[0010] The visual language navigation module of modality-aligned action prompts, the action prompt set passes through the prompt encoding module, and the output prompt features are connected with the output instruction features of the text encoding module; the instruction features based on the prompts and the output visual features of the visual encoding module are provided to the multi-layer transformer for action decision;
[0011] The learning modules, namely the modality alignment loss module and the continuity consistency loss module, are optimized to achieve effective action cue learning.
[0012] Furthermore, the visual language navigation module of the modal alignment action prompt includes:
[0013] Text encoding module: This module receives the input of language information and encodes it using a multi-layer transformer neural network to obtain the corresponding feature vector.
[0014] The prompt decoding module consists of two unimodal sub-prompt encoders and one multimodal prompt encoder. The image sub-prompt and text sub-prompt are respectively obtained through the corresponding unimodal autoencoders to obtain sub-prompt features, which are then connected and input into the multimodal prompt encoder to obtain prompt features.
[0015] The visual encoding module receives the input of visual observation information, encodes it through the visual encoder, and obtains the corresponding feature vector.
[0016] Furthermore, the optimization learning module includes:
[0017] Modality alignment loss module: when an action cue already has matching image and text sub-cues, InfoNCE loss is used to align them in feature space, so that the action cue can become more discriminative;
[0018] The sequential consistency loss module encourages the agent to sequentially attend to relevant action cues in the retrieved cue set based on its observations.
[0019] A visual language navigation method based on modality-aligned action prompts includes the following steps:
[0020] S1: At the beginning of navigation, the agent obtains instructions and retrieves the action prompt set related to the instructions from the action prompt library through the action prompt generation module;
[0021] S2: Through the visual encoding module and the text encoding module, the neural network encodes the input image information and instruction information respectively to obtain visual encoding, instruction encoding, and state features respectively;
[0022] S3: Through the prompt encoding model, the image sub-prompt and the text sub-prompt in the action prompt set are respectively obtained through the corresponding single-modal autoencoder to obtain sub-prompt features, which are then connected and input into the multi-modal prompt encoder to obtain prompt features;
[0023] S4: Connect the above instruction code and the prompt code to obtain a prompt-based instruction feature, and connect the above state feature with the visual code to obtain a state visual feature;
[0024] S5: Visual-language navigation module with modality-aligned action cues. The state visual feature is updated based on the cross-modal attention between itself and the cue-based instruction feature. The attention is decomposed into two parts. The first part weights the instruction encoding to update the state feature. The second part weights the image and text sub-cue features to calculate the sequential consistency loss. The state visual feature is input into another self-attention module to obtain the attention score of the state feature on the visual feature, that is, the action prediction probability based on the cue.
[0025] S6: By optimizing the learning model, combining the commonly used imitation learning loss and reinforcement learning loss, as well as the modal alignment loss and continuous consistency loss unique to the present invention, a weighted sum is performed to obtain the total training goal, the model is updated and optimized, and the navigation performance and generalization ability of the intelligent agent are improved.
[0026] Furthermore, the step S1 includes the following sub-steps:
[0027] S100: Construction of action prompt library. In order to align images and action phrases and form action prompts, a two-branch scheme is designed to collect image and text sub-prompts: First, for an instruction path instance in the training dataset, a pre-created visual object / location vocabulary is used to find the visual object / location mentioned in the instruction. For each visual object / location, the relevant image and text sub-prompts are obtained respectively. CLIP with excellent 0-shot cross-modal alignment capability is used to locate the image related to the object / location. In order to adapt to the reasoning process of CLIP, the token {CLASS} in the phrase "a photo of {CLASS}" is replaced by a visual object / location with a category label of c. The probability that an image B belongs to class c in the action sequence is calculated by the following method:
[0028]
[0029] Where τ1 is the temperature parameter, sim is the cosine similarity, b, w c The image features and phrase features generated by CLIP are respectively, M is the size of the vocabulary, and then the image with the greatest similarity to the phrase is selected as the image sub-cue. In order to obtain the text sub-cue, a simple nearest verb search scheme is used, that is, to find the nearest verb before a specific object / position word, which is in the pre-built verb vocabulary. Finally, the image and text sub-cue with the same visual object / position and action form an aligned action cue;
[0030] S101: Retrieval of action prompt set. At the beginning of navigation, the agent retrieves action prompts related to the instruction from the action prompt library, and calculates the sentence similarity between each object / location-related action phrase and the text sub-prompt in the prompt library to retrieve the action prompt set related to the instruction. Where N is the size of the set.
[0031] Furthermore, the step S2 includes the following sub-steps:
[0032] S200: Encoding of visual input, for each image view O in the candidate view at time step t t,i , will use a pre-trained convolutional neural network CNN or transformer to extract image features v t,i , then v t,i The visual encoder F v Mapping to visual encoding:
[0033] V t,i =F v (v t,i θv )
[0034] where θ v F v Parameters, a set Represents the candidate visual encoding at time t;
[0035] S201: Encoding of language input. During initialization, the instruction encoding X and the initialized state feature s0 are obtained by inputting the instruction sequence I and [CLS] and [SEP] tokens to the self-attention module in the transformer:
[0036]
[0037] Concat(·) represents the concatenation operation. Represents the parameters of the self-attention module, s0 will be updated to s at time step t t .
[0038] Furthermore, the step S3 includes the following steps:
[0039] use Get the prompt code through the prompt encoder The cue encoder consists of two unimodal sub-cue encoders and a multimodal cue encoder. The image sub-prompt and text sub-prompt are and and First, the sub-cue features are obtained through a unimodal sub-cue encoder and
[0040]
[0041]
[0042] Where E i (·) Using parameter θ i , E u (·) Using parameter θ u , respectively represent the image sub-cue encoder and the text sub-cue encoder, and then and Sent to the multimodal prompt encoder E p (·), get the prompt code
[0043]
[0044] where θ pFor E p (·) parameter, Concat(·) is the concatenation operation, encoder E i (·), E u (·) and E p (·) Consists of a linear layer followed by a dropout operation to reduce overfitting.
[0045] Furthermore, the step S4 includes the following sub-steps:
[0046] In the prompt code and instruction encoding X, by simply replacing X with Connect them together to get the prompt-based instruction feature X p .
[0047] Furthermore, the step S5 includes the following sub-steps:
[0048] State visual feature K t Based on K t and X p Cross-modal attention between renew:
[0049]
[0050] Then Decompose into and Obtain different features based on attention mechanism enhancement and participate in instruction features By performing Weighted image sub-cue features enhanced by attention mechanism and text sub-prompt features enhanced by attention mechanism Through and conduct Weighted gain, and Used to calculate the sequential consistency loss L c , like the baseline agent, Used to update the state characteristics, and finally, enter Get the action prediction probability based on the prompt
[0051] Furthermore, the step S6 includes the following sub-steps:
[0052] S600: Modality alignment loss, which encourages action cues to have matching image and text sub-cues aligned in feature space. Following the contrastive learning paradigm used in CLIP, it makes paired image and text features similar, while unpaired image and text features alienated. InfoNCE loss is used to promote feature alignment of image and text sub-cues in each action cue:
[0053]
[0054] Where τ2 is the temperature parameter, Indicates action prompt p n The features of paired image and text sub-cues, Representing unpaired sub-cues, action cues can become more discriminative through modality alignment loss, thus learning action-level modality alignment;
[0055] S601: Sequential consistency loss. Since instructions usually point to different visual landmarks sequentially, the retrieved action prompt set {p n The action prompts in} are also related to different objects / positions. In order to encourage the agent to pay attention to the relevant action prompts in the retrieved prompt set in sequence according to its observations, a sequential consistency loss is proposed, which is the sum of two unimodal consistency losses; at each time step t, the text sub-prompt features enhanced by the attention mechanism And the guidance features enhanced by the attention mechanism Must be close to:
[0056]
[0057] definition, Used to improve image sub-cue features based on attention mechanism enhancement and the similarity between visual features enhanced by the attention mechanism, the sequential consistency loss L c for:
[0058]
[0059] S602: Total target uses navigation loss L n , that is, the imitation loss L IL and reinforcement learning loss L RL , the overall training goal is:
[0060] L=L RL +λ1L IL +λ2L c +λ3L a
[0061] where λ1, λ2 and λ3 are the loss weights for balancing the loss.
[0062] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0063] This paper proposes modality-aligned action prompts to force the agent to explicitly learn cross-modal action knowledge to improve action decisions during navigation, and develops prompt-based navigation in visual language navigation tasks; develops a modality alignment loss and a continuous consistency loss to achieve effective learning of action prompts. The contrastive language-image pre-training (CLIP) model is used to ensure the quality of action prompts; effectively improves the navigation performance of R2R and RxR-based agents, and has good interpretability and generalization capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 This is an architecture diagram of a visual language navigation system based on modality-aligned action prompts of the present invention;
[0065] Figure 2 A flowchart of the steps of a visual language navigation method based on modality-aligned action prompts of the present invention;
[0066] Figure 3 This is an example diagram of a visual language navigation module for modal alignment action prompts in a specific embodiment of the present invention;
[0067] Figure 4 An example diagram of the construction of an action prompt library of an action prompt set generation module in a specific embodiment of the present invention;
[0068] Figure 5 The following is a comparative display of the result samples of the visual language navigation method and the baseline method in the specific embodiment of the present invention. DETAILED DESCRIPTION
[0069] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;
[0070] In order to better illustrate the present embodiment, some parts in the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product;
[0071] It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0072] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0073] Example 1
[0074] like Figure 1 As shown, a visual language navigation system based on modality-aligned action prompts includes:
[0075] The action prompt collection module 10 uses a recently developed contrastive language image pre-trained CLIP model with powerful cross-modal object / position level alignment capabilities to locate object / position related images in order to produce a high-quality action prompt library. In order to better align images and action phrases to form action prompts, a two-branch scheme is designed to collect image and text sub-prompts. First, for an instruction path instance in the training data set, a pre-created visual object / position vocabulary is used to find the visual objects / positions mentioned in the instruction. Then for each visual object / position, the relevant image and text sub-prompts are obtained respectively. At the beginning of navigation, the instruction is input into the action prompt set generation module, and the intelligent agent retrieves the action prompts related to the instruction from the pre-built action prompt library to form an action prompt set.
[0076] The visual language navigation module 11 of the modality-aligned action prompt obtains the prompt feature through a prompt encoder, and connects it with the output instruction feature of the text encoding module to obtain the prompt-based instruction feature. This feature and the output visual feature of the visual encoding module are provided to the multi-layer transformer for action decision making.
[0077] The learning module 12, namely the modality alignment loss module and the continuity consistency loss module, is optimized to achieve effective action cue learning.
[0078] In a specific embodiment of the present invention, specifically, the visual language navigation module 11 of the modal alignment action prompt further includes:
[0079] The text encoding module 110 receives the input of language information, encodes it using a self-supervised neural network, and obtains corresponding text feature vectors and state features.
[0080] The prompt encoding module 111 is composed of two unimodal sub-prompt encoders and one multimodal prompt encoder. The image sub-prompt and the text sub-prompt obtain sub-prompt features through the corresponding unimodal autoencoders respectively, and then connected and input into the multimodal prompt encoder to obtain the prompt features.
[0081] The visual encoding module 112 receives the input of visual observation information, encodes it through a pre-trained visual feature encoder, and obtains a corresponding feature vector.
[0082] In a specific embodiment of the present invention, specifically, the optimization learning module 12 further includes:
[0083] Modality Alignment Loss Module 120,When an action cue already has matching image and text sub-cues, they may not be aligned in feature space. To address this issue, following the contrastive learning paradigm used in CLIP, which makes paired image and text features similar and unpaired image and text features distant, an infoNCE loss is used to promote feature alignment of image and text sub-cues in each action cue. Through the modality alignment loss, action cues can become more discriminative, thus learning action-level modality alignment.
[0084] Continuous consistency loss module 121, since instructions usually point to different visual landmarks sequentially, the action prompts in the retrieved action prompt set are also related to different objects / positions. In order to encourage the agent to pay attention to the relevant action prompts in the retrieved prompt set in sequence according to its observations, a sequential consistency loss is proposed, which is the sum of two unimodal consistency losses. Taking the text modality as an example, at each time step t, the text sub-prompt features and the guidance features must be close; similar losses are defined in the image modality to improve the similarity between the image sub-prompt features and the visual features.
[0085] Example 2
[0086] like Figure 2 As shown, a visual language navigation method based on modality-aligned action prompts includes the following steps:
[0087] Step S1, searching for a relevant action prompt set according to input instruction information.
[0088] Specifically, step S1 further includes:
[0089] Step S100, construction of action prompt library. In order to better align images and action phrases to form action prompts, a two-branch scheme is designed to collect image and text sub-prompts. First, for an instruction path instance in the training dataset, a pre-created visual object / location vocabulary is used to find the visual objects / locations mentioned in the instruction. Then for each visual object / location, the relevant image and text sub-prompts are obtained respectively, as described below.
[0090] Note that the ground-truth path sequence contains a set of single-view images, each of which represents an action that needs to be performed at a specific time step. Therefore, to derive image sub-cues in action cues, only images related to objects / locations, which themselves contain action information, are retrieved from the ground-truth path sequence. Instead of resorting to existing object classifiers or detectors trained on a fixed set of item categories, CLIP, with its excellent 0-shot cross-modal alignment capabilities, is used to locate images related to objects / locations. To adapt the reasoning process of CLIP, the token {CLASS} in the phrase “a photo of {CLASS}” is replaced with a visual object / location whose class label is c. The probability that an image B belongs to class c in the action sequence is calculated as follows:
[0091]
[0092] Where τ1 is the temperature parameter, sim is the cosine similarity, b, w c are the image features and phrase features generated by CLIP, M is the size of the vocabulary, and then the image with the greatest similarity to the phrase is selected as the image sub-cue.
[0093] To obtain textual subcues, a simple nearest verb search scheme is used, i.e., finding the nearest verb (in a pre-built verb vocabulary) before a specific object / location word. Finally, image and textual subcues with the same visual object / location and action form an aligned action cue.
[0094] Step S101, retrieval of action prompt set. At the beginning of navigation, the agent retrieves action prompts related to the instruction from the action prompt library. The sentence similarity between each action phrase related to the object / location and the text sub-prompt in the prompt library is calculated to retrieve the action prompt set related to the instruction. Where N is the size of the set.
[0095] Step S2, encoding the input image information and instruction information respectively through a neural network.
[0096] Specifically, step S2 further includes:
[0097] Step S200, encoding of visual input, for each image view O in the candidate view at time step t t,i , will use a pre-trained convolutional neural network CNN or transformer to extract image features v t,i , then v t,i The visual encoder F vMapping to visual encoding:
[0098] V t,i =F v (v t,i θ v )
[0099] where θ v F v Parameters, a set Represents the candidate visual encoding at time t.
[0100] Step S201, encoding of language input. During initialization, the instruction encoding X and the initialized state feature s0 are obtained by inputting the instruction sequence I and [CLS] and [SEP] tokens to the self-attention module in the transformer:
[0101]
[0102] Concat(·) represents the concatenation operation. Represents the parameters of the self-attention module, s0 will be updated to s at time step t t .
[0103] Step S3, encoding the action prompt set through a modal encoder. Get the prompt code through the prompt encoder The cue encoder consists of two unimodal sub-cue encoders and a multimodal cue encoder. The image sub-prompt and text sub-prompt are and and First, the sub-cue features are obtained through a unimodal sub-cue encoder and
[0104]
[0105]
[0106] Where E i (·) Using parameter θ i , E u (·) Using parameter θ u , respectively represent the image sub-cue encoder and the text sub-cue encoder, and then and Sent to the multimodal prompt encoder E p (·), get the prompt code
[0107]
[0108] where θ p For E p (·) parameter, Concat(·) is the concatenation operation, encoder E i (·), E u (·) and E p (·) Consists of a linear layer followed by a dropout operation to reduce overfitting.
[0109] Step S4, in the prompt code and instruction encoding X, by simply replacing X with Connect them together to get the prompt-based instruction feature X p .
[0110] Step S5: state visual feature K t Based on K t and X p Cross-modal attention between renew:
[0111]
[0112] Then Decompose into and Obtain different features based on attention mechanism enhancement and participate in instruction features By performing Weighted image sub-cue features enhanced by attention mechanism and text sub-prompt features enhanced by attention mechanism Through and conduct Weighted gain, and Used to calculate the sequential consistency loss L c , like the baseline agent, Used to update the state characteristics, and finally, enter Get the action prediction probability based on the prompt
[0113] Step S6, calculate the weighted sum of the losses and update and optimize the model to improve the navigation performance and generalization ability of the intelligent agent.
[0114] Specifically, step S6 further includes:
[0115] Step S600, modality alignment loss, forces action cues to align already matched image and text sub-cues in feature space, following the contrastive learning paradigm used in CLIP, making paired image and text features similar, while unpaired image and text features distant, using infoNCE loss to promote feature alignment of image and text sub-cues in each action cue:
[0116]
[0117] Where τ2 is the temperature parameter, Indicates action prompt p n The features of paired image and text sub-cues, Representing unpaired sub-cues, action cues can become more discriminative through the modality alignment loss, thus learning action-level modality alignment.
[0118] Step S601, sequential consistency loss. Since instructions usually point to different visual signs sequentially, the retrieved action prompt set {p n The action prompts in} are also related to different objects / positions. In order to encourage the agent to pay attention to the relevant action prompts in the retrieved prompt set in sequence according to its observations, a sequential consistency loss is proposed, which is the sum of two unimodal consistency losses; at each time step t, the text sub-prompt features enhanced by the attention mechanism And the guidance features enhanced by the attention mechanism Must be close to:
[0119]
[0120] definition, Used to improve image sub-cue features based on attention mechanism enhancement and the similarity between visual features enhanced by the attention mechanism, the sequential consistency loss L c for:
[0121]
[0122] Step S602, the total target uses the navigation loss L n , that is, the imitation loss L IL and reinforcement learning loss L RL , the overall training goal is:
[0123] L=L RL +λ1L IL +λ2L c +λ3L a
[0124] where λ1, λ2 and λ3 are the loss weights for balancing the loss.
[0125] Example 3
[0126] like Figure 1 As shown, a visual language navigation system based on modality-aligned action prompts includes:
[0127] The action prompt collection module 10 uses a recently developed contrastive language image pre-trained CLIP model with powerful cross-modal object / position level alignment capabilities to locate object / position related images in order to produce a high-quality action prompt library. In order to better align images and action phrases to form action prompts, a two-branch scheme is designed to collect image and text sub-prompts. First, for an instruction path instance in the training data set, a pre-created visual object / position vocabulary is used to find the visual objects / positions mentioned in the instruction. Then for each visual object / position, the relevant image and text sub-prompts are obtained respectively. At the beginning of navigation, the instruction is input into the action prompt set generation module, and the intelligent agent retrieves the action prompts related to the instruction from the pre-built action prompt library to form an action prompt set.
[0128] The visual language navigation module 11 of the modality-aligned action prompt obtains the prompt feature through a prompt encoder, and connects it with the output instruction feature of the text encoding module to obtain the prompt-based instruction feature. This feature and the output visual feature of the visual encoding module are provided to the multi-layer transformer for action decision making.
[0129] The learning module 12, namely the modality alignment loss module and the continuity consistency loss module, is optimized to achieve effective action cue learning.
[0130] In a specific embodiment of the present invention, specifically, the visual language navigation module 11 of the modal alignment action prompt further includes:
[0131] The text encoding module 110 receives the input of language information, encodes it using a self-supervised neural network, and obtains corresponding text feature vectors and state features.
[0132] The prompt encoding module 111 is composed of two unimodal sub-prompt encoders and one multimodal prompt encoder. The image sub-prompt and the text sub-prompt obtain sub-prompt features through the corresponding unimodal autoencoders respectively, and then connected and input into the multimodal prompt encoder to obtain the prompt features.
[0133] The visual encoding module 112 receives the input of visual observation information, encodes it through a pre-trained visual feature encoder, and obtains a corresponding feature vector.
[0134] In a specific embodiment of the present invention, specifically, the optimization learning module 12 further includes:
[0135] Modality Alignment Loss Module 120,When an action cue already has matching image and text sub-cues, they may not be aligned in feature space. To address this issue, following the contrastive learning paradigm used in CLIP, which makes paired image and text features similar and unpaired image and text features distant, an infoNCE loss is used to promote feature alignment of image and text sub-cues in each action cue. Through the modality alignment loss, action cues can become more discriminative, thus learning action-level modality alignment.
[0136] Continuous consistency loss module 121, since instructions usually point to different visual landmarks sequentially, the action prompts in the retrieved action prompt set are also related to different objects / positions. In order to encourage the agent to pay attention to the relevant action prompts in the retrieved prompt set in sequence according to its observations, a sequential consistency loss is proposed, which is the sum of two unimodal consistency losses. Taking the text modality as an example, at each time step t, the text sub-prompt features and the guidance features must be close; similar losses are defined in the image modality to improve the similarity between the image sub-prompt features and the visual features.
[0137] like Figure 2 As shown, the navigation method of the visual language navigation system based on the action prompt of modal alignment comprises the following steps:
[0138] Step S1, searching for a relevant action prompt set according to input instruction information.
[0139] Specifically, step S1 further includes:
[0140] Step S100, construction of action prompt library. In order to better align images and action phrases to form action prompts, a two-branch scheme is designed to collect image and text sub-prompts. First, for an instruction path instance in the training dataset, a pre-created visual object / location vocabulary is used to find the visual objects / locations mentioned in the instruction. Then for each visual object / location, the relevant image and text sub-prompts are obtained respectively, as described below.
[0141] Note that the ground-truth path sequence contains a set of single-view images, each of which represents an action that needs to be performed at a specific time step. Therefore, to derive image sub-cues in action cues, only images related to objects / locations, which themselves contain action information, are retrieved from the ground-truth path sequence. Instead of resorting to existing object classifiers or detectors trained on a fixed set of item categories, CLIP, with its excellent 0-shot cross-modal alignment capabilities, is used to locate images related to objects / locations. To adapt the reasoning process of CLIP, the token {CLASS} in the phrase “a photo of {CLASS}” is replaced with a visual object / location whose class label is c. The probability that an image B belongs to class c in the action sequence is calculated as follows:
[0142]
[0143] Where τ1 is the temperature parameter, sim is the cosine similarity, b, w c are the image features and phrase features generated by CLIP, M is the size of the vocabulary, and then the image with the greatest similarity to the phrase is selected as the image sub-cue.
[0144] To obtain textual subcues, a simple nearest verb search scheme is used, i.e., finding the nearest verb (in a pre-built verb vocabulary) before a specific object / location word. Finally, image and textual subcues with the same visual object / location and action form an aligned action cue.
[0145] Step S101, retrieval of action prompt set. At the beginning of navigation, the agent retrieves action prompts related to the instruction from the action prompt library. The sentence similarity between each action phrase related to the object / location and the text sub-prompt in the prompt library is calculated to retrieve the action prompt set related to the instruction. Where N is the size of the set.
[0146] Step S2, encoding the input image information and instruction information respectively through a neural network.
[0147] Specifically, step S2 further includes:
[0148] Step S200, encoding of visual input, for each image view O in the candidate view at time step t t,i , will use a pre-trained convolutional neural network CNN or transformer to extract image features v t,i , then v t,i The visual encoder F vMapping to visual encoding:
[0149] V t,i =F v (v t,i θ v )
[0150] where θ v F v Parameters, a set Represents the candidate visual encoding at time t.
[0151] Step S201, encoding of language input. During initialization, the instruction encoding X and the initialized state feature s0 are obtained by inputting the instruction sequence I and [CLS] and [SEP] tokens to the self-attention module in the transformer:
[0152]
[0153] Concat(·) represents the concatenation operation. Represents the parameters of the self-attention module, s0 will be updated to s at time step t t .
[0154] Step S3, encoding the action prompt set through a modal encoder. Get the prompt code through the prompt encoder The cue encoder consists of two unimodal sub-cue encoders and a multimodal cue encoder. The image sub-prompt and text sub-prompt are and and First, the sub-cue features are obtained through a unimodal sub-cue encoder and
[0155]
[0156]
[0157] Where E i (·) Using parameter θ i , E u (·) Using parameter θ u , respectively represent the image sub-cue encoder and the text sub-cue encoder, and then and Sent to the multimodal prompt encoder E p (·), get the prompt code
[0158]
[0159] where θ p For E p (·) parameter, Concat(·) is the concatenation operation, encoder E i (·), E u (·) and E p (·) Consists of a linear layer followed by a dropout operation to reduce overfitting.
[0160] Step S4, in the prompt code and instruction encoding X, by simply replacing X with Connect them together to get the prompt-based instruction feature X p .
[0161] Step S5: state visual feature K t Based on K t and X p Cross-modal attention between renew:
[0162]
[0163] Then Decompose into and Obtain different features based on attention mechanism enhancement and participate in instruction features By performing Weighted image sub-cue features enhanced by attention mechanism and text sub-prompt features enhanced by attention mechanism Through and conduct Weighted gain, and Used to calculate the sequential consistency loss L c , like the baseline agent, Used to update the state characteristics, and finally, enter Get the action prediction probability based on the prompt
[0164] Step S6, calculate the weighted sum of the losses and update and optimize the model to improve the navigation performance and generalization ability of the intelligent agent.
[0165] Specifically, step S6 further includes:
[0166] Step S600, modality alignment loss, forces action cues to align already matched image and text sub-cues in feature space, following the contrastive learning paradigm used in CLIP, making paired image and text features similar, while unpaired image and text features distant, using infoNCE loss to promote feature alignment of image and text sub-cues in each action cue:
[0167]
[0168] Where τ2 is the temperature parameter, Indicates action prompt p n The features of paired image and text sub-cues, Representing unpaired sub-cues, action cues can become more discriminative through the modality alignment loss, thus learning action-level modality alignment.
[0169] Step S601, sequential consistency loss. Since instructions usually point to different visual signs sequentially, the retrieved action prompt set {p n The action prompts in} are also related to different objects / positions. In order to encourage the agent to pay attention to the relevant action prompts in the retrieved prompt set in sequence according to its observations, a sequential consistency loss is proposed, which is the sum of two unimodal consistency losses; at each time step t, the text sub-prompt features enhanced by the attention mechanism And the guidance features enhanced by the attention mechanism Must be close to:
[0170]
[0171] definition, Used to improve image sub-cue features based on attention mechanism enhancement and the similarity between visual features enhanced by the attention mechanism, the sequential consistency loss L c for:
[0172]
[0173] Step S602, the total target uses the navigation loss L n , that is, the imitation loss L IL and reinforcement learning loss L RL , the overall training goal is:
[0174] L=L RL +λ1L IL +λ2L c +λ3L a
[0175] where λ1, λ2 and λ3 are the loss weights for balancing the loss.
[0176] Figure 3 This is an example diagram of a visual language navigation module for modal alignment action prompts in a specific embodiment of the present invention.
[0177] This figure shows the comparison of action decisions between the baseline agent and the proposed agent. With the help of the action prompt related to "walking towards the stairs", the proposed agent can choose the correct action to successfully navigate in the given observation.
[0178] Figure 4 This is an example diagram of the construction of an action prompt library of an action prompt set generation module in a specific embodiment of the present invention.
[0179] The present invention uses a two-branch scheme to collect image and text sub-cues. First, for an instruction path instance in the training dataset, a recently developed contrastive language image pre-trained CLIP model with powerful cross-modal object / location level alignment capabilities is used to replace the token {CLASS}token in the phrase "a photo of{CLASS}" with a visual object / location with a class label of c. The probability that an image B belongs to class c in the action sequence is calculated, and then the image with the greatest similarity to the phrase is selected as the image sub-cue. For text sub-cues, a nearest verb search scheme is used, that is, the nearest verb (in the pre-built verb vocabulary) before a specific object / location word is found.
[0180] Figure 5 The result sample comparison of the visual language navigation method and the baseline method in the specific embodiment of the present invention is presented. By introducing action prompts, the present invention can make accurate action decisions and complete successful navigation. With the help of action prompts related to "walking past the window", the present invention performs the correct "walking past the window" action in the first two navigation steps. However, the baseline agent failed to perform the "walking past the window" action during navigation, resulting in an incorrect trajectory.
[0181] The same or similar reference numerals correspond to the same or similar components;
[0182] The positional relationships described in the drawings are only for illustrative purposes and should not be construed as limiting the present patent.
[0183] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A visual language navigation system based on modality-aligned action prompts, characterized in that: include: The action prompt set generation module inputs the instruction to the action prompt set generation module, and the intelligent agent retrieves the action prompt set related to the instruction from the action prompt library before starting navigation; The visual language navigation module of modality-aligned action prompts, the action prompt set passes through the prompt encoding module, and the output prompt encoding is connected with the output instruction features of the text encoding module; the instruction features based on the prompts and the output visual features of the visual encoding module are provided to the multi-layer transformer for action decision; Optimize the learning modules, namely the modality alignment loss module and the continuity consistency loss module, to achieve effective action cue learning; The visual language navigation module of the modal alignment action prompt includes: The text encoding module receives the input of language information and encodes it using a multi-layer transformer neural network to obtain the corresponding feature vector; The prompt encoding module consists of two unimodal sub-prompt encoders and one multimodal prompt encoder. The image sub-prompt and text sub-prompt are respectively obtained through the corresponding unimodal autoencoders to obtain sub-prompt features, which are then connected and input into the multimodal prompt encoder to obtain prompt codes. The visual encoding module receives the input of visual observation information, encodes it through the visual encoder, and obtains the corresponding feature vector.
2. The visual language navigation system based on modality-aligned action prompts according to claim 1, characterized in that: The optimization learning module includes: Modality alignment loss module: when an action cue already has matching image and text sub-cues, InfoNCE loss is used to align them in feature space, so that the action cue can become more discriminative; The sequential consistency loss module encourages the agent to sequentially attend to relevant action cues in the retrieved cue set based on its observations.
3. A visual language navigation method using the system of claim 2, characterized in that: The following steps are involved: S1: At the beginning of navigation, the agent obtains instructions and retrieves action prompt sets related to the instructions from the action prompt library through the action prompt set generation module; S2: Encode the input visual observation information and language information through the visual encoding module and the text encoding module, respectively, to obtain visual encoding, instruction encoding, and state features; S3: Through the prompt encoding module, the image sub-prompt and the text sub-prompt in the action prompt set are respectively obtained through the corresponding single-modal autoencoder to obtain sub-prompt features, which are then connected and input into the multi-modal prompt encoder to obtain the prompt code; S4: connect the instruction code and the prompt code to obtain the prompt-based instruction feature, and connect the state feature with the visual code to obtain the state visual feature; S5: Visual-language navigation module with modality-aligned action cues. The state visual feature is updated based on the cross-modal attention between itself and the cue-based instruction feature. The attention is decomposed into two parts. The first part weights the instruction encoding to update the state feature. The second part weights the image and text sub-cue features to calculate the sequential consistency loss. The state visual feature is input into another self-attention module to obtain the attention score of the state feature on the visual feature, that is, the action prediction probability based on the cue. S6: By optimizing the learning module, combining the imitation learning loss and reinforcement learning loss, as well as the modality alignment loss and the continuous consistency loss, a weighted sum is performed to obtain the total training goal, and the model is updated and optimized to improve the navigation performance and generalization ability of the intelligent agent.
4. The visual language navigation method according to claim 3, characterized in that: The step S1 comprises the following sub-steps: S100: Construction of action prompt library. In order to align images and action phrases and form action prompts, a two-branch scheme is designed to collect image and text sub-prompts: First, for an instruction path instance in the training dataset, a pre-created visual object / location vocabulary is used to find the visual object / location mentioned in the instruction. For each visual object / location, the relevant image and text sub-prompts are obtained respectively. CLIP with excellent 0-shot cross-modal alignment capability is used to locate the image related to the object / location. In order to adapt to the reasoning process of CLIP, the token {CLASS} in the phrase "a photo of{CLASS}" is replaced with a visual object / location with a category label of c. The probability that an image B belongs to class c in the action sequence is calculated by the following method: Where τ1 is the temperature parameter, sim is the cosine similarity, b, w c The image features and phrase features generated by CLIP are respectively, M is the size of the vocabulary, and then the image with the greatest similarity to the phrase is selected as the image sub-cue. In order to obtain the text sub-cue, a simple nearest verb search scheme is used, that is, to find the nearest verb before a specific object / position word, which is in the pre-built verb vocabulary. Finally, the image and text sub-cue with the same visual object / position and action form an aligned action cue; S101: Retrieval of action prompt set. At the beginning of navigation, the agent retrieves action prompts related to the instruction from the action prompt library, and calculates the sentence similarity between each object / location-related action phrase and the text sub-prompt in the prompt library to retrieve the action prompt set related to the instruction. Where N is the size of the set.
5. The visual language navigation method according to claim 3, characterized in that: The step S2 comprises the following sub-steps: S200: Encoding of visual input, for each image view O in the candidate view at time step t t,i , will use a pre-trained convolutional neural network CNN or transformer to extract image features v t,i , then v t,i The visual encoder F v Mapping to visual encoding: V t,i =F v (v t,i ;θ v ) where θ v F v Parameters, a set Represents the candidate visual encoding at time t; S201: Encoding of language input. During initialization, the instruction encoding X and the initialized state feature s0 are obtained by inputting the instruction sequence I and [CLS] and [SEP] tokens to the self-attention module in the transformer: Concat(·) represents the concatenation operation. Represents the parameters of the self-attention module, s0 will be updated to s at time step t t .
6. The visual language navigation method according to claim 3, characterized in that: The step S3 comprises the steps of: use Get the prompt code through the prompt encoder The cue encoder consists of two unimodal sub-cue encoders and a multimodal cue encoder. The image sub-prompt and text sub-prompt are and and First, the sub-cue features are obtained through a unimodal sub-cue encoder and Where E i (·) Using parameter θ i , E u (·) Using parameter θ u , respectively represent the image sub-cue encoder and the text sub-cue encoder, and then and Sent to the multimodal prompt encoder E p (·), get the prompt code where θ p For E p (·) parameter, Concat(·) is the concatenation operation, encoder E i (·), E u (·) and E p (·) Consists of a linear layer followed by a dropout operation to reduce overfitting.
7. The visual language navigation method according to claim 3, characterized in that: The step S4 comprises the following sub-steps: In the prompt code and instruction encoding X, by simply replacing X with Connect them together to get the prompt-based instruction feature X p .
8. The visual language navigation method according to claim 3, characterized in that: The step S5 comprises the following sub-steps: State visual feature K t Based on K t and X p Cross-modal attention between renew: Then Decompose into and Obtain different features based on attention mechanism enhancement and participate in instruction features By performing Weighted image sub-cue features enhanced by attention mechanism and text sub-prompt features enhanced by attention mechanism Through and conduct Weighted gain, and Used to calculate the sequential consistency loss L c , like the baseline agent, Used to update the state characteristics, and finally, enter Get the action prediction probability based on the prompt 9. The visual language navigation method according to claim 3, characterized in that: The step S6 comprises the following sub-steps: S600: Modality alignment loss, which encourages action cues to have matching image and text sub-cues aligned in feature space. Following the contrastive learning paradigm used in CLIP, it makes paired image and text features similar, while unpaired image and text features alienated. InfoNCE loss is used to promote feature alignment of image and text sub-cues in each action cue: Where τ2 is the temperature parameter, Indicates action prompt p n The features of paired image and text sub-cues, Representing unpaired sub-cues, action cues can become more discriminative through modality alignment loss, thus learning action-level modality alignment; S601: Sequential consistency loss. Since instructions usually point to different visual landmarks sequentially, the retrieved action prompt set {p n The action prompts in} are also related to different objects / positions. In order to encourage the agent to pay attention to the relevant action prompts in the retrieved prompt set in sequence according to its observations, a sequential consistency loss is proposed, which is the sum of two unimodal consistency losses; at each time step t, the text sub-prompt features enhanced by the attention mechanism And the guidance features enhanced by the attention mechanism Must be close to: definition, Used to improve image sub-cue features based on attention mechanism enhancement and the similarity between visual features enhanced by the attention mechanism, the sequential consistency loss L c for: S602: Total target uses navigation loss L n , that is, the imitation loss L IL and reinforcement learning loss L RL , the overall training goal is: L=L RL +λ1L IL +λ2L c +λ3L a where λ1, λ2 and λ3 are the loss weights for balancing the loss.
Citation Information
Patent Citations
Visual language navigation system and method based on dynamic enhancement instruction attack module
CN113804200A