Continuous visual language navigation model and method based on visual representation of knowledge and historical perception

By introducing visual representations of knowledge and historical perception into visual language navigation, the problem of insufficient utilization of visual information in the prior art is solved, and the navigation success rate and generalization ability of agents in continuous environments is improved.

CN120368980APending Publication Date: 2025-07-25SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510515052.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing visual language navigation methods are difficult to effectively utilize the complex information of the current observation in a continuous environment, making it difficult for the agent to pay attention to the most effective visual information, affecting the accuracy of navigation action prediction.

Method used

Introduce visual representations of knowledge and history perception, and through topology map construction modules, cross-modal planning modules and path control modules, RGB image features, depth image features, knowledge features and historical features are extracted and fused to form multi-perception fusion features, enhancing the cross-modal alignment capabilities of language instructions and visual observations.

Benefits of technology

It improves the navigation success rate and generalization ability of the agent in an unseen environment, and enhances the perceived coherence of the navigation process and the accuracy of navigation instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120368980A_ABST
    Figure CN120368980A_ABST
Patent Text Reader

Abstract

The invention discloses a visual representation continuous vision language navigation model and method based on knowledge and historical perception, the continuous vision language navigation model comprises a topological graph construction module, a cross-modal planning module and a path control module, the topological graph construction module comprises an extraction module, a filtering module, an interaction module and an aggregation module; the extraction module is used for extracting an RGB image feature fr, a depth image feature fd, a first knowledge feature fk, a first historical feature fh and a navigation instruction fi; the filtering module calculates a correlation matrix among a first knowledge feature fk, a first historical feature fh and a navigation instruction fi to obtain a weighted second knowledge feature # imgabs0 # and a weighted second historical feature # imgabs1 #, and the interaction module interacts the second knowledge feature # imgabs2 # and the second historical feature # imgabs3 # with the instruction to obtain a multi-perception fusion feature fusion; the aggregation module aggregates the multi-perception fusion feature fusion, the RGB image feature fr and the depth image feature fd to obtain a visual representation fimg; according to the visual language navigation method provided by the invention, rich visual representations related to the navigation instruction are constructed, and the navigation capability of the intelligent agent is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention relates to the field of vision-language navigation, and particularly to a continuous vision-language navigation model and method based on knowledge and history-aware visual representation. Background Art:

[0002] Vision-and-Language Navigation (VLN) aims to enable an agent to navigate in a dynamic environment according to given natural language instructions and finally reach the target location. In the initial task setting, a navigation map is provided for each scene, and the agent only needs to teleport between a predefined few views, which is called discrete VLN [1]. To transition to a more practical scenario, Visual and Language Navigation in Continuous Environment (VLN-CE) [2] adopts a continuous setting with low-level actions, which more closely reflects the real-world situation.

[0003] Because many simplifying assumptions in discrete VLN are removed, the continuous VLN task is more challenging than the discrete VLN task, and many works have been proposed to bridge the gap between the two. Early VLN-CE works [3, 4] are end-to-end training systems that directly predict a low-level action. With the emergence and popularity of the locus predictor [5], many subsequent works [6, 7, 8] have applied this pre-trained model to predict candidate reachable viewpoints of the current viewpoint and modularize the whole system to predict actions.

[0004] Existing methods use the image information of the current observation of the viewpoint as the visual representation [7, 9, 10]. However, due to the intertwined redundant information and key information in the current observation, it is difficult for the agent to focus on the most effective visual information, which is insufficient for proper action prediction. Summary of the Invention:

[0005] Aiming at the technical problems existing in the prior art, the present invention proposes a continuous vision-language navigation method for an agent based on knowledge and history-aware visual representation. By introducing knowledge and historical information in the navigation view, the present invention constructs a visual representation that contains richer content and is most relevant to the navigation instruction, strengthens the cross-modal alignment ability between the language instruction and the visual observation, and the perceptual coherence between views during the navigation process, enabling the agent to make more accurate decisions.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A continuous vision-language navigation model based on knowledge and history-aware visual representation, the continuous vision-language navigation model includes a topological map construction module, a cross-modal planning module, and a path control module, and the topological map construction module includes an extraction module, a filtering module, an interaction module, and an aggregation module; wherein:

[0008] The extraction module is used to extract the RGB image feature f r , the depth image feature f d , the first knowledge feature f k , the first historical feature f h and the navigation instruction f i ;

[0009] The filtering module calculates the correlation matrices between the first knowledge feature f k , the first historical feature f h and the navigation instruction f i respectively, and obtains the weighted second knowledge feature and the second historical feature

[0010] The interaction module interacts the second knowledge feature the second historical feature with the instruction to obtain the multi-sensory fusion feature f fusion ;

[0011] The aggregation module aggregates the multi-sensory fusion feature f fusion , the RGB image feature f r and the depth image feature f d to obtain the visual representation f img .

[0012] Furthermore, the process of the extraction module for extracting the RGB image feature f r , the depth image feature f d , the first knowledge feature f k , the first historical feature f h and the navigation instruction f i includes:

[0013] At each step t, the current observed RGB image and depth image are obtained, and the pre-trained ViT-B / 32 model is used to encode the RGB image to obtain the RGB feature f r , and the pre-trained ResNet-50 model is used to encode the depth image to obtain the depth feature f d ;

[0014] Retrieve in the knowledge base KB, select the m pieces of knowledge with the highest cosine similarity to the RGB image o rgb and concatenate them to obtain, and calculate the knowledge feature f k according to the following formula:

[0015]

[0016] Where: kn represents knowledge encoding, and cos represents cosine similarity calculation;

[0017] Through the topological graph Map t The historical feature f is obtained by calculating the node view feature of the previous step, the node access time step encoding of the previous step, and the relative position encoding of the previous node and all nodes in the graph according to the following formula h ;

[0018]

[0019] Where: img, step, and pos represent view, access time step, and relative position respectively, and t - 1 represents the previous step of the current navigation step;

[0020] The topological graph Map t contains three types of nodes: the current node, the visited nodes, and the navigable nodes; the nodes in the topological graph record the observation information of the agent at each step, and the edges record the distance information between the nodes, which is updated during the navigation process.

[0021] Furthermore, the filtering module calculates the correlation matrices between the first knowledge feature f k , the first historical feature f h and the navigation instruction f i respectively, and obtains the weighted second knowledge feature and the second historical feature through the softmax layer, including;

[0022] Calculate the correlation matrices between the first knowledge feature f k , the first historical feature f h and the navigation instruction f i respectively. The calculation formula is as follows:

[0023] M k = f k W k (f i W i ) T

[0024] M h = f h W h (f i W i ) T

[0025] Where: k, h, and i represent knowledge, history, and instruction respectively, and W k , W h , W i are learnable parameters;

[0026] Based on the correlation matrix, the normalized exponential layer assigns correlation scores to the knowledge features and historical features, and this score will serve as a weight reflecting the association strength between the first knowledge feature and the first historical feature and the navigation instruction. Thus, the filtered weighted second knowledge feature can be generated. and the second historical feature The calculation formula is as follows:

[0027]

[0028] where softmax represents the normalized exponential layer and d represents the dimension of the features.

[0029] Furthermore, the interaction module obtains the multi-sensory fusion feature f by interacting the second knowledge feature the second historical feature fusion with the instruction. The process includes:

[0030] Calculating the knowledge-historical feature f through a multi-layer cross-modal encoder for the second knowledge feature the second historical feature kh

[0031]

[0032] Interacting the knowledge-historical feature f kh with the instruction according to the following formula to obtain the multi-sensory fusion feature f fusion :

[0033]

[0034] where: W kh , W cls are learnable parameters, and I0 represents the encoded head token, representing the overall semantic representation of the instruction text.

[0035] Furthermore, each layer of the multi-layer cross-modal encoder contains a cross-attention sub-layer, a self-attention sub-layer, and two feed-forward networks.

[0036] Furthermore, the aggregation module passes the multi-sensory fusion feature f fusion , the RGB image feature f r and the depth image feature f d through three feed-forward neural networks respectively, and aggregates the results to form the visual representation f img The calculation formula is as follows:

[0037] f img = FFN fusion (f fusiion ​) + FFN rgb (f r ) + FFN depth (f d )

[0038] Wherein: FFN represents a feed - forward neural network.

[0039] Beneficial effects

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] The present invention introduces external knowledge without specific rules, enabling the agent to pay more attention to the objects related to the current instruction, strengthening the cross - modal alignment ability between language instructions and visual observations, and improving the generalization ability of the agent.

[0042] The present invention explicitly adds past historical view information to the visual representation, providing navigation memory for the current visual observation to comprehensively describe the navigation view in the VLN - CE task, and improving the navigation success rate. Description of the drawings:

[0043] Figure 1 is a framework diagram of a continuous visual - language navigation model based on knowledge - and - history - aware visual representation according to the present invention;

[0044] Figure 2 is a structural diagram of a continuous visual - language navigation model based on knowledge - and - history - aware visual representation according to the present invention. Detailed implementation manners:

[0045] The present invention will be described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0046] As Figure 1 shown, the present invention provides a continuous visual - language navigation model based on knowledge - and - history - aware visual representation; the topological graph construction module includes an extraction module, a filtering module, an interaction module, and an aggregation module; wherein:

[0047] The extraction module is used to extract RGB image feature f r , depth image feature f d , first knowledge feature f k , first history feature f h , and navigation instruction f i ;

[0048] The filtering module respectively calculates the first knowledge feature f k , history feature f h , and navigation instruction f iThe correlation matrix between them is used to obtain the weighted second knowledge feature and the second historical feature

[0049] The interaction module obtains the multi-sensory fusion feature f by interacting the second knowledge feature the second historical feature with the instruction fusion ;

[0050] The aggregation module aggregates the multi-sensory fusion feature f fusion , the RGB image feature f r and the depth image feature f d to obtain the visual representation f img . Wherein:

[0051] At each step t, the agent obtains the current observed RGB image and depth image, and uses the pre-trained ViT-B / 32 model to encode the RGB image to obtain the RGB feature f r , and uses the pre-trained ResNet-50 model to encode the depth image to obtain the depth feature f d .

[0052] The first knowledge feature f k is obtained by retrieving in the knowledge base KB, calculating the cosine similarity between each piece of knowledge kn and the RGB image o rgb , selecting the top m pieces of knowledge and concatenating them. According to the experience of existing research, m = 5 is selected, and the calculation formula is as follows:

[0053]

[0054]

[0055] where; cos represents the calculation of cosine similarity.

[0056] The first historical feature f h is obtained by adding the node view feature of the previous step, the encoding of the access time step of the previous step, and the relative position encoding of the node where the agent is located and all nodes in the topological graph Map t . The calculation formula is as follows;

[0057]

[0058] where: img, step, pos represent the view, access time step, and relative position respectively, and t-1 represents the previous step of the current navigation step.

[0059] The RGB feature f r , the depth feature f d , the knowledge feature fk The historical feature f h all pass through different linear layers to unify the feature dimensions.

[0060] The topological graph Map t belongs to the topological graph construction part in the model structure. The graph contains three types of nodes: the current node, the visited node, and the navigable node; the nodes in the topological graph record the observation information of the agent at each step, and the edges record the distance information between the nodes, which is updated as the navigation process progresses.

[0061] The filtering module purifies the first knowledge feature and the first historical feature, removing redundant information that is ineffective for the current navigation task, and obtaining features highly relevant to the navigation instruction.

[0062] First, calculate the correlation matrices between the first knowledge feature f k , the first historical feature f h and the navigation instruction f i respectively. The calculation formula is as follows:

[0063] M k = f k W k (f i W i ) T

[0064] M h = f h W h (f i W i ) T

[0065] where: k, h, i represent knowledge, history, and instruction respectively, and W k , W h , W i are learnable parameters;

[0066] Then, based on the correlation matrix, assign correlation scores to the knowledge feature and the historical feature through the softmax layer. This score will be used as the weight reflecting the association strength between the first knowledge feature and the first historical feature and the navigation instruction, and thus the weighted second knowledge feature and the second historical feature can be generated. The calculation formula is as follows:

[0067]

[0068] where softmax represents the softmax layer and d represents the dimension of the feature.

[0069] The interaction module uses a multi-layer cross-modal encoder to obtain the knowledge-history feature fkh , the calculation formula is as follows:

[0070]

[0071] The multi - layer cross - modal encoder is a multi - layer network structure, and each layer contains a cross - attention sub - layer, a self - attention sub - layer, and two feed - forward networks.

[0072] Interact the knowledge - history feature f kh with the instruction, and calculate the correlation matrix between the knowledge - history feature f kh and the head token I0 after encoding the instruction. Assign correlation scores through the softmax layer to obtain the multi - perception fusion feature f fusion , the calculation formula is as follows:

[0073]

[0074] where W kh , W cls are learnable parameters, I0 represents the head token after encoding, representing the overall semantic representation of the instruction text.

[0075] The aggregation module passes the multi - perception fusion feature f fusion , the RGB image feature f r and the depth image feature f d through three feed - forward neural networks respectively, and aggregates the results to form the visual representation f img , the calculation formula is as follows:

[0076] f img = FFN fusion (f fusion ) + FFN rgb (f r ) + FFN depth (f d )

[0077] The structure diagram of the continuous visual - language navigation model is as shown in Figure 2 . The entire continuous visual - language navigation model is divided into three modules: constructing a topological map, cross - modal planning, and path control; the topological mapping module constructs a time - expanding topological map of the navigation process. At each step of the navigation, the agent processes the new observation information of the current node, uses the locus predictor [5] to obtain navigable nodes, and updates the topological map; the cross - modal planning module sends the encoded topological map and the instruction text into the cross - modal encoder, uses a feed - forward network to predict the long - term goal, and formulates a path plan according to the distance information stored in the topological map; the path control module adopts a low - level action control agent to execute the plan, and the low - level actions include turning left / right by 15°, moving forward by 0.25 meters, and stopping.

[0078] Specifically, this embodiment conducts experiments on the Habitat simulator

[12] using the VLN-CE dataset [2]. During the experiment, this embodiment mainly focuses on the following five performance evaluation metrics:

[0079] (1) Trajectory Length (TL): The average total length of the agent's navigation trajectory;

[0080] (2) Navigation Error (NE): The average distance between the agent's stopping position and the target position;

[0081] (3) Overall Success Rate (OS): The proportion of navigation episodes in which the point closest to the target position in the agent's predicted trajectory is within 3 meters of the target position among the total number of navigation episodes;

[0082] (4) Success Rate (SR): The proportion of navigation episodes in which the agent's stopping position is within 3 meters of the target position among the total number of navigation episodes;

[0083] (5) Success Rate with Path Length (SPL): Combining SR and TL, the success rate is weighted by the path to measure the accuracy and efficiency of navigation.

[0084] This embodiment uses the pre-trained model of previous work [7] and continues to train the model using the method proposed in the present invention. The model uses the AdamW optimizer with a learning rate of 1e-5 and is trained for 25,000 iterations on 1 NVIDIA RTX 3090 GPU, with 8 environments running simultaneously on each GPU.

[0085] This application is compared with the performance of existing models, and the experimental results are shown in Table 1:

[0086]

[0087] Table 1 Comparison of Results with the Prior Art

[0088] As can be seen in Table 1, the agent continuous visual language navigation method based on knowledge and history-aware visual representation is superior to other methods in the NE, SR, and SPL evaluation metrics of the unseen validation set, demonstrating the strong generalization ability of the agent continuous visual language navigation method based on knowledge and history-aware visual representation in unseen environments. In particular, on the unseen validation set, the present invention improves the SR and SPL metrics by 3.2% and 2.4% respectively compared to the strong baseline model ETP, and reduces the NE metric by 4.0%. The present invention completes the task with higher accuracy and reliability, demonstrating the effectiveness of knowledge and history-aware visual representation.

[0089] To verify the importance of key components in the agent's continuous vision-language navigation method based on knowledge and history-aware visual representations, this embodiment conducted ablation experiments on the unseen validation set of the VLN-CE dataset. The experimental results are shown in Table 2. First, the results in the second and third rows are better than those in the first row in all metrics except SPL, indicating that adding knowledge information or historical information to the visual representation is beneficial for navigation. Second, the comprehensive results in the fourth row are better than those in the second and third rows, especially in the SPL metric, demonstrating that integrating knowledge and historical information into the visual representation simultaneously performs better than adding only one of the elements, thus verifying the effectiveness of the knowledge-history co-aware visual representation.

[0090]

[0091] Table 2 Ablation Experiments on Key Components

[0092] Furthermore, this embodiment verified the effectiveness of the key modules for processing visual representations. The experimental results are shown in Table 3. Using all modules produced the best results on the unseen validation set. The success rates in the second to fourth rows are all better than those of the baseline model, but the navigation performance is worse than that in the fifth row, indicating that individual modules have limited improvement on navigation performance and need to cooperate with other modules. Using the filtering module and the interaction module simultaneously yields the best performance, especially in the SPL metric, indicating that the complete model has greater advantages in the efficiency and accuracy of path planning.

[0093]

[0094] Table 3 Ablation Experiments on Key Modules

[0095] It should be noted that the above specific implementation manners are only the preferred embodiments of the present invention and the applied technical principles. Those skilled in the art should understand that various modifications, equivalent replacements, and changes can be made to the present invention. However, as long as these transformations do not deviate from the spirit of the present invention, they should be within the protection scope of the present invention. In addition, some terms used in the specification and claims of this application are not restrictive and are only for the convenience of description.

[0096] References

[0097] [1] Anderson P, Wu Q, Teney D, et al. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2018:3674-3683.

[0098] [2] Krantz J, Wijmans E, Majumdar A, et al. Beyond the nav-graph: Vision-and-language navigation in continuous environments[C] / / Computer Vision–ECCV2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVIII 16. Springer International Publishing, 2020:104-120.

[0099] [3] Krantz J, Gokaslan A, Batra D, et al. Waypoint models for instruction-guided navigation in continuous environments[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021:15162-15171.

[0100] [4] Raychaudhuri S, Wani S, Patel S, et al. Language-aligned waypoint (law) supervision for vision-and-language navigation in continuous environments[J]. arXiv preprint arXiv:2109.15207, 2021.

[0101] [5]Hong Y,Wang Z,Wu Q,et al.Bridging the gap between learning indiscrete and continuous environments for vision-and-language navigation[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition.2022:15439-15449.

[0102] [6]Krantz J,Lee S.Sim-2-Sim Transfer for Vision-and-LanguageNavigation in Continuous Environments[C] / / European Conference on ComputerVision.Cham:Springer Nature Switzerland,2022:588-603.

[0103] [7]An D,Wang H,Wang W,et al.Etpnav:Evolving topological planning forvision-language navigation in continuous environments[J].IEEE Transactions onPattern Analysis and Machine Intelligence,2024.

[0104] [8]An D,Qi Y,Li Y,et al.Bevbert:Multimodal map pre-training forlanguage-guided navigation[C] / / Proceedings of the IEEE / CVF InternationalConference on Computer Vision.2023:2737-2748.

[0105] [9] Georgakis G, Schmeckpeper K, Wanchoo K, et al. Cross-modal map learning for vision and language navigation[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2022:15460-15470.

[0106]

[10] Chen P, Ji D, Lin K, et al. Weakly-supervised multi-granularity map learning for vision-and-language navigation[J]. Advances in Neural Information Processing Systems, 2022, 35:38149-38161.

[0107]

[11] Krishna R, Zhu Y, Groth O, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations[J]. International journal of computer vision, 2017, 123:32-73.

[0108]

[12] Savva M, Kadian A, Maksymets O, et al. Habitat: A platform for embodied ai research[C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2019:9339-9347.

[0109]

[13] Chen K, Chen JK, Chuang J, et al. Topological planning with transformers for vision-and-language navigation[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021:11276-11286.

[0110]

[14] Wang Z, Li X, Yang J, et al. Gridmm: Grid memory map for vision-and-language navigation[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2023:15625-15636.

Claims

1. A continuous vision-language navigation model based on knowledge and history-aware visual representations, characterized in that The continuous vision-language navigation model includes a topological map construction module, a cross-modal planning module, and a path control module. The topological map construction module includes an extraction module, a filtering module, an interaction module, and an aggregation module; specifically: The extraction module is used to extract RGB image feature f r , depth image feature f d , first knowledge feature f k , first historical feature f h and navigation instruction f i ; The filtering module respectively calculates the correlation matrix between the first knowledge feature f k , the first historical feature f h and the navigation instruction f i to obtain the weighted second knowledge feature and the second historical feature The interaction module obtains the multi-sensory fusion feature f by interacting the second knowledge feature with the second historical feature and the instruction fusion ; The aggregation module aggregates the multi-sensory fusion feature f fusion , the RGB image feature f r and the depth image feature f d to obtain the visual representation f img .

2. The continuous vision-language navigation model based on knowledge and history-aware visual representation according to claim 1, wherein The extraction module is used to extract the RGB image feature f r , the depth image feature f d , the first knowledge feature f k , the first historical feature f h and the navigation instruction f i The process includes: At each step t, the current observed RGB image and depth image are obtained. The pre-trained ViT-B / 32 model is used to encode the RGB image to obtain the RGB feature f r , and the pre-trained ResNet-50 model is used to encode the depth image to obtain the depth feature f d ; Retrieve in the knowledge base KB and select m pieces of knowledge with the highest cosine similarity to the RGB image o rgb and connect them. The knowledge feature f is calculated according to the following formula k : Specifically: kn represents knowledge encoding, and cos represents cosine similarity calculation; Through the topological graph Map t The node view feature of the previous step, the node access time step encoding of the previous step, and the relative position encoding of the node where the previous step is located and all nodes in the graph are calculated according to the following formula to obtain the historical feature f h ; Specifically: img, step, and pos respectively represent the view, access time step, and relative position, and t-1 represents the previous step of the current navigation step; The topological graph Map t contains three types of nodes: the current node, the visited node, and the navigable node; the nodes in the topological graph record the observation information of each step of the agent, and the edges record the distance information between the nodes, which is updated during the navigation process.

3. A continuous vision-language navigation model based on knowledge and history-aware visual representation according to claim 1, characterized in that, The filtering module calculates the correlation matrices between the first knowledge feature f k , the first historical feature f h , and the navigation instruction f i respectively, and obtains the weighted second knowledge feature and the second historical feature through the softmax layer. The process includes: Calculate the first knowledge feature f separately k and the first historical feature f h respectively, and the correlation matrix between the navigation instruction f i is as follows. The calculation formula is as follows: M k = f k W k (f i W i ) T M h = f h W h (f i W i ) T where: k, h, i respectively represent knowledge, history, instruction, W k , W h , W i are learnable parameters; Based on the correlation matrix, the normalized exponential layer assigns correlation scores to the knowledge features and historical features, and this score will be used as a weight reflecting the association strength between the first knowledge feature and the first historical feature and the navigation instruction, whereby the filtered weighted second knowledge feature can be generated and the second historical feature The calculation formula is as follows: Wherein, softmax represents the normalization exponential layer, and d represents the dimension of the feature.

4. A continuous vision-language navigation model based on knowledge and history-aware visual representation according to claim 1, characterized in that The interaction module obtains the multi-sensory fusion feature f through interacting the second knowledge feature with the second historical feature and the instruction. The process includes: fusion ​ The second knowledge feature is processed by a multi-layer cross-modal encoder The second historical feature Calculate the knowledge-historical feature f kh The knowledge-historical feature f kh interacts with the instruction according to the following formula to obtain the multi-sensory fusion feature f fusion : Where: W kh , W cls are learnable parameters, and I0 represents the encoded head token, representing the overall semantic representation of the instruction text.

5. A continuous vision-language navigation model based on knowledge and history-aware visual representation according to claim 4, wherein Each layer of the multi-layer cross-modal encoder includes a cross-attention sub-layer, a self-attention sub-layer, and two feed-forward networks.

6. The intelligent agent continuous visual language navigation method based on knowledge and history-aware visual representation according to claim 1, characterized in that, The aggregation module aggregates the multi-sensory fusion feature f fusion , the RGB image feature f r and the depth image feature f d through three feed-forward neural networks respectively, and aggregates the results to form the visual representation f img . The calculation formula is as follows: f img = FFN fusion (f fusion ) + FFN rgb (f r ) + FFN depth (f d ) Specifically: FFN represents a feed-forward neural network.

7. A continuous vision-language navigation method based on knowledge and history-aware visual representation, characterized in that, The method is implemented based on the continuous vision-language navigation model according to any one of claims 1-6.

Citation Information

Cited By

  • Visual language navigation method and device fusing semantic enhancement and hierarchical decision

    CN122360519A