An Indoor Visual Navigation Method Based on Causal Attention
Correcting the error correlation of the self-attention mechanism through the causal attention mechanism has improved the accuracy of navigation prediction of indoor visual navigation models in unknown environments, and solved the problem of insufficient generalization ability in traditional methods.
Patent Information
- Application Number
- CN202211273306.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-18
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-10-18
AI Technical Summary
The existing Transformer-based indoor visual navigation method has excessive attention to false correlations between features, resulting in insufficient generalization ability of the model in unknown environments.
The indoor visual navigation method based on causal attention is adopted, and by constructing a clustering center, visual and positional characteristics are extracted, combined with self-attention and causal attention mechanism, the error correlation captured by self-attention mechanism is corrected, and the prediction accuracy of the model is improved in unknown environments.
It improves the accuracy of the navigation prediction of the model in unknown environments, reduces the dependence on false correlations, and enhances the generalization ability of the model.
Smart Images

Figure CN115512214B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to visual navigation technology, and specifically to an indoor visual navigation method based on causal attention. Background Art
[0002] Indoor visual navigation is a navigation task involving indoor visual environments, aiming to predict and execute navigation actions based on visual images observed from the environment to reach a specified destination and complete the navigation goal. Existing methods for solving indoor visual navigation generally include two steps. One is the understanding of the visual environment state, and the other is the prediction of navigation actions.
[0003] Visual state understanding methods focus on understanding the information of observed visual images and analyzing the environmental state, and extract the environmental visual state features and historical state features at each moment by constructing a representation model with complex structures and mechanisms.
[0004] Navigation action prediction methods aim to predict navigation actions based on visual state features, and formulate the best navigation action sequence by constructing an effective path planning strategy, environmental exploration mode and reward feedback mechanism to reach a specified destination and complete the navigation task.
[0005] Due to the complex high-dimensional state space in the indoor visual navigation environment and the development of technologies such as representation learning and large-scale pre-trained models, most existing work focuses on visual environment state understanding methods. In existing indoor visual navigation methods based on Transformer, visual environment state understanding methods have significantly improved the prediction performance of navigation models by constructing a representation model with strong feature representation ability and obtaining prior knowledge from large-scale image pre-trained models. However, affected by hidden environmental factors, such methods have the problem of over-focusing on false correlations, and the generalization prediction effect in unknown environments is relatively average. Summary of the Invention
[0006] The technical problem to be solved by the present invention is: to propose an indoor visual navigation method based on causal attention to solve the problem that traditional indoor visual navigation schemes over-focus on false correlations between features and reduce the generalization ability of the model.
[0007] The technical solution adopted by the present invention to solve the above technical problems is:
[0008] An indoor visual navigation method based on causal attention, comprising the following steps:
[0009] A. Data Preparation
[0010] Obtain an indoor visual image dataset, where the indoor visual image dataset includes a set of navigation trajectory data. Each navigation trajectory data respectively includes a navigation trajectory composed of a sequence of positions and a sequence of visual images at each position on the navigation trajectory. Each sequence of visual images respectively includes images in each observation direction at the corresponding position;
[0011] And based on the navigation trajectory data, construct a navigation image sequence composed of the images corresponding to the navigation directions at each position on the navigation trajectory before reaching the end point. The navigation direction corresponding image is an image determined from the sequence of visual images at the corresponding position according to the direction from the corresponding position to the next position on the navigation trajectory; Then, perform visual feature extraction and clustering on the navigation image sequences of all navigation trajectory data to obtain the clustering centers;
[0012] B. Execute the indoor visual navigation task through the indoor visual navigation model:
[0013] B1. Use the navigation start position as the initial current position and randomly initialize the historical state features;
[0014] B2. Observe each observation direction at the current position to obtain the sequence of visual images at the current position, extract the visual features of each image in the sequence of visual images at the current position, and encode to obtain the position features of each observation direction. And according to the distances between the visual features of each image and each clustering center, obtain the global features of each image;
[0015] B3. Incorporate the historical state features into the visual features of each image in the sequence of visual images at the current position respectively to obtain the visual image features of each image;
[0016] Fuse the visual image features and their position features of each image, and through the self-attention mechanism, calculate the self-attention features of each image in the sequence of visual images at the current position;
[0017] Fuse the visual image features and position features of each image to construct a query vector; According to the global features of each image, construct key vectors and value vectors. Then, based on the constructed query vector, key vector and value vector, through the causal attention mechanism, calculate the causal attention features of each image in the sequence of visual images at the current position;
[0018] Then, fuse the self-attention features and their causal attention features of each image to obtain the visual environment state features of each image in the sequence of visual images at the current position;
[0019] B4. According to the preset navigable directions, calculate the correlation between the visual features of the images in the navigable directions in the sequence of visual images at the current position and their corresponding visual environment state features, and predict the navigation action at the current position according to the correlation;
[0020] B5. Determine the next position of navigation based on the navigation action at the current position, and determine whether the end point is reached or the preset maximum number of navigation steps is reached. If so, end the navigation; otherwise, execute step B6.
[0021] B6. Update the historical state features based on the visual environment state features of the current position obtained in step B3 and the navigation action predicted in step B4; use the next position determined by the navigation action at the current position and the updated historical state features as inputs, and return to step B2.
[0022] Further, train the indoor visual navigation model according to the following steps:
[0023] C1. Use the indoor visual image dataset as the training dataset and calculate to obtain the clustering centers.
[0024] C2. Extract a navigation trajectory data from the training dataset, and use all or part of it as the navigation trajectory data for this round of training.
[0025] C3. From the input navigation trajectory data, extract the visual image sequence of its starting point as the initial input visual image sequence, and randomly initialize the historical state features.
[0026] C4. Use the position corresponding to the input visual image sequence as the current position, extract the visual features of each image in the visual image sequence of the current position, encode to obtain the position features of each observation direction, and obtain the global features of each image according to the distances between the visual features of each image and each clustering center.
[0027] C5. Incorporate the historical state features into the visual features of each image in the visual image sequence of the current position respectively to obtain the visual image features of each image; then, calculate the self-attention feature and the causal attention feature of the current position, and fuse the self-attention feature and its causal attention feature to obtain the visual environment state features.
[0028] C6. According to the preset navigable directions, calculate the correlation between the visual features of the images in the navigable directions in the visual image sequence of the current position and their corresponding visual environment state features, and predict the navigation action of the current position based on the correlation.
[0029] C7. Determine whether the end point of the input navigation trajectory data is reached. If so, execute step C9; otherwise, execute step C8.
[0030] C8. Update the historical state features based on the visual environment state features of the current position obtained in step C5 and the navigation action predicted in step C6; extract the visual image sequence of the next position of the navigation trajectory from the navigation trajectory data, and use the visual image sequence and the updated historical state features as inputs, and return to step C4.
[0031] C9. Calculate the loss based on the preset expert navigation actions and the predicted navigation actions at each position, and update the parameters of the indoor visual navigation model according to the cumulative loss;
[0032] C10. Repeat steps C2 - C9 for iterative training until the training termination condition is met, and obtain the trained indoor visual navigation model.
[0033] Furthermore, in step B, initially, use the clustering centers obtained during training, and use the navigation trajectory data of the indoor visual image dataset during training as the initial historical navigation trajectory data; after performing the indoor visual navigation task, collect the navigation trajectory data of the actually completed navigation tasks. After the collection reaches the set number, update the historical navigation trajectory data according to the collected navigation trajectory data, and update the clustering centers based on the updated historical navigation trajectory data.
[0034] Furthermore, in step C9, the cumulative loss is calculated according to the following loss function:
[0035] L = w1L il + w2L rl
[0036] where w1 and w2 are both trainable parameters, L il represents the loss generated by imitation learning, and L rl represents the loss generated by reinforcement learning. The reinforcement learning adopts an actor - critic framework, where the actor network is the indoor visual navigation model and the critic network is a feed - forward neural network;
[0037] where L il and L rl are calculated according to the following formulas respectively:
[0038]
[0039]
[0040] where a t represents the predicted navigation action at the position at time t, represents the preset expert navigation action at the position at time t, π t represents the correlation between the visual feature of the visual image sequence and the corresponding visual environment state feature at the position at time t, G t represents the cumulative reward of the actor network at the position at time t, and TD t is the output of the critic network at the position at time t and is calculated according to the following formula:
[0041] TD t = max(0, πt W TD1 )W TD2
[0042] Among them, W TD1 and W TD2 are trainable parameters.
[0043] Furthermore, according to the following formula, calculate the cumulative reward G of the executor network t :
[0044]
[0045]
[0046] where p cur represents the position at the next moment corresponding to the predicted navigation action at the position at time t, p goal represents the position at the next moment corresponding to the expert navigation action at time t, dis(·) represents the Euclidean distance, and γ t represents the decay factor at time t.
[0047] Specifically, the calculation of the cluster center includes:
[0048] D1. Extract the visual features of each image in the navigation image sequence of each navigation trajectory data, and form a global visual feature dataset with all the extracted visual features;
[0049] D2. Set K cluster centers and initialize them;
[0050] D3. According to the global visual feature dataset, calculate the Euclidean distance between each visual feature and each cluster center respectively;
[0051] D4. Classify each visual feature based on the minimum distance between each visual feature and each cluster center;
[0052] D5. Update the value of the cluster center according to the following formula:
[0053]
[0054] where g k represents the value of the k-th cluster center, and C k represents the set of visual features included in the k-th cluster center;
[0055] D6. Repeat the above steps D3 to D5 to iteratively update the value of the cluster center until the change in the value of all cluster centers is less than the preset threshold or exceeds the preset number of iteration rounds.
[0056] Specifically, according to the distances between the visual features of each image in the visual image sequence at the current position and each cluster center, the global feature is obtained. Among them, represents the global feature of the image in the i-th observation direction, N is the number of observation directions, and it is calculated according to the following steps:
[0057] Calculate the distances between the visual features of the image in the i-th observation direction and the K cluster centers respectively, and take the mean of the distances from it to the K cluster centers as its global feature.
[0058] Specifically, the historical state features are respectively incorporated into the visual features of each image in the visual image sequence at the current position to obtain the visual image features of each image, including:
[0059] First, the visual feature F t ={f1, f2, … f i , …, f N} is respectively subjected to global average pooling;
[0060] Then, in the form of vector splicing, the historical state feature H t-1 is respectively incorporated into each visual feature after global average pooling to obtain the visual image feature C t ={c1, c2, … c i , …, c N} of each image, where t represents the current position and t - 1 represents the previous position of the current position.
[0061] Specifically, the position feature is subjected to absolute position encoding using a pre-trained BERT model.
[0062] Specifically, in each step, a residual neural network is used to extract the visual features of the image.
[0063] Specifically, the visual image features of each image and their position features are fused, and through a self-attention mechanism, the self-attention features of each image in the visual image sequence at the current position are calculated, including:
[0064] First, the visual image features and their position features are fused by splicing, and then, through a multi-layer perceptron network with different parameters, the fused features are converted into a query vector Q s , a key vector K s and a value vector V s :
[0065] Q s = max(0, (C t + PE t ) W qs + b qs )
[0066] K s = max(0, (C t + PE t )W ks + b ks )
[0067] V s = max(0, (C t + PE t )W vs + b vs )
[0068] Wherein, C t represents the visual image feature at the current position, PE t represents the position feature at the current position, W qs , b qs , W ks , b ks , W vs and b vs are all parameters of the multi-layer perceptron network;
[0069] Then, calculate the attention weight a s :
[0070]
[0071] Wherein, dim is the dimension of the multi-layer perceptron network, and T represents matrix transpose;
[0072] Finally, through the attention weight and the value vector, calculate and obtain the self-attention feature:
[0073] SA t = softmax(a s V s )
[0074] Wherein, SA t represents the self-attention feature at the current position.
[0075] Specifically, fuse the visual image features and position features of each image to construct a query vector; according to the global features of each image, construct a key vector and a value vector, and then, based on the constructed query vector, key vector and value vector, through the causal attention mechanism, calculate the causal attention features of each image in the visual image sequence at the current position, including:
[0076] First, fuse the visual image feature and its position feature by splicing, and then, through the multi-layer perceptron network, convert the fused feature into a query vector Q c :
[0077] Q c = max(0, (Ct +PE t )W qc +b qc )
[0078] And through a multi - layer perceptron network with different parameters, convert the global features of the visual image sequence corresponding to the current position into the key vector K c and the value vector V c :
[0079]
[0080]
[0081] Among them, C t represents the visual image features of the current position, PE t represents the position features of the current position, represents the global features, W qc , b qc , W kc , b kc , W vc and b vc are all parameters of the multi - layer perceptron network;
[0082] Then, calculate the attention weight a c :
[0083]
[0084] Among them, dim is the dimension of the multi - layer perceptron network, and T represents matrix transpose;
[0085] Finally, through the attention weight and the value vector, calculate and obtain the causal attention feature:
[0086] CA t = softmax(a c V c )
[0087] Among them, CA t represents the causal attention feature of the current position.
[0088] Specifically, fuse the self - attention features and causal attention features of each image to obtain the visual environment state features of each image in the visual image sequence at the current position, including:
[0089] First, through the way of vector concatenation, fuse the self - attention feature SA t and the causal attention feature CA t , to obtain the fusion feature [SA t , CA t ;
[0090] Then, a feedforward neural network is adopted to convert the fused features [SA t , CA t into the visual environment state feature S t :
[0091] S t = max(0, [SA t , CA t W ffn1 + b ffn1 )W ffn2 + b ffn2
[0092] where are all parameters of the feedforward neural network, dim is the dimension of the encoding network for constructing query vectors, key vectors, and value vectors in the attention calculation, and N is the number of observation directions.
[0093] Furthermore, the navigation trajectory data further includes navigable direction labels at each position of the navigation trajectory. In step C6, only the directions with navigable direction labels are regarded as navigable directions; in step B4, all observation directions are regarded as navigable directions.
[0094] Specifically, according to the preset navigable directions, the correlation between the visual features of the images of the navigable directions in the current position visual image sequence and their corresponding visual environment state features is calculated, and the navigation action at the current position is predicted, including:
[0095] First, calculate the correlation π t between the visual features of the images of each navigable direction in the current position visual image sequence and their corresponding visual environment state features:
[0096]
[0097] where represents the visual feature of the image of the navigable direction in the current position visual image sequence, and S t represents the visual environment state feature of the image of the navigable direction in the current position visual image sequence;
[0098] Then, according to the correlation π t predict the navigation action a t at the current position:
[0099] a t = argmax m π t,m
[0100] where π t,m represents v tThe correlation of the m-th direction in the sequence.
[0101] Specifically, according to the visual environment state features at the current position and the predicted navigation actions at the current position, update the historical state features, including:
[0102] First, filter the key features of the visual environment state features S t at the current position and the predicted navigation actions a r at the current position through the reset gate, and fuse them into the historical state features H t-1 at the previous moment of the current position:
[0103] r t = σ(W r H t-1 + U r [S t , π t , a t )
[0104]
[0105] where π t represents the correlation between the visual features of each navigable direction image in the visual image sequence at the current position and their corresponding visual environment state features, r t represents the forgetting gate weight, and W r , U r , W g and U g are all trainable parameters, σ(·) and tanh(·) represent activation functions, ⊙ represents the Hadamard product operation, t represents the current position, and t - 1 represents the previous position of the current position;
[0106] Then, filter the valid historical information z t to be retained through the update gate, and fuse it into the historical state features H t-1 at the previous moment of the current position to update the historical state features:
[0107] z t = σ(W z H t-1 + U z [S t , π t , a t )
[0108]
[0109] where z t represents the update gate weight, and W z and U z are all trainable parameters.
[0110] The beneficial effects of the present invention are as follows:
[0111] Existing indoor visual language navigation methods based on Transformer capture the correlation between visual features through the self-attention mechanism to predict navigation actions. However, the correlation calculation of the self-attention mechanism is limited by the co-occurrence frequency of features in the training dataset, and it is prone to capturing false correlations, resulting in the model trained only performing well in the training dataset and poorly in other datasets.
[0112] The causal attention mechanism proposed by the present invention corrects the wrong correlation by means of intervention, that is, mapping the current feature to other features to determine whether the correlation still exists in other scenarios, so as to improve the generalization ability of the model in unknown environments. Specifically, the present invention constructs a clustering center according to historical navigation trajectory data, obtains the global features of each observation direction at the current position according to the clustering center, and then, through the causal attention mechanism, corrects the wrong correlation captured by the self-attention mechanism to improve the prediction accuracy of the model in unknown test environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0113] Figure 1 It is a flowchart of the training of the indoor visual navigation model in an embodiment of the present invention;
[0114] Figure 2 It is a diagram of the process of extracting visual environment state features in an embodiment of the present invention;
[0115] Figure 3 It is a diagram of the process of predicting navigation actions in an embodiment of the present invention;
[0116] Figure 4 It is a diagram of the process of updating historical state features in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0117] The present invention aims to propose an indoor visual navigation method based on causal attention to solve the problem that traditional indoor visual navigation schemes over-focus on false correlations between features and reduce the generalization ability of the model. The indoor visual navigation method based on causal attention includes two major parts: training of the indoor visual navigation model and using the model to execute navigation tasks, but the processes are similar. The following part will mainly describe the training of the indoor visual navigation model.
[0118] During the training process of the indoor visual navigation model, first, according to the visual image dataset, cluster the images corresponding to the navigation directions at each position of each navigation trajectory and calculate the cluster centers. Next, extract the visual features and position features of the image at the current moment, and calculate the global features based on the distance between the visual features and the cluster centers. Then, fuse the visual features with the historical state features and position features, calculate the self-attention features through the self-attention mechanism, and construct a query vector based on the visual features fused with the historical state features and position features. Use the global image features to construct the key vector and value vector, and calculate the causal attention features through the causal attention mechanism. Then, fuse the self-attention features and causal attention features to obtain the visual environment state features of each image. Then, predict the navigation action at the current position by calculating the correlation between the visual features at the current position and the corresponding visual environment state features. Finally, update the historical state features according to the predicted navigation action at the current position and the visual environment state features at the current position. Use the updated historical state features and the image at the next navigation position as the input, and perform iterative training using the cumulative loss after completing the navigation instance task to obtain a trained indoor visual navigation model.
[0119] The solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0120] For the sake of easy understanding, first, the technical terms that may be involved in this embodiment are described:
[0121] Residual Network (Resnet): A convolutional neural network model for image recognition, mainly composed of several stacked residual layers. Currently, in various computer vision tasks, it is often used to extract the visual features of input images.
[0122] Attention mechanism: A mechanism for selectively processing features, mainly composed of a query vector, a key vector, a value vector, and an attention operation. Currently, it has become an indispensable basic component in most deep learning models.
[0123] Transformer: An encoder-decoder model based on the self-attention mechanism, initially applied to sequence conversion tasks such as machine translation and sequence modeling, and has become the main deep learning model in the field of natural language processing. Due to its powerful performance, Transformer has gradually been widely applied in the field of computer vision to extract the visual features of images.
[0124] Actor-Critic (AC): A most commonly used method for solving the optimal strategy in reinforcement learning. It combines both policy gradient and value estimation methods for solving strategies, mainly composed of a policy network and a value evaluation network.
[0125] Front Door Adjustment (FDA): It is a method to implement intervention in causal inference. By blocking the front door path, the intervention distribution is estimated, and even when it is impossible to effectively observe hidden confounders, the causal relationship between feature variables can still be analyzed.
[0126] Embodiment:
[0127] Among them, the model training process, as Figure 1 shown, is specifically described as follows:
[0128] S1. Data preprocessing of the training dataset
[0129] The training uses an indoor visual image dataset as the training dataset. The indoor visual image dataset includes a set of navigation trajectory data. Each navigation trajectory data includes a navigation trajectory composed of a position sequence and a visual image sequence at each position on the navigation trajectory. Each visual image sequence includes images in each observation direction at the corresponding position.
[0130] Assume that the current position is t. Then, the visual image sequence at position t can be expressed as V t ={v1, v2, … v i , …, v N}, where N represents the number of observation directions, and v i represents the image obtained by observing in the i-th observation direction at position t. The image format is RGB image, which can be expressed as H and W respectively represent the height and width of the image.
[0131] Then, based on the navigation trajectory data, a navigation image sequence composed of the images corresponding to the navigation directions at each position on the navigation trajectory before reaching the end point is constructed. The image corresponding to the navigation direction is the image determined from the visual image sequence at the corresponding position according to the direction from the corresponding position to the next position on the navigation trajectory; then, visual feature extraction and clustering are performed on the navigation image sequences of all navigation trajectory data to obtain the clustering centers.
[0132] The calculation of the clustering centers includes:
[0133] a1. Through the Resnet-164 residual neural network, extract the visual features of each image in the navigation image sequences of each navigation trajectory data, and form a global visual feature dataset with all the extracted visual features. In addition to the residual neural network, other existing methods can also be used for visual feature extraction, such as Transformer.
[0134] a2. Set K clustering centers, and randomly sample K visual features from the global visual feature dataset as the initial values of the K clustering centers. The initialization of the clustering centers can also be done in other ways, such as random assignment or manual assignment.
[0135] a3. Calculate the Euclidean distances between each visual feature in the global visual feature dataset and each clustering center respectively.
[0136] a4. Classify each visual feature based on the minimum distance between each visual feature and each clustering center.
[0137] a5. Update the values of the clustering centers according to the following formula:
[0138]
[0139] where, g k represents the value of the k-th clustering center, and C k represents the set of visual features included in the k-th clustering center;
[0140] a6. Repeat the above steps a3 - a5 to iteratively update the values of the clustering centers until the changes in all clustering center values are less than a preset threshold or exceed a preset number of iteration rounds. If the preset number of iteration rounds is exceeded, it indicates that the calculation has failed and the clustering calculation should be performed again.
[0141] S2. Extract navigation trajectory data and train an indoor visual navigation model
[0142] S21. Extract a navigation trajectory data from the training dataset as the input for training. If the extracted navigation trajectory data has a large number of navigation steps, it can also be segmented and input, that is, during the training process, only a part of it is extracted as the input.
[0143] S22. Initialization: Extract the visual image sequence at the starting point from the input navigation trajectory data as the initial input visual image sequence, and randomly initialize the historical state features.
[0144] S23. Take the position corresponding to the input visual image sequence as the current position, extract the visual features of each image in the visual image sequence at the current position, encode to obtain the position features in each observation direction, and obtain the global features of each image according to the distances between the visual features of each image and each clustering center.
[0145] For this embodiment, the extraction of various features is specifically described as follows:
[0146] I. Visual feature extraction
[0147] For the visual image sequence V t ={v1, v2,... vi , …, v N}, use the Resnet-164 residual neural network to extract the visual feature F t = {f1, f2, … f i , …, f N}, f i represents the visual feature in the i-th direction at the t position.
[0148] II. Position Feature Encoding
[0149] Since in the subsequent process of extracting visual environment state features, the position relationship of each image cannot be recognized through visual image features, a position encoding vector is required to represent the directional position information of the image. Therefore, the present invention represents the directional position information of the image through position features, and the position features have the same dimension as the subsequent visual image features.
[0150] In the embodiment, the position features are subjected to absolute position encoding using a pre-trained BERT model, and the encoding process is as follows:
[0151] First, initialize the position feature PE t = {pe1, pe2, … pe i , …, pe N}, and its initialization can adopt any existing method. In the embodiment, PE t = {[1, 1,.., 1], [2, 2.., 2] …, [N, N,.., N]}, where pe i represents the position feature in the i-th observation direction, N is the number of observation directions, and t represents the current position;
[0152] Then, input the initialized position features into the pre-trained BERT model to obtain the absolute position encoding of each position through learning.
[0153] The above pre-trained BERT model comes from the Google paper "Pre-training of Deep Bidirectional Transformers for Language Understanding", which uses the Encoder module of Transformer. BERT is the abbreviation of "Bidirectional Encoder Representations from Transformers".
[0154] III. Global Feature Extraction
[0155] Obtain the global feature according to the distance between the visual features of each image in the current position visual image sequence and each cluster center Among them, Represents the global feature of the image in the i-th observation direction. N is the number of observation directions and is calculated as follows:
[0156] Calculate the distances between the visual features of the image in the i-th observation direction and the K clustering centers respectively, and take the mean of the distances from it to the K clustering centers as its global feature
[0157] S24. Calculate the visual environment state feature of the image
[0158] The visual environment state feature, as Figure 2 shown, is obtained by fusing the self-attention feature and the causal attention feature of the image, and is used to capture the correlation and causal relationship between visual features.
[0159] In this step, first, the historical state features are respectively incorporated into the visual features of each image in the visual image sequence at the current position to obtain the visual image features of each image.
[0160] Then, fuse the visual image features and their position features of each image, and through the self-attention mechanism, calculate the self-attention features of each image in the visual image sequence at the current position. Fuse the visual image features and position features of each image to construct a query vector; according to the global features of each image, construct a key vector and a value vector, and then, based on the constructed query vector, key vector and value vector, calculate the causal attention features of each image in the visual image sequence at the current position through the causal attention mechanism.
[0161] Finally, fuse the self-attention features and the causal attention features of each image to obtain the visual environment state features of each image in the visual image sequence at the current position.
[0162] The specific description is as follows:
[0163] I. Calculate the visual image features
[0164] First, for the convenience of vector splicing, the visual feature F t ={f1, f2, … f i , …, f N} is respectively subjected to global average pooling to reduce the tensor to a vector;
[0165] Then, in the form of vector splicing, the historical state feature H t-1 is respectively incorporated into each visual feature after global average pooling to obtain the visual image feature C t ={c1, c2, … c i , …, c N} of each image, where t represents the current position and t-1 represents the previous position of the current position.
[0166] II. Calculate self-attention features
[0167] First, fuse the visual image features and their position features by concatenation, and then convert the fused features into query vector Q through a multi-layer perceptron network with different parameters s , key vector K s and value vector V s :
[0168] Q s = max(0, (C t + PE t ) W qs + b qs )
[0169] K s = max(0, (C t + PE t ) W ks + b ks )
[0170] V s = max(0, (C t + PE t ) W vs + b vs )
[0171] where C t represents the visual image features at the current position, PE t represents the position features at the current position, and W qs , b qs , W ks , b ks , W vs and b vs are all parameters of the multi-layer perceptron network;
[0172] Then, calculate the attention weight a s :
[0173]
[0174] where dim is the dimension of the multi-layer perceptron network, and T represents matrix transpose;
[0175] Finally, calculate the self-attention features through the attention weight and the value vector:
[0176] SA t = softmax(a s V s )
[0177] where SA t represents the self-attention features at the current position.
[0178] III. Calculate Causal Attention Features
[0179] First, fuse the visual image features and their position features by concatenation, and then convert the fused features into a query vector Q through a multi-layer perceptron network c :
[0180] Q c = max(0, (C t + PE t ) W qc + b qc )
[0181] And convert the global features of the visual image sequence corresponding to the current position into a key vector K and a value vector V through multi-layer perceptron networks with different parameters c and value vector V c :
[0182]
[0183]
[0184] where C t represents the visual image features of the current position, PE t represents the position features of the current position, represents the global features, W qc , b qc , W kc , b kc , W vc and b vc are all parameters of the multi-layer perceptron network;
[0185] Then, calculate the attention weight a c :
[0186]
[0187] where dim is the dimension of the multi-layer perceptron network, and T represents matrix transpose;
[0188] Finally, calculate the causal attention features through the attention weight and the value vector:
[0189] CA t = softmax(a c V c )
[0190] where CA t represents the causal attention features of the current position.
[0191] The causal attention mechanism is a front-door adjustment method based on causal reasoning. By blocking the front-door path, intervening in variable inputs, and analyzing the causal relationships between feature variables, it corrects the spurious correlations established by the self-attention mechanism in known training data. In the actual implementation process, if all the navigation trajectory data of the training data set is used for intervention in sequence, it will consume a large amount of computing resources. Therefore, the present invention uses global features for replacement. Therefore, in order to ensure the representativeness of the global features and the generalization performance of the model, when performing the indoor visual navigation task, initially, the clustering centers obtained during training are adopted, and the navigation trajectory data of the indoor visual image data set during training is used as the initial historical navigation trajectory data; after performing the indoor visual navigation task, the navigation trajectory data of the actually completed navigation tasks is collected. After the collected data reaches the set quantity, the historical navigation trajectory data is updated according to the collected navigation trajectory data, and the clustering centers are updated based on the updated historical navigation trajectory data according to steps a1 to a6.
[0192] IV. Fusing Self-Attention Features and Causal-Attention Features
[0193] First, the self-attention feature SA t and the causal-attention feature CA t are fused by means of vector concatenation to obtain the fused feature [SA t , CA t ;
[0194] Then, a feed-forward neural network is used to convert the fused feature [SA t , CA t into the visual environment state feature S t :
[0195] S t = max(0, [SA t , CA t W ffn1 + b ffn1 )W ffn2 + b ffn2
[0196] where are all parameters of the feed-forward neural network, dim is the dimension of the encoding network for constructing query vectors, key vectors, and value vectors in attention calculation, and N is the number of observation directions.
[0197] S25. Predicting the Navigation Action at the Current Location
[0198] In this step, first, according to the preset navigable directions, the correlation between the visual features of the images in the navigable directions in the current location visual image sequence and their corresponding visual environment state features is calculated, and then the navigation action at the current location is predicted according to the correlation. The process is as followsFigure 3 As shown in the figure, it specifically includes:
[0199] First, calculate the correlation π between the visual features of the images in each navigable direction of the current position visual image sequence and their corresponding visual environment state features t :
[0200]
[0201] Among them, represents the visual features of the images in the navigable directions in the current position visual image sequence, and S t represents the visual environment state features of the images in the navigable directions in the current position visual image sequence;
[0202] Then, based on the correlation π t predict the navigation action a at the current position t :
[0203] a t = argmax m π t,m
[0204] Among them, π t,m represents the correlation of the m-th direction in the π t sequence.
[0205] The above-mentioned navigable directions can be all the observation directions. However, in order to narrow the exploration space and improve the training efficiency, during the training process of the embodiment, navigable direction labels are set for annotation. That is, the navigation trajectory data also includes the navigable direction labels at each position of the navigation trajectory. In the above steps during training, only the directions with navigable direction labels are regarded as navigable directions; while in an unfamiliar environment where the indoor visual navigation task is actually executed, all the observation directions are regarded as navigable directions. Specifically, at position t, the visual image sequence obtained in each observation direction is V t = {v1, v2,... v i ,..., v N}, and the navigable direction label corresponds to a mask vector Mask t = {0, 1,... 1,..., 0}, where the assignment of 1 indicates navigable. At this time, the images in the navigable directions are O t = {v2,..., v i ,...}. While when actually executing the indoor visual navigation task, for an unfamiliar environment, the mask vector can be all set to 1, that is, Mask t = {1, 1,... 1,..., 1}.
[0206] S26. Iterative training
[0207] Since the navigation is in the form of step - by - step, only completing the action prediction of the current step does not mean that it has completed the current round of navigation instance task. Therefore, in this step, first, it is determined whether the end point of the input navigation trajectory data is reached. If so, an iterative input is constructed and the iteration is returned. Otherwise, the loss is calculated and the parameters are updated.
[0208] Among them, constructing the iterative input and returning the iteration is as follows: According to the visual environment state feature of the current position obtained in step S24 and the navigation action of the current position predicted in step S25, the historical state feature is updated; from the navigation trajectory data, the visual image sequence of the next position of the navigation trajectory is extracted, and this visual image sequence and the updated historical state feature are used as inputs, and step S23 is returned.
[0209] Furthermore, the historical state feature represents the historical information of the completed navigation process. Its update, that is, fusing the information of the current step and the historical information before the current step. Therefore, a gated network can be used to fuse the visual environment state feature and navigation action of the current position with the historical state feature of the current position.
[0210] In this embodiment, as Figure 4 shown, it specifically includes:
[0211] First, through the reset gate, the key features of the visual environment state feature S t of the current position and the predicted navigation action a t of the current position are screened, and they are fused into the historical state feature H t-1 of the previous moment of the current position:
[0212] r t =σ(W r H t-1 +U r [S t ,π t ,a t )
[0213]
[0214] Among them, π t represents the correlation between the visual features of each navigable direction image of the visual image sequence of the current position and their corresponding visual environment state features, r t represents the forgetting gate weight, W r , U r , W g and U g are all trainable parameters, σ(·) and tanh(·) represent activation functions, ⊙ represents the Hadamard product operation, t represents the current position, and t - 1 represents the previous position of the current position;
[0215] Then, through the update gate, the effective historical information z to be retained is filtered r , and it is fused into the historical state feature H at the previous moment at the current position t-1 to update the historical state feature:
[0216] z t = σ(W z H t-1 + U z [S t , π t , a t )
[0217]
[0218] where z t represents the update gate weight, and both W z and U z are trainable parameters.
[0219] The above loss calculates and updates the parameters, and then trains according to the cumulative loss of completing the current round of navigation instance tasks.
[0220] In this embodiment, the training method includes two parts: namely, reinforcement learning training and imitation learning training.
[0221] 1) Reinforcement learning training: For the indoor visual navigation method, the cumulative reward feedback from the environment is used as the supervision signal, and the navigation model parameters are trained using this signal. Through the reinforcement learning training method, the model is guided to output actions with high potential reward benefits, which can promote the model to predict the correct navigation trajectory related to the task as much as possible.
[0222] 2) Imitation learning training:
[0223] The training of the indoor visual navigation model depends on effective feedback rewards. However, the complex and large state space of the environment makes it usually difficult for the model to explore the correct positive reward trajectory, increasing the training difficulty. Therefore, the model is guided to predict expert actions through imitation learning training, and as much as possible explore the positive reward trajectory similar to the expert data to quickly learn the navigation prior knowledge.
[0224] Specifically, the cumulative loss is calculated according to the following loss function:
[0225] L = w1L il + w2L rl
[0226] where w1 and w2 are both trainable parameters, L il represents the loss generated by imitation learning, L rlDenote the loss generated by reinforcement learning, where the reinforcement learning adopts an actor-critic framework, with the actor network being an indoor visual navigation model and the critic network being a feedforward neural network;
[0227] Among them, L il and L rl are calculated according to the following formulas respectively:
[0228]
[0229]
[0230] Among them, a t represents the predicted navigation action at the position at time t, represents the preset expert navigation action at the position at time t, π t represents the correlation between the visual feature of the visual image sequence and the corresponding visual environment state feature at the position at time t, G t represents the cumulative reward of the actor network at the position at time t, TD t is the output of the critic network at the position at time t and is calculated according to the following formula:
[0231] TD t =max(0,π t W TD1 )W TD2
[0232] In the reinforcement learning method, the navigation model will obtain a feedback reward r t from the environment at each moment, and the feedback reward is used to measure the navigation effect and can be defined according to the actual environment. In the embodiment, the cumulative reward G t of the above-mentioned actor network is calculated according to the following formula:
[0233]
[0234]
[0235] Among them, p cur represents the position at the next moment corresponding to the predicted navigation action at the position at time t, p goal represents the position at the next moment corresponding to the expert navigation action at time t, dis(·) represents the Euclidean distance, and γ t represents the attenuation factor at time t.
[0236] After the loss function calculation is completed, the model parameters are updated through backpropagation according to the training loss. The model parameters include the parameters of the residual neural network for extracting visual features, the pre-trained BERT model parameters for position encoding, the parameters of each multi-layer perceptron network in the two attention calculations, the parameters of the gated network for updating the historical state features, the feed-forward neural network parameters for calculating the visual environment state features, and the feed-forward neural network parameters of the critic network.
[0237] After the parameter update is completed, it is determined whether the training is completed. If so, the training ends and the trained indoor visual navigation model is obtained; otherwise, steps S21 - S26 are repeated for iterative training until the training termination condition is met. The training termination condition includes model convergence or reaching the set maximum number of training iterations. If the set maximum number of training iterations is reached and the model still does not converge, training should be restarted.
[0238] After obtaining the trained model, the model can be used for practical applications. Specifically, it includes:
[0239] Step 1: Observe the visual image sequence of the current position in each observation direction, and predict the navigation action of the current position according to the trained indoor visual navigation model and the clustering center.
[0240] Step 2: Determine the next position of the navigation according to the navigation action of the current position, and determine whether the end point is reached or the preset maximum number of navigation steps is reached. If so, the navigation ends; otherwise, the next position determined by the navigation action of the current position is used as the input and step 1 is returned.
[0241] For the case where the preset maximum number of navigation steps is reached, the position at the end of the navigation can be used as a new starting point for navigation; or, after retraining the model, navigation is performed again.
[0242] Although the present invention has been described herein with reference to the embodiments of the present invention, the above embodiments are only the preferred embodiments of the present invention, and the embodiments of the present invention are not limited by the above embodiments. It should be understood that those skilled in the art can design many other modifications and embodiments, and these modifications and embodiments will fall within the scope of the principles and spirit disclosed in this application.
Claims
1. An indoor visual navigation method based on causal attention, characterized in that, Including the following steps: A. Data preparation Obtain an indoor visual image dataset, where the indoor visual image dataset includes a set of navigation trajectory data. Each navigation trajectory data respectively includes a navigation trajectory composed of a position sequence and a visual image sequence at each position on the navigation trajectory. Each visual image sequence respectively includes images in each observation direction at the corresponding position; And based on the navigation trajectory data, construct a navigation image sequence composed of navigation direction corresponding images at each position on the navigation trajectory before reaching the end point. The navigation direction corresponding image is an image determined from the visual image sequence at the corresponding position according to the direction from the corresponding position to the next position on the navigation trajectory; then, perform visual feature extraction and clustering on the navigation image sequences of all navigation trajectory data to obtain clustering centers; B. Execute the indoor visual navigation task through the indoor visual navigation model: B1. Use the navigation start position as the initial current position and randomly initialize the historical state features; B2. Observe each observation direction at the current position to obtain the visual image sequence at the current position, extract the visual features of each image in the visual image sequence at the current position, and encode to obtain the position features of each observation direction, and obtain the global features of each image according to the distances between the visual features of each image and each clustering center; B3. Incorporate the historical state features into the visual features of each image in the visual image sequence at the current position respectively to obtain the visual image features of each image; Fuse the visual image features and position features of each image, and calculate the self-attention features of each image in the visual image sequence at the current position through the self-attention mechanism; Fuse the visual image features and position features of each image to construct a query vector; construct a key vector and a value vector according to the global features of each image, and then, based on the constructed query vector, key vector and value vector, calculate the causal attention features of each image in the visual image sequence at the current position through the causal attention mechanism; Then, fuse the self-attention features and causal attention features of each image to obtain the visual environment state features of each image in the visual image sequence at the current position; B4. Calculate the correlation between the visual features of the images in the navigable direction in the visual image sequence at the current position and their corresponding visual environment state features according to the preset navigable directions, and predict the navigation action at the current position according to the correlation; B5. Determine the next position of the navigation according to the navigation action at the current position, and determine whether the end point is reached or whether the preset maximum navigation steps are reached. If so, end the navigation, otherwise, execute step B6; B6. Update the historical state features according to the visual environment state features at the current position obtained in step B3 and the navigation action at the current position predicted in step B4; use the next position determined by the navigation action at the current position and the updated historical state features as inputs, and return to step B2.
2. The indoor visual navigation method based on causal attention according to claim 1, characterized in that Train the indoor visual navigation model according to the following steps: C1. Use the indoor visual image dataset as the training dataset and calculate to obtain the clustering centers; C2. Extract a navigation trajectory data from the training dataset and use all or part of it as the navigation trajectory data for this round of training; C3. Extract the visual image sequence of the starting point from the input navigation trajectory data as the initial input visual image sequence, and randomly initialize the historical state features; C4. Take the corresponding position of the input visual image sequence as the current position, extract the visual features of each image in the visual image sequence of the current position, encode to obtain the position features of each observation direction, and obtain the global features of each image according to the distances between the visual features of each image and each cluster center; C5. Incorporate the historical state features into the visual features of each image in the visual image sequence of the current position respectively to obtain the visual image features of each image; Then, calculate the self-attention feature and the causal attention feature of the current position, and fuse the self-attention feature and its causal attention feature to obtain the visual environment state feature; C6. Calculate the correlation between the visual features of the images in the navigable direction in the visual image sequence of the current position and their corresponding visual environment state features according to the preset navigable directions, and predict the navigation action of the current position based on the correlation; C7. Determine whether the end point of the input navigation trajectory data is reached. If so, execute step C9; otherwise, execute step C8; C8. Update the historical state features according to the visual environment state feature of the current position obtained in step C5 and the navigation action of the current position predicted in step C6; Extract the visual image sequence of the next position of the navigation trajectory from the navigation trajectory data, and use this visual image sequence and the updated historical state features as the input, and return to step C4; C9. Calculate the loss according to the preset expert navigation actions and the predicted navigation actions at each position, and update the parameters of the indoor visual navigation model according to the cumulative loss; C10. Repeat steps C2 - C9 for iterative training until the training termination condition is met to obtain the trained indoor visual navigation model.
3. An indoor visual navigation method based on causal attention according to claim 2, characterized in that, In step B, initially, use the cluster centers obtained during training, and use the navigation trajectory data of the indoor visual image dataset during training as the initial historical navigation trajectory data; After performing the indoor visual navigation task, collect the navigation trajectory data of the actually completed navigation tasks. After the collection reaches the set quantity, update the historical navigation trajectory data according to the collected navigation trajectory data, and update the cluster centers based on the updated historical navigation trajectory data.
4. An indoor visual navigation method based on causal attention according to claim 2, characterized in that, In step C9, the cumulative loss is calculated according to the following loss function: L = w1L il + w2L rl where w1 and w2 are both trainable parameters, and L il represents the loss generated by imitation learning, and L rl represents the loss generated by reinforcement learning. The reinforcement learning adopts an actor-critic framework, where the actor network is an indoor visual navigation model and the critic network is a feedforward neural network; wherein, L il and L rl are calculated according to the following formulas respectively: Among them, a t represents the predicted navigation action of the position at time t, represents the preset expert navigation action of the position at time t, π t represents the correlation between the visual feature of the visual image sequence and the corresponding visual environment state feature of the position at time t, G t represents the cumulative reward of the executor network at the position at time t, TD t is the output of the critic network at the position at time t and is calculated by the following formula: TD t = max(0, π t W TD1 )W TD2 Among them, W TD1 and W TD2 are trainable parameters.
5. An indoor visual navigation method based on causal attention according to claim 4, characterized in that, Calculate the cumulative revenue G of the executor network according to the following formula t : Among them, p cur represents the position at the next moment corresponding to the predicted navigation action at time t, p goal represents the position at the next moment corresponding to the expert navigation action at time t, dis(·) represents the Euclidean distance, γ t represents the attenuation factor at time t.
6. A method for indoor visual navigation based on causal attention according to any one of claims 1, 2 or 3, characterized in that, The calculation of the cluster centers includes: D1. Extract the visual features of each image in the navigation image sequence of each navigation trajectory data, and form a global visual feature dataset with all the extracted visual features; D2. Set and initialize K cluster centers; D3. Calculate the Euclidean distances between each visual feature in the global visual feature dataset and each cluster center respectively; D4. Classify each visual feature based on the minimum distance between each visual feature and each cluster center; D5. Update the values of the cluster centers according to the following formula: Among them, g k represents the value of the k-th clustering center, and C k represents the set of visual features included in the k-th clustering center; D6. Repeat the above steps D3 - D5 to iteratively update the values of the cluster centers until the change in all cluster center values is less than a preset threshold or exceeds a preset number of iteration rounds.
7. A method for indoor visual navigation based on causal attention according to any one of claims 1 or 2, characterized in that, Obtain the global feature according to the distance between the visual features of each image in the visual image sequence of the current position and each clustering center wherein, represents the global feature of the image in the i-th observation direction, N is the number of observation directions, and it is calculated according to the following steps: Calculate the distances between the visual features of the image in the $i$-th observation direction and the $K$ cluster centers respectively, and take the mean of the distances to the $K$ cluster centers as its global feature 8. A method for indoor visual navigation based on causal attention according to any one of claims 1 or 2, characterized in that Integrate the historical state features into the visual features of each image in the current position visual image sequence respectively to obtain the visual image features of each image, including: First, perform global average pooling on the visual feature F t ={f1, f2, … f i , …, f N} separately; Then, in the form of vector splicing, the historical state feature H t-1 is respectively incorporated into each visual feature after global average pooling to obtain the visual image feature C t ={c1, c2, … c i , …, c N}, where t represents the current position and t-1 represents the previous position of the current position.
9. A method for indoor visual navigation based on causal attention according to any one of claims 1 or 2, characterized in that, The position feature is encoded with absolute position using a pre-trained BERT model.
10. A method for indoor visual navigation based on causal attention according to any one of claims 1 or 2, characterized in that, In each step, a residual neural network is used to extract the visual features of the image.
11. A method for indoor visual navigation based on causal attention according to any one of claims 1 or 2, characterized in that, Fuse the visual image features and their position features of each image, and calculate the self-attention features of each image in the current position visual image sequence through a self-attention mechanism, including: First, the visual image features and their position features are fused by splicing. Then, through a multi-layer perceptron network with different parameters, the features obtained by fusion are converted into a query vector Q s , key vector K s and value vector V s : Q s = max(0, (C t + PE t ) W qs + b qs ) K s = max(0, (C t + PE t )W ks + b ks ) V s = max(0, (C t + PE t )W vs + b vs ) Among them, C t represents the visual image feature at the current position, and PE t represents the position feature at the current position. W qs , b qs , W ks , b ks , W vs and b vs are all parameters of the multi-layer perceptron network; Then, calculate the attention weight a s : where dim is the dimension of the multi-layer perceptron network, and T represents matrix transpose; Finally, calculate the self-attention features through the attention weights and value vectors: SA t = softmax(a s V s ) Among them, SA t represents the self-attention feature of the current position.
12. A method for indoor visual navigation based on causal attention according to any one of claims 1 or 2, characterized in that, Fuse the visual image features and position features of each image to construct a query vector; construct a key vector and a value vector according to the global features of each image, and then, based on the constructed query vector, key vector and value vector, calculate the causal attention features of each image in the current position visual image sequence through a causal attention mechanism, including: First, the visual image features and their position features are fused by splicing, and then, through a multi-layer perceptron network, the features obtained by fusion are converted into a query vector Q c : Q c = max(0, (C t + PE t ) W qc + b qc ) And through a multi-layer perceptron network with different parameters, the global features of the visual image sequence corresponding to the current position are converted into a key vector K c and a value vector V c : Among them, C t represents the visual image feature at the current position, and PE t represents the position feature at the current position, represents the global feature, and W qc , b qc , W kc , b kc , W vc and b vc are all parameters of the multi-layer perceptron network; Then, calculate the attention weight a c : where dim is the dimension of the multi-layer perceptron network, and T represents matrix transpose; Finally, calculate the causal attention features through the attention weights and value vectors: CA t = softmax(a c V c ) Among them, CA t represents the causal attention feature at the current position.
13. A method for indoor visual navigation based on causal attention according to any one of claims 1 or 2, characterized in that Fuse the self-attention features and their causal attention features of each image to obtain the visual environment state features of each image in the current position visual image sequence, including: First, through the method of vector splicing, fuse the self-attention feature SA t and the causal attention feature CA t , to obtain the fused feature [SA t , CA t ; Then, a feedforward neural network is used to convert the fused features [SA t , CA t into the visual environment state feature S t : S t = max(0, [SA t , CA t W ffn1 + b ffn1 )W ffn2 + b ffn2 Among them, are all parameters of the feedforward neural network. dim is the dimension of the encoding network for constructing query vectors, key vectors, and value vectors in the attention calculation, and N is the number of observation directions.
14. The indoor visual navigation method based on causal attention according to claim 2, wherein The navigation trajectory data also includes the navigable direction labels at each position of the navigation trajectory. In step C6, only the directions with navigable direction labels are regarded as navigable directions; in step B4, all observed directions are regarded as navigable directions.
15. A method for indoor visual navigation based on causal attention according to any one of claims 1, 2 or 14, characterized in that, Calculate the correlation between the visual features of the images in the navigable directions in the current position visual image sequence and their corresponding visual environment state features according to the preset navigable directions, and predict the navigation action at the current position, including: First, calculate the correlation π between the visual features of the images in each navigable direction of the current position visual image sequence and the corresponding visual environment state features t : Among them, represents the visual features of the images of the navigable directions in the current position visual image sequence, S t represents the visual environment state features of the images of the navigable directions in the current position visual image sequence; Then, according to the relevance π t predict the navigation action a at the current position t : a t = argmax m π t,m Among them, π t,m represents the correlation of the m-th direction in the π t sequence.
16. A method for indoor visual navigation based on causal attention according to any one of claims 1 or 2, characterized in that, Update the historical state features according to the visual environment state features at the current position and the predicted navigation action at the current position, including: First, by resetting the gate, screen the visual environment state feature S at the current position t and the predicted navigation action a at the current position t of the key features, and fuse them into the historical state feature H at the previous moment at the current position t-1 : r t = σ(W r H t-1 + U r [S t , π t , a t ) Among them, π t represents the correlation between the visual features of each navigable direction image in the visual image sequence at the current position and the visual environment state features corresponding thereto, r t represents the forgetting gate weight, W r 、U r 、W g and U g are all trainable parameters, σ(·) and tanh(·) represent activation functions, ⊙ represents the Hadamard product operation, t represents the current position, and t - 1 represents the previous position of the current position; Then, through the update gate, filter the valid historical information z to be retained t , and fuse it into the historical state feature H at the previous moment of the current position t-1 , and update the historical state feature: z t = σ(W z H t-1 + U z [S t , π t , a t ) where z t represents the update gate weights, W z and U z are both trainable parameters.
Citation Information
Patent Citations
Visual language indoor navigation method and system, terminal and application
CN112710310A
Self-adaptive target navigation method and system for service robot
CN114460943A