A visual-inertial odometry feature fusion method based on deep learning

Through the deep learning visual inertial odometer feature fusion method, the characteristics are screened using spatial channels and cross attention mechanisms, and combined with the decoder to adjust the weight, the efficient feature fusion of the visual inertial odometer is achieved, the positioning accuracy and robustness are improved, and the sensor calibration and synchronization problems are solved.

CN116975780BActive Publication Date: 2025-08-19GUANGZHOU KANGQI INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310956646.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-31
Publication Date
2025-08-19
Estimated Expiration
2043-07-31

AI Technical Summary

Technical Problem

During the feature fusion process of the existing visual inertial odometry method, the positioning accuracy is low and affected by sensor calibration and time synchronization, resulting in insufficient data utilization efficiency and accuracy.

Method used

The visual inertial odometry feature fusion method based on deep learning is adopted to filter image features through the spatial channel attention module, the cross attention module filters inertial navigation features, and the decoder structure is used to adjust the fusion feature weight to achieve efficient fusion of images and inertial navigation features.

Benefits of technology

It improves the positioning accuracy of the visual inertial mileage calculation method, reduces the average translation error and rotation error, solves the impact of synchronization problems and sensor data degradation, and improves positioning performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116975780B_ABST
    Figure CN116975780B_ABST
Patent Text Reader

Abstract

The present invention discloses a visual inertial odometry feature fusion method based on deep learning, comprising the following steps: S1, obtaining two adjacent frames of images captured by a camera, and inputting the two adjacent frames of images into an image feature extraction network to extract image features; S2, obtaining inertial navigation information between the two adjacent frames of images, and inputting the inertial navigation information into an inertial navigation feature extraction network to extract inertial navigation features; S3, according to the image features extracted in step S1, inputting the image features into a spatial channel attention module to screen the image features; S4, according to the image features extracted in step S1 and the inertial navigation features extracted in step S2, inputting the image features and the inertial navigation features into a cross attention module to screen the inertial navigation features; S5, splicing the screened image features and the inertial navigation features, and adjusting the weights of the fused features using a decoder structure to obtain final fused features. The method efficiently fuses the screened image and inertial navigation features, thereby improving the positioning accuracy of the visual inertial odometry method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous positioning, and in particular to a visual-inertial odometry feature fusion method based on deep learning. Background Art

[0002] Visual inertial odometry is a system that utilizes both visual and inertial navigation information to achieve positioning. With the advancement of hardware, the deployment cost of visual inertial odometry systems continues to decrease, and their application is becoming increasingly widespread across various fields. Visual inertial odometry systems primarily input data collected by two types of sensors: cameras and IMUs. They extract features from this data and calculate pose to achieve positioning. Therefore, effectively integrating image and inertial navigation features to improve positioning accuracy has long been a key research topic in the positioning field.

[0003] Traditional visual-inertial odometry methods extract low-level features such as points and lines during feature extraction. Feature fusion is achieved through loose coupling and tight coupling, respectively, at the decision-level and feature set. The loose coupling approach fuses data at the decision-level, processing only the final result and underutilizing features. While the tight coupling approach fuses at the feature level, its results are affected by sensor calibration and time synchronization.

[0004] In recent years, the application of deep learning in computer vision has gradually increased. In the field of visual inertial odometry, many deep learning-based models have emerged. These models can extract high-dimensional motion features from images and inertial navigation, while effectively solving the difficulty of sensor parameter calibration. However, during the feature fusion process, these models only perform splicing or simple selection, without effectively integrating and utilizing the two types of features, resulting in inaccurate positioning results. As a result, positioning methods suffer from low data utilization efficiency and low accuracy. Summary of the Invention

[0005] In response to the above-mentioned deficiencies in the prior art, the present invention provides a visual-inertial odometry feature fusion method based on deep learning. This method optimizes the feature fusion algorithm to efficiently fuse the screened image features with the inertial navigation features, effectively improving the positioning accuracy of the visual-inertial odometry method while solving synchronization problems and the impact of sensor data degradation to a certain extent.

[0006] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0007] A deep learning-based visual-inertial odometry feature fusion method includes the following steps:

[0008] S1, obtain two adjacent frames of pictures captured by the camera, and input the two adjacent frames of pictures into the image feature extraction network to extract image features;

[0009] S2. Obtain inertial navigation information between two adjacent frames of images, and input the inertial navigation information into an inertial navigation feature extraction network to extract inertial navigation features;

[0010] S3, according to the image features extracted in step S1, input the image features into the spatial channel attention module to screen the image features;

[0011] S4, according to the image features extracted in step S1 and the inertial navigation features extracted in step S2, inputting the image features and the inertial navigation features into a cross attention module to screen the inertial navigation features;

[0012] S5. Combine the filtered image features with the inertial navigation features, and use the decoder structure to adjust the weight of the fusion features to obtain the final fusion features.

[0013] Furthermore, step S1 specifically includes the following steps:

[0014] S11, inputting two adjacent frames of images captured by the camera into a convolutional neural network based on a step structure;

[0015] S12, stacking two adjacent frames of input images together;

[0016] S13. Extract image features based on two stacked adjacent frames using a convolutional neural network based on a step structure.

[0017] Furthermore, the step-structured convolutional neural network in step S11 adopts an encoder structure of an optical flow network, specifically including:

[0018] The 7×7 convolution kernel of 64 channels, the 5×5 convolution kernel of 128 channels, the 5×5 convolution kernel of 256 channels, the 3×3 convolution kernel of 256 channels, and the four 3×3 convolution kernels of 512 channels are connected in sequence, and a step structure is added to the 5×5 convolution kernel of 128 channels and 256 channels and the two 3×3 convolution kernels of 512 channels in the encoder structure;

[0019] The step structure adds the image features before the 5×5 convolution kernel extraction to the extracted image features and feeds them into the 3×3 convolution kernel to extract the image features.

[0020] Furthermore, step S2 specifically includes:

[0021] The matrix information composed of the inertial measurement unit data between two adjacent frames of images is obtained, the inertial navigation information is obtained using the matrix information, and the inertial navigation information is input into the long short-term memory network to extract the inertial navigation features.

[0022] Furthermore, step S3 specifically includes the following steps:

[0023] S31, input the extracted image features into the spatial channel attention module;

[0024] S32. Based on the input image features, the spatial channel attention module is used to calculate the attention weights reflecting the feature correlation from the two dimensions of spatial attention and channel attention.

[0025] S33. Multiply the obtained attention weight by the input image feature to obtain the filtered image feature.

[0026] Furthermore, the calculation formula for image feature screening in step S33 is as follows:

[0027]

[0028] in, Represents element-wise multiplication, F represents the input image features, M C represents the one-dimensional channel attention weight, F′ represents the image features after the channel attention network is filtered, and M S represents the two-dimensional spatial attention weight, and F″ represents the image features after filtering by the spatial attention network.

[0029] Furthermore, step S4 specifically includes the following steps:

[0030] S41, inputting the extracted image features and inertial navigation features into a cross attention module;

[0031] S42. According to the correlation between the input image features and the inertial navigation features, a set of attention weights are obtained through cross-attention network learning;

[0032] S43. Multiply the obtained attention weight by the inertial navigation feature to obtain the filtered inertial navigation feature.

[0033] Furthermore, the calculation formula for inertial navigation feature screening in step S43 is as follows:

[0034] O=softmax((W Q S2)(W K S1) T W V S1)

[0035] Among them, O represents the filtered inertial navigation feature, softmax(·) represents the normalization function, S1 represents the input inertial navigation feature, S2 represents the input image feature, and W Q Represents the matrix corresponding to Q in the cross attention mechanism, W K Represents the matrix corresponding to K in the cross attention mechanism, W V Represents the matrix corresponding to V in the cross attention mechanism, and T represents the matrix transpose operation.

[0036] Furthermore, step S5 specifically includes the following steps:

[0037] S51, splicing the filtered image features and inertial navigation features along the channel direction to obtain fusion features;

[0038] S52, adjusting the weight of the fused feature using a decoder structure according to the obtained fused feature to obtain an adjusted weight;

[0039] S53: Multiply the obtained fusion feature by the adjusted weight to obtain the final fusion feature.

[0040] Furthermore, the calculation formula of the fusion feature in step S51 is:

[0041]

[0042] Among them, F1 represents the fusion feature, F v Represents the filtered image features, F i represents the filtered inertial navigation characteristics, Indicates splicing along the channel direction.

[0043] Furthermore, the calculation formula of the adjusted weight in step S52 is:

[0044] W=σ(f(wF1+b))

[0045] Among them, W represents the adjusted weight, σ represents the sigmoid activation function, f represents the function relu corresponding to the decoder, w and b represent the parameters learned during network training, and F1 represents the fusion feature.

[0046] Furthermore, the calculation formula of the final fusion feature in step S53 is:

[0047]

[0048] Among them, F1 represents the fusion feature, W represents the adjusted weight, and F2 represents the final fusion feature. Represents element-wise multiplication.

[0049] The present invention has the following beneficial effects:

[0050] The present invention adopts a spatial channel attention mechanism to screen image features, fully considering the relationship between feature space and channels. It also adopts a cross-attention mechanism to screen inertial navigation features, introducing image features to guide the selection of inertial navigation features, fully leveraging their complementarity. Finally, a self-attention mechanism is used to process the fused information, and feature fusion is completed by utilizing the correlation between image and inertial navigation features. This method reduces the average translation error and average rotation error of positioning performance by 10% compared with the fusion method, improving the positioning accuracy of the visual inertial odometry method while, to a certain extent, addressing synchronization issues and the effects of sensor data degradation. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is a flow chart of a deep learning-based visual-inertial odometry feature fusion method of the present invention;

[0052] Figure 2 Schematic diagram of image feature extraction;

[0053] Figure 3 This is an overall schematic diagram of a deep learning-based visual-inertial odometry feature fusion method of the present invention. DETAILED DESCRIPTION

[0054] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0055] like Figure 1 As shown, a visual inertial odometry feature fusion method based on deep learning includes the following steps S1-S4:

[0056] S1. Obtain two adjacent frames of images captured by the camera, and input the two adjacent frames of images into the image feature extraction network to extract image features.

[0057] In an optional embodiment of the present invention, this embodiment uses a convolutional neural network based on a step structure to extract image features. The convolutional neural network based on the step structure adopts the encoder structure of the FlowNetSimple optical flow network. The encoder structure can effectively extract image features and add a step structure at the 5×5 convolution kernel and the 3×3 convolution kernel of the encoder structure. The added step structure is used to extract multi-scale information of the image.

[0058] like Figure 2 As shown, step S1 includes the following steps S11-S13:

[0059] S11. Input two adjacent frames of images captured by the camera into a convolutional neural network based on a step structure.

[0060] S12: stack two adjacent input frames together.

[0061] S13. Extract image features based on two stacked adjacent frames using a convolutional neural network based on a step structure.

[0062] Specifically, the step-structure-based convolutional neural network in step S11 adopts an encoder structure of an optical flow network, specifically including:

[0063] The 7×7 convolution kernel of 64 channels, the 5×5 convolution kernel of 128 channels, the 5×5 convolution kernel of 256 channels, the 3×3 convolution kernel of 256 channels, and four 3×3 convolution kernels of 512 channels are connected in sequence, and a step structure is added to the 5×5 convolution kernel of 128 channels and 256 channels and the two 3×3 convolution kernels of 512 channels in the encoder structure.

[0064] The step structure adds the image features before the 5×5 convolution kernel extraction to the extracted image features and feeds them into the 3×3 convolution kernel to extract the image features.

[0065] S2. Obtain inertial navigation information between two adjacent frames of images, and input the inertial navigation information into an inertial navigation feature extraction network to extract inertial navigation features.

[0066] In an optional embodiment of the present invention, the function of the inertial navigation feature extraction network in this embodiment is to extract features related to the timing of the inertial measurement unit data. In deep learning, recurrent neural networks are generally used to extract inertial navigation features. However, recurrent neural networks cannot retain information over a large time range, resulting in the loss of previous information after a certain period of time. In addition, recurrent neural networks also have the problem of large number of parameters and difficulty in training. Therefore, this embodiment uses a long short-term memory network instead of a recurrent neural network to extract the inertial navigation information of the inertial navigation sequence.

[0067] In this embodiment, the results of packaging 10 groups of inertial measurement units between two adjacent frames of images captured by the camera are used as the input of the inertial navigation feature extraction network. Based on the matrix information composed of the 10 frames of inertial measurement unit data, inertial navigation information is analyzed and obtained, and inertial navigation features are extracted using a long short-term memory network based on the obtained inertial navigation information.

[0068] like Figure 3As shown, in an optional embodiment of the present invention, this embodiment screens and fuses images and inertial navigation features from different sources in multiple dimensions based on the spatial channel attention mechanism, the cross-attention mechanism, and the self-attention mechanism, making full use of the spatial information in the pathological full-slice image and the temporal information in the inertial navigation data, while considering the complementarity of the image features and the inertial navigation features, thereby improving the positioning accuracy of the visual inertial odometry method, and being able to adapt to situations where data is not strictly synchronized or sensor degradation occurs, artificial simulation data is time-out, image occlusion, etc., and relatively accurate positioning can still be achieved after feature screening and fusion, making this method have good practicality and application prospects.

[0069] S3. According to the image features extracted in step S1, the image features are input into the spatial channel attention module to screen the image features.

[0070] In an optional embodiment of the present invention, this embodiment inputs the image features extracted by the convolutional neural network into the spatial channel attention module to perform adaptive feature screening on the image features.

[0071] Specifically, step S3 includes the following steps S31-S33:

[0072] S31. Input the extracted image features into the spatial channel attention module.

[0073] S32. Based on the input image features, the spatial channel attention module is used to calculate the attention weights of the response feature correlation from the two dimensions of spatial attention and channel attention.

[0074] S33. Multiply the obtained attention weight by the input image feature to obtain the filtered image feature.

[0075] Specifically, the calculation formula for image feature screening in step S33 is as follows:

[0076]

[0077] in, Represents element-wise multiplication, F represents the input image features, M C represents the one-dimensional channel attention weight and M C ∈R C×1×1 , F′ represents the image features after filtering by the channel attention network, M S represents the two-dimensional spatial attention weight and M S ∈R 1×H×W , F″ represents the image features after filtering by the spatial attention network, R C×1×1 Represents the one-dimensional channel tensor dimension, R 1 ×H×WRepresents the dimension of a two-dimensional spatial tensor, R represents the matrix, C represents the number of channels, and H and W represent the height and width respectively.

[0078] S4. According to the image features extracted in step S1 and the inertial navigation features extracted in step S2, the image features and the inertial navigation features are input into a cross attention module to screen the inertial navigation features.

[0079] In an optional embodiment of the present invention, the cross-attention module in this embodiment can accept different feature sequences and guide the selection of inertial navigation features through the correlation between image features and inertial navigation features.

[0080] Specifically, step S4 includes the following steps S41-S43:

[0081] S41. Input the extracted image features and inertial navigation features into the cross attention module.

[0082] S42. According to the correlation between the input image features and the inertial navigation features, a set of attention weights are obtained through cross-attention network learning.

[0083] S43. Multiply the obtained attention weight by the inertial navigation feature to obtain the filtered inertial navigation feature.

[0084] Specifically, the calculation formula for inertial navigation feature screening in step S43 is as follows:

[0085] O=softmax((W Q S2)(W K S1) T W V S1)

[0086] Among them, O represents the filtered inertial navigation feature, softmax(·) represents the normalization function, S1 represents the input inertial navigation feature, S2 represents the input image feature, and W Q Represents the matrix corresponding to Q in the cross attention mechanism, W K Represents the matrix corresponding to K in the cross attention mechanism, W V Represents the matrix corresponding to V in the cross attention mechanism, and T represents the matrix transpose operation.

[0087] S5. Combine the filtered image features with the inertial navigation features, and use the decoder structure to adjust the weight of the fusion features to obtain the final fusion features.

[0088] In an optional embodiment of the present invention, the decoder structure in this embodiment first processes the spliced fusion features through a linear function, and uses the relu function and the sigmoid function to generate a set of weights, and finally multiplies the generated weights with the spliced fusion features to obtain the final fusion features, thereby completing feature fusion.

[0089] Specifically, step S5 includes the following steps S51-S53:

[0090] S51. Concatenate the filtered image features and the inertial navigation features along the channel direction to obtain fused features.

[0091] S52. Adjust the weight of the fusion feature using a decoder structure according to the obtained fusion feature to obtain an adjusted weight.

[0092] S53: Multiply the obtained fusion feature by the adjusted weight to obtain the final fusion feature.

[0093] Specifically, the calculation formula for the fusion feature in step S51 is:

[0094]

[0095] Among them, F1 represents the fusion feature, F v Represents the filtered image features, F i represents the filtered inertial navigation characteristics, Indicates splicing along the channel direction.

[0096] Specifically, the calculation formula of the adjusted weight in step S52 is:

[0097] W=σ(f(wF1+b))

[0098] Among them, W represents the adjusted weight, σ represents the sigmoid activation function, f represents the function relu corresponding to the decoder, w and b represent the parameters learned during network training, which are used for preliminary processing of fusion features, and F1 represents the fusion feature.

[0099] Specifically, the calculation formula of the final fusion feature in step S53 is:

[0100]

[0101] Among them, F1 represents the fusion feature, W represents the adjusted weight, and F2 represents the final fusion feature. Represents element-wise multiplication.

[0102] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

[0103] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.

Claims

1. A visual inertial odometry feature fusion method based on deep learning, characterized in that: The following steps are involved: S1. Obtain two adjacent frames of images captured by a camera, and input the two adjacent frames of images into an image feature extraction network to extract image features; wherein the image feature extraction network includes a convolutional neural network based on a step structure; The convolutional neural network based on the step structure adopts the encoder structure of the optical flow network, specifically including: a 7×7 convolution kernel with a size of 64 channels, a 5×5 convolution kernel with a size of 128 channels, a 5×5 convolution kernel with a size of 256 channels, a 3×3 convolution kernel with a size of 256 channels, and four 3×3 convolution kernels with a size of 512 channels, which are connected in sequence. A step structure is added at the 5×5 convolution kernels with a size of 128 channels and 256 channels, and at the two 3×3 convolution kernels with a size of 512 channels in the encoder structure; the step structure adds the image features before extraction by the 5×5 convolution kernel to the extracted image features and feeds them into the 3×3 convolution kernel to extract the image features; S2. Obtain inertial navigation information between two adjacent frames of images, and input the inertial navigation information into an inertial navigation feature extraction network to extract inertial navigation features; S3, according to the image features extracted in step S1, input the image features into the spatial channel attention module to screen the image features; S4, according to the image features extracted in step S1 and the inertial navigation features extracted in step S2, inputting the image features and the inertial navigation features into a cross attention module to screen the inertial navigation features; S5. Combine the filtered image features with the inertial navigation features, and use the decoder structure to adjust the weight of the fusion features to obtain the final fusion features, specifically: S51, splicing the filtered image features and inertial navigation features along the channel direction to obtain fusion features; S52, adjusting the weight of the fused feature using a decoder structure according to the obtained fused feature to obtain an adjusted weight; S53: Multiply the obtained fusion feature by the adjusted weight to obtain the final fusion feature.

2. The method for visual inertial odometry feature fusion based on deep learning according to claim 1, characterized in that: Step S1 specifically includes the following steps: S11, inputting two adjacent frames of images captured by the camera into a convolutional neural network based on a step structure; S12, stacking two adjacent frames of input images together; S13. Extract image features based on two stacked adjacent frames using a convolutional neural network based on a step structure.

3. The method for fusion of visual inertial odometry features based on deep learning according to claim 1, characterized in that: Step S2 specifically includes: The matrix information composed of the inertial measurement unit data between two adjacent frames of images is obtained, the inertial navigation information is obtained using the matrix information, and the inertial navigation information is input into the long short-term memory network to extract the inertial navigation features.

4. The method for fusion of visual inertial odometry features based on deep learning according to claim 1, characterized in that: Step S3 specifically includes the following steps: S31, input the extracted image features into the spatial channel attention module; S32. Based on the input image features, the spatial channel attention module is used to calculate the attention weights reflecting the feature correlation from the two dimensions of spatial attention and channel attention. S33. Multiply the obtained attention weight by the input image feature to obtain the filtered image feature.

5. The method for fusion of visual inertial odometry features based on deep learning according to claim 4, characterized in that: The calculation formula for image feature screening in step S33 is as follows: , in, represents element-wise multiplication, represents the input image features, represents the one-dimensional channel attention weight, represents the image features after filtering by the channel attention network, represents the two-dimensional spatial attention weight, Represents the image features after filtering by the spatial attention network.

6. The method for fusion of visual inertial odometry features based on deep learning according to claim 1, characterized in that: Step S4 specifically includes the following steps: S41, inputting the extracted image features and inertial navigation features into a cross attention module; S42. According to the correlation between the input image features and the inertial navigation features, a set of attention weights are obtained through cross-attention network learning; S43. Multiply the obtained attention weight by the inertial navigation feature to obtain the filtered inertial navigation feature.

7. The method for fusion of visual inertial odometry features based on deep learning according to claim 6, characterized in that: The calculation formula for inertial navigation feature screening in step S43 is as follows: in, represents the filtered inertial navigation characteristics, represents the normalization function, represents the input inertial navigation characteristics, represents the input image features, Indicates the cross attention mechanism The corresponding matrix, Indicates the cross attention mechanism The corresponding matrix, Indicates the cross attention mechanism The corresponding matrix, Represents a matrix transpose operation.

8. The method for visual inertial odometry feature fusion based on deep learning according to claim 1, characterized in that: The calculation formula of the fusion feature in step S51 is: in, represents the fusion feature, represents the filtered image features, represents the filtered inertial navigation characteristics, Indicates splicing along the channel direction; The calculation formula of the adjusted weight in step S52 is: in, represents the adjusted weight, represents the sigmoid activation function, Represents the function relu corresponding to the decoder, 、 Both represent the parameters learned during network training, represents fusion features; The calculation formula of the final fusion feature in step S53 is: in, represents the fusion feature, represents the adjusted weight, represents the final fusion feature, Represents element-wise multiplication.

Citation Information

Patent Citations

  • Monocular vision inertial navigation positioning method based on self-supervised deep learning

    CN114526728A

  • LTC-DNN-based visual inertial navigation combined navigation system and self-learning method

    WO2022262878A1