Six-degree-of-freedom visual inertial odometer method and equipment based on double attention and mixed basis function

Through the visual inertial odometer method of double attention and mixed basis function, the accuracy and robustness of the position estimation of the visual inertial odometer in complex environments is solved, and high-precision and stable position estimation are achieved.

CN120232428APending Publication Date: 2025-07-01HEFEI UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510462087.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing visual inertial odometry methods are prone to drift in low-texture or dynamic environments, and deep learning methods have problems of redundant information loss and insufficient feature extraction when dealing with complex nonlinear motion, making it difficult to achieve high-precision and robust pose estimation.

Method used

The six-degree of freedom visual inertial odometry method based on double attention and mixed basis functions are adopted to extract visual features through deep convolutional neural networks, combine with one-dimensional convolutional neural network to process inertial data, and use the policy network and the median Kolmogolov-Arnold network to perform adaptive data fusion and nonlinear feature enhancement to achieve high-precision pose estimation.

Benefits of technology

High-precision and robust six-degree-of-freedom pose estimation is achieved in complex motion scenarios, which improves feature discrimination ability and nonlinear motion modeling effect, reduces the computational burden, and enhances the ability to capture complex motion dynamics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120232428A_ABST
    Figure CN120232428A_ABST
Patent Text Reader

Abstract

The invention relates to a six-degree-of-freedom visual inertial odometer method based on double attention and a mixed basis function. The method comprises the following steps: collecting visual image data and inertial data and carrying out preprocessing; obtaining a visual feature vector; extracting an inertial feature vector; generating a final fusion feature; and inputting the final fusion features into a long and short term memory network for time sequence modeling, carrying out nonlinear feature enhancement by using a median Colmogorov-Arnold network, and outputting six-degree-of-freedom pose estimation in a target environment through a full-connection regression layer. According to the method, through a self-adaptive regulation and control mechanism of vision-inertia data, redundant information is compressed, and meanwhile key motion features are reserved; a double attention module is innovatively introduced to improve the feature discrimination ability, the mixed basis function modeling ability of the median Colmogolov-Arnold network is combined to realize the accurate representation of the complex nonlinear motion dynamics, and finally stable and high-precision six-degree-of-freedom pose estimation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of visual inertial odometry, and in particular to a six-degree-of-freedom visual inertial odometry method and device based on dual attention and hybrid basis functions. Background Art

[0002] Existing visual inertial odometry methods mainly estimate the six-degree-of-freedom pose of a device by fusing visual information obtained by a camera and motion data collected by an inertial measurement unit (IMU). Traditional geometry-based VIO methods, such as algorithms based on keyframe optimization or filtering, rely on manually extracted features and strict model assumptions, and are prone to problems such as drift or decreased estimation accuracy in low-texture or dynamic environments.

[0003] In recent years, with the rapid development of deep learning technology, data-driven VIO methods have gradually become a research hotspot. They automatically learn robust feature representations through end-to-end network training, and to a certain extent improve the accuracy and robustness of pose estimation. However, existing deep learning methods still have two major deficiencies: on the one hand, high-dimensional convolutional features often contain a large amount of redundant information, and directly compressing or fusing them may lose key details; on the other hand, traditional recurrent neural networks, such as LSTM, have limitations in capturing complex non-linear motion dynamics and cannot fully express the mutations and non-Gaussian characteristics in motion. Especially in practical applications, due to the influence of factors such as illumination changes and motion blur on the quality of visual data, how to adaptively regulate the contributions of visual and inertial data to achieve more accurate and stable pose estimation is still a key technical problem to be solved urgently.

[0004] Therefore, a visual inertial odometry method that can effectively fuse visual and inertial information and strengthen the ability to capture complex motion dynamics through an adaptive mechanism is needed to meet the requirements of high precision and robustness. Summary of the Invention

[0005] To solve the problems of deficiencies in feature extraction, data fusion, and non-linear motion modeling of traditional visual inertial odometry, the primary object of the present invention is to provide a six-degree-of-freedom visual inertial odometry method based on dual attention and hybrid basis functions that improves feature discrimination ability and non-linear motion modeling effect and realizes high-robustness and high-precision pose estimation in complex motion scenarios.

[0006] To achieve the above object, the present invention adopts the following technical solution: A six-degree-of-freedom visual inertial odometry method based on dual attention and hybrid basis functions, the method comprising the following steps in sequence:

[0007] (1) Acquisition and preprocessing: Using a device equipped with a monocular camera and an inertial measurement unit, visual image data and inertial data are collected in the target motion environment, and the visual image data and inertial data are preprocessed to obtain preprocessed visual image data and preprocessed inertial data;

[0008] (2) Visual feature extraction: A deep convolutional neural network is used to perform multi-scale feature extraction on the preprocessed visual image data, and a dual attention module is introduced. By fusing local details and global context information, a visual feature vector v is obtained t ;

[0009] (3) Inertial feature extraction: One-dimensional convolutional neural network is used to perform temporal modeling on the preprocessed inertial data, and an inertial feature vector i is extracted t ;

[0010] (4) Adaptive data fusion: Using a policy network, the inertial feature vector i t and the previous moment's LSTM hidden state are used as the input of the policy network. The visual feature vector v t is adaptively gated and dynamically fused to generate the final fused feature f t ;

[0011] (5) Temporal modeling and pose estimation: The final fused feature f t is input into a long short-term memory network for temporal modeling, and a median Kolmogorov - Arnold network combined with a hybrid basis function is used to perform non-linear feature enhancement on the output of the long short-term memory network. The six-degree-of-freedom pose estimation in the target environment is output through a fully connected regression layer.

[0012] Step (1) specifically refers to: The visual image data is collected by a high-resolution monocular camera with an image frame rate not lower than 10 Hz. The visual image data is preprocessed, that is, grayscale conversion, denoising, and geometric correction are performed to enhance the image quality, and preprocessed visual image data is obtained; The inertial data is collected by an inertial measurement unit at a fixed sampling frequency. The inertial data is preprocessed, that is, low-pass filtering, time synchronization processing are performed, and then weighted according to the statistical distribution of the motion rotation information to obtain preprocessed inertial data.

[0013] Step (2) specifically includes the following steps:

[0014] (2a) Two consecutive frames of preprocessed visual images I t and I t+1 are concatenated along the channel dimension to obtain a composite image V t :

[0015]

[0016] Among them, cat represents the image stitching operation, and H and W are the height and width of the image respectively;

[0017] (2b) Input the composite image V t into a deep convolutional neural network composed of multiple layers of convolution, batch normalization, and LeakyReLU activation function for visual feature extraction. The output feature map of the l-th layer satisfies:

[0018]

[0019] Among them, the initial feature is the composite image V t , Conv (l) represents the l-th layer convolution operation, BN represents the batch normalization operation, and σ(·) is the LeakyReLU activation function; after being processed by the deep convolutional neural network, the final output feature map is denoted as Among them, C represents the number of channels, and H' and W' represent the height and width of the final output feature map respectively;

[0020] (2c) Integrate a dual attention module in the deep convolutional neural network to enhance the expression ability of visual features. The dual attention module includes a local-global information fusion module and a channel recalibration module;

[0021] Input the final output feature map F t into the local-global information fusion module. The local-global information fusion module processes F t using three parallel branches:

[0022] The first branch uses a 1×1 convolution to achieve channel remapping, and the output feature is denoted as

[0023] The second branch uses a 3×3 convolution to capture local spatial information, and the output feature is denoted as

[0024] The third branch performs adaptive average pooling and then generates a global attention map A through a 1×1 convolution and Sigmoid activation global :

[0025]

[0026] Among them, W g is a learnable parameter matrix, |Ω| represents the spatial region area of the feature map, and x is the spatial coordinate;

[0027] Subsequently, multiply A global element-wise with F t and then output the output feature Output feature A global ⊙F t Add them together and apply the SiLU activation function to the result to obtain the fused feature map F'. t :

[0028]

[0029] Among them, ⊙ represents element-wise multiplication;

[0030] The processing of the channel recalibration module is as follows:

[0031] For the fused feature map F' t Perform global average pooling, then generate the channel weights denoted as A through one-dimensional convolution and Sigmoid activation channel ;

[0032] Multiply A channel element-wise with F' t to obtain the recalibrated feature map

[0033]

[0034] Finally, Flatten and map through a fully connected layer to output the visual feature vector enhanced by the dual attention module d v represents the dimension of the visual feature vector v t .

[0035] Step (3) specifically includes the following steps:

[0036] (3a) Reconstruct the preprocessed inertial data into a format suitable for one-dimensional convolution processing;

[0037] (3b) Apply multi-layer one-dimensional convolution operations to the reconstructed inertial data, and each layer of convolution combines batch normalization and the LeakyReLU activation function to fully capture the temporal dynamic characteristics in the inertial signal;

[0038] (3c) Flatten the feature sequence processed by the multi-layer convolutional layer and map it through a fully connected layer to obtain the inertial feature vector d i represents the dimension of the inertial feature vector i t .

[0039] Step (4) specifically includes the following steps:

[0040] (4a) Construct the composite vector x t to guide the adaptive regulation of the policy network. The composite vector x t is composed of the inertial feature vector it Combined with the previous moment's LSTM hidden state , d r represents the dimension of the LSTM hidden state, and the composite vector x t has the following expression:

[0041]

[0042] (4b) Process x t using the policy network to output the decision vector d t . The decision vector is discretized using the Gumbel-Softmax method to dynamically regulate the weights of the visual feature vector v t ; Subsequently, the regulated visual feature vector is concatenated with the inertial feature vector i t to form the final fusion feature f t .

[0043] Step (5) specifically includes the following steps:

[0044] (5a) Input the final fusion feature into the long short-term memory network to capture temporal correlations, thereby obtaining the LSTM hidden state h t :

[0045] h t = LSTM(f t , h t-1 )

[0046] where is the hidden state of the previous moment; d f is the dimension of the final fusion feature f t ;

[0047] (5b) Apply the median Kolmogorov-Arnold network, i.e., the MKAN network, to the LSTM hidden state h t to achieve non-linear feature expansion. The MKAN network is obtained by improving the Kolmogorov-Arnold network. Specifically: Mix the Gaussian radial basis function and the Lorenz function to obtain a mixed basis function, and replace the Gaussian radial basis function of the Kolmogorov-Arnold network with the mixed basis function; The mixed basis function is:

[0048]

[0049] where g is the kernel center; δ is the scale parameter; ɑ ∈ [0, 1] is a learnable weight, and α = 1 is set at the beginning of training and then gradually adjusted adaptively; Apply the mixed basis function to each component of h t and form a deep non-linear mapping by stacking L layers of MKAN through a linear mapping based on spline functions, and output the enhanced hidden state

[0050]

[0051] Among them, represents the conversion operation of the l-th layer of MKAN;

[0052] (5c) The enhanced hidden state generated by stacking MKAN layers is mapped by the fully connected regression layer to obtain the final six-degree-of-freedom pose estimation

[0053]

[0054] Among them, and are the weight matrix and bias parameter of the regression layer respectively, and d m is the dimension of the MKAN output.

[0055] Another object of the present invention is to provide an electronic device, including:

[0056] a processor; and

[0057] a memory, in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor executes the six-degree-of-freedom visual inertial odometry method based on dual attention and hybrid basis functions as described above.

[0058] The present invention also provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor executes the six-degree-of-freedom visual inertial odometry method based on dual attention and hybrid basis functions as described above.

[0059] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: First, the present invention proposes a visual inertial odometry framework based on deep learning, which realizes high-precision and robust six-degree-of-freedom pose estimation in complex motion scenarios by adaptively balancing the contributions of visual data and inertial data, and reduces the computational burden when the motion is gentle; Second, the dual attention module proposed by the present invention enhances the feature extraction process by adaptively recalibrating the channel response. This module improves the feature discriminability while suppressing redundancy, overcoming the limitations of traditional channel compression techniques in deep networks; Third, by integrating the median Kolmogorov-Arnold network into the pose estimation link, the output of the long short-term memory network is nonlinearly extended, thereby enhancing the ability to capture complex nonlinear motion dynamics. Description of the Drawings

[0060] Figure 1It is the method flow chart of the present invention;

[0061] Figure 2 It is the schematic diagram of trajectory estimation obtained by the present invention on the 10th sequence of the KITTI dataset. Specific embodiments

[0062] As Figure 1 shown, a six-degree-of-freedom visual inertial odometry method based on dual attention and hybrid basis functions, the method includes the following steps in sequence:

[0063] (1) Acquisition and preprocessing: Using a device equipped with a monocular camera and an inertial measurement unit, collect visual image data and inertial data in the target motion environment, and preprocess the visual image data and inertial data to obtain preprocessed visual image data and preprocessed inertial data;

[0064] (2) Visual feature extraction: Use a deep convolutional neural network to perform multi-scale feature extraction on the preprocessed visual image data, and introduce a dual attention module to obtain a visual feature vector v by fusing local details and global context information t ;

[0065] (3) Inertial feature extraction: Use a one-dimensional convolutional neural network to perform temporal modeling on the preprocessed inertial data, and extract an inertial feature vector i t ;

[0066] (4) Adaptive data fusion: Use a policy network, take the inertial feature vector i t and the previous moment's LSTM hidden state as the input of the policy network, perform adaptive gating and dynamic fusion on the visual feature vector v t to generate a final fusion feature f t ;

[0067] (5) Temporal modeling and pose estimation: Input the final fusion feature f t into a long short-term memory network for temporal modeling, and use a median Kolmogorov-Arnold network combined with hybrid basis functions to perform non-linear feature enhancement on the output of the long short-term memory network, and output the six-degree-of-freedom pose estimation in the target environment through a fully connected regression layer.

[0068] Step (1) specifically refers to: The visual image data is collected by a high-resolution monocular camera with an image frame rate not lower than 10 Hz. The visual image data is preprocessed, that is, grayscaled, denoised, and geometrically corrected to enhance the image quality, obtaining the preprocessed visual image data; The inertial data is collected by an inertial measurement unit at a fixed sampling frequency. The inertial data is preprocessed, that is, low-pass filtered and time-synchronized to ensure data stability and time consistency, and then weighted according to the statistical distribution of the motion rotation information to optimize the sample balance and improve the model training effect, obtaining the preprocessed inertial data.

[0069] Step (2) specifically includes the following steps:

[0070] (2a) Concatenate two consecutive frames of preprocessed visual images I t and I t+1 along the channel dimension to obtain a composite image V t :

[0071]

[0072] where cat represents the image concatenation operation, and H and W are the height and width of the image respectively;

[0073] (2b) Input the composite image V t into a deep convolutional neural network composed of multiple layers of convolution, batch normalization, and LeakyReLU activation function for visual feature extraction. The output feature map of the l-th layer satisfies:

[0074]

[0075] where the initial feature is the composite image V t , Conv (l) represents the l-th layer convolution operation, BN represents the batch normalization operation, and σ(·) is the LeakyReLU activation function; After being processed by the deep convolutional neural network, the final output feature map is denoted as where C represents the number of channels, and H' and W' represent the height and width of the final output feature map respectively;

[0076] (2c) Integrate a dual attention module in the deep convolutional neural network to enhance the expression ability of visual features. The dual attention module includes a local-global information fusion module and a channel recalibration module;

[0077] Input the final output feature map F t into the local-global information fusion module. The local-global information fusion module processes F t using three parallel branches:

[0078] The first branch uses a 1×1 convolution to achieve channel remapping, and the output feature is denoted as

[0079] The second branch uses a 3×3 convolution to capture local spatial information, and the output feature is denoted as

[0080] After the third branch performs adaptive average pooling, it is then passed through a 1×1 convolution and Sigmoid activation to generate the global attention map A global :

[0081]

[0082] where, W g is a learnable parameter matrix, |Ω| represents the area of the spatial region of the feature map, and x is the spatial coordinate;

[0083] Subsequently, multiply A global element-wise with F t and then add the output feature output feature A global ⊙F t and apply the SiLU activation function to the result to obtain the fused feature map F' t :

[0084]

[0085] where, ⊙ represents element-wise multiplication;

[0086] The processing of the channel recalibration module is as follows:

[0087] Perform global average pooling on the fused feature map F' t and then pass it through a one-dimensional convolution and Sigmoid activation to generate the channel weights denoted as A channel ;

[0088] Multiply A channel element-wise with F' t to obtain the recalibrated feature map

[0089]

[0090] Finally, flatten and map it through a fully connected layer to output the visual feature vector enhanced by the dual attention module d v represents the dimension of the visual feature vector v t of.

[0091] Step (3) specifically includes the following steps:

[0092] (3a) Reconstruct the preprocessed inertial data into a format suitable for one-dimensional convolutional processing;

[0093] (3b) Apply multi-layer one-dimensional convolutional operations to the reconstructed inertial data. Each layer of convolution combines batch normalization and the LeakyReLU activation function to fully capture the temporal dynamic characteristics in the inertial signal;

[0094] (3c) Flatten the feature sequence processed by the multi-layer convolutional layer and perform mapping through a fully connected layer to obtain the inertial feature vector d i represents the dimension of the inertial feature vector i t .

[0095] Step (4) specifically includes the following steps:

[0096] (4a) Construct a composite vector x t to guide the adaptive regulation of the policy network. The composite vector x t is composed of the inertial feature vector i t and the previous moment's LSTM hidden state . d r represents the dimension of the LSTM hidden state. The expression of the composite vector x t is:

[0097]

[0098] (4b) Process x t using the policy network and output the decision vector d t . The decision vector is discretized using the Gumbel-Softmax method to dynamically regulate the weight of the visual feature vector v t ; subsequently, the regulated visual feature vector and the inertial feature vector i t are concatenated to form the final fusion feature f t .

[0099] Step (5) specifically includes the following steps:

[0100] (5a) Input the final fusion feature into the long short-term memory network to capture the temporal correlation, thereby obtaining the LSTM hidden state h t :

[0101] h t = LSTM(f t , h t-1 )

[0102] where is the hidden state at the previous moment; d f is the final fused feature f t dimension;

[0103] (5b) Apply the median Kolmogorov - Arnold network, i.e., the MKAN network, to the LSTM hidden state h t to achieve non - linear feature expansion. The MKAN network is obtained by improving the Kolmogorov - Arnold network. Specifically: Mix the Gaussian radial basis function and the Lorenz function to obtain a mixed basis function, and replace the Gaussian radial basis function of the Kolmogorov - Arnold network with the mixed basis function; the mixed basis function is:

[0104]

[0105] where g is the kernel center; δ is the scale parameter; α ∈ [0, 1] is the learnable weight, and α = 1 is set at the beginning of training and then gradually adjusted adaptively; apply the mixed basis function to each component of h t and form a deep non - linear mapping by stacking L layers of MKAN through a linear mapping based on the spline function, and output the enhanced hidden state

[0106]

[0107] where represents the transformation operation of the l - th layer of MKAN;

[0108] (5c) Map the enhanced hidden state generated by the stacked MKAN layers through the fully - connected regression layer to obtain the final six - degree - of - freedom pose estimation

[0109]

[0110] where and are the weight matrix and bias parameter of the regression layer respectively, and d m is the dimension of the MKAN output.

[0111] MKAN uses a mixed basis function to replace the single basis function. The design basis is that although the Gaussian radial basis function has smoothness and can capture local gradual changes, it is not sensitive enough to motion mutations and outliers; while the Lorenz function has sharp peaks and heavy - tail characteristics and can more sensitively reflect abnormal changes. Therefore, the present invention proposes to fuse the two in a linear combination manner.

[0112] Such as Figure 2As shown, the red curve represents the true trajectory, the blue curve represents the estimation result of the present invention, and the black dot is the starting position; in this sequence, the relative translation error of the present invention is 2.80%, and the relative rotation error is 0.54°.

[0113] In summary, through the adaptive regulation mechanism of visual-inertial data, the present invention retains key motion features while compressing redundant information; innovatively introduces a dual attention module to enhance feature discrimination ability, and combines the mixed basis function modeling ability of the median Kolmogorov-Arnold network to achieve accurate characterization of complex non-linear motion dynamics, and finally achieves stable and high-precision six-degree-of-freedom pose estimation.

[0114] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection required by the present invention is defined by the appended claims and their equivalents.

Claims

1. A six-degree-of-freedom visual-inertial odometry method based on dual attention and mixed basis functions, characterized by: The method comprises the following steps in order: (1) Acquisition and preprocessing: Using a device equipped with a monocular camera and an inertial measurement unit, visual image data and inertial data are collected in a target motion environment, and the visual image data and inertial data are preprocessed to obtain preprocessed visual image data and preprocessed inertial data; (2) Visual feature extraction: A deep convolutional neural network is used to extract multi-scale features from the preprocessed visual image data, and a dual attention module is introduced to obtain the visual feature vector v by fusing local details with global context information. t ; (3) Inertial feature extraction: A one-dimensional convolutional neural network is used to perform time series modeling on the preprocessed inertial data to extract the inertial feature vector i t ; (4) Adaptive data fusion: Using the strategy network, the inertial feature vector i t The LSTM hidden state at the previous moment is used as the input of the policy network to the visual feature vector v t Perform adaptive gating and dynamic fusion to generate the final fusion feature f t ; (5) Time series modeling and pose estimation: The final fusion feature f t The long short-term memory network is input for time series modeling, and the nonlinear features of the output of the long short-term memory network are enhanced using the median Kolmogorov-Arnold network combined with the mixed basis function. The six-degree-of-freedom pose estimation in the target environment is output through the fully connected regression layer.

2. The six-degree-of-freedom visual-inertial odometry method based on dual attention and mixed basis function according to claim 1, characterized in that: Step (1) specifically refers to: the visual image data is collected by a high-resolution monocular camera with an image frame rate of not less than 10 Hz, the visual image data is preprocessed, that is, grayscale, denoising and geometric correction are performed to enhance the image quality to obtain preprocessed visual image data; the inertial data is collected by an inertial measurement unit at a fixed sampling frequency, the inertial data is preprocessed, that is, low-pass filtering and time synchronization are performed, and then weighted according to the statistical distribution of motion rotation information to obtain preprocessed inertial data.

3. The six-degree-of-freedom visual-inertial odometer method based on dual attention and mixed basis function according to claim 1, characterized in that: Step (2) specifically includes the following steps: (2a) The two frames of preprocessed visual images I t and I t+1 Splice along the channel dimension to get the composite image V t : Among them, cat represents the image stitching operation, H and W are the height and width of the image respectively; (2b) The composite image V t The input is a deep convolutional neural network composed of multi-layer convolution, batch normalization and LeakyReLU activation function for visual feature extraction. The lth layer outputs the feature map satisfy: Among them, the initial features is the composite image V t , Conv (l) represents the l-th layer convolution operation, BN represents the batch normalization operation, and σ(·) is the LeakyReLU activation function. After being processed by the deep convolutional neural network, the final output feature map is recorded as Among them, C represents the number of channels, H' and W' represent the height and width of the final output feature map respectively; (2c) Integrating a dual attention module into a deep convolutional neural network to improve the expressiveness of visual features. The dual attention module includes a local-global information fusion module and a channel recalibration module. The final output feature map F t Input local-global information fusion module, local-global information fusion module uses three parallel branches to t To process: The first branch uses 1×1 convolution to achieve channel remapping, and the output feature is recorded as The second branch uses 3×3 convolution to capture local spatial information, and the output feature is recorded as The third branch performs adaptive average pooling, and then generates a global attention map A through 1×1 convolution and Sigmoid activation. global : Among them, W g is the learnable parameter matrix, |Ω| represents the spatial area of ​​the feature map, and x is the spatial coordinate; Then, A global With F t Multiply element by element and then output features Output Features A global ⊙F t Add them together and apply SiLU activation function to the result to get the fusion feature map F' t : Among them, ⊙ represents element-by-element multiplication; The channel recalibration module processes as follows: For the fusion feature map F' t Perform global average pooling, and then generate channel weights through one-dimensional convolution and Sigmoid activation, denoted as A channel ; A channel With F' t Multiply element by element to get the recalibrated feature map Finally, Flattened and mapped through a fully connected layer, outputting a visual feature vector enhanced by a dual attention module d v Represents the visual feature vector v t Dimension.

4. The six-degree-of-freedom visual-inertial odometer method based on dual attention and mixed basis function according to claim 1, characterized in that: Step (3) specifically includes the following steps: (3a) The preprocessed inertial data Reconstruct into a format suitable for one-dimensional convolution processing; (3b) Multi-layer one-dimensional convolution operations are applied to the reconstructed inertial data, and each convolution layer is combined with batch normalization and LeakyReLU activation function to fully capture the temporal dynamic characteristics of the inertial signal; (3c) Flatten the feature sequence after multi-layer convolutional layer processing and map it through the fully connected layer to obtain the inertia feature vector d i represents the inertial eigenvector i t Dimension.

5. The six-degree-of-freedom visual-inertial odometer method based on dual attention and mixed basis function according to claim 1, characterized in that: Step (4) specifically includes the following steps: (4a) Construct a composite vector x t To guide the adaptive regulation of the policy network, the composite vector x t By the inertia eigenvector i t and the LSTM hidden state at the previous moment Combination, d r Represents the dimension of the LSTM hidden state, the composite vector x t The expression is: (4b) Using the policy network to t Process and output the decision vector d t , the decision vector is discretized using the Gumbel-Softmax method to dynamically control the visual feature vector v t The weight of the regulated visual feature vector is then combined with the inertial feature vector i t Splice to form the final fusion feature f t .

6. The six-degree-of-freedom visual-inertial odometer method based on dual attention and mixed basis function according to claim 1, characterized in that: Step (5) specifically includes the following steps: (5a) The final fusion feature Input the long short-term memory network to capture the temporal correlation, thus obtaining the LSTM hidden state h t : h t =LSTM(f t ,h t-1 ) in, is the hidden state at the previous moment; d f is the final fusion feature f t Dimensions; (5b) For the LSTM hidden state h t The median Kolmogorov-Arnold network, i.e., the MKAN network, is applied to realize nonlinear feature expansion. The MKAN network is obtained by improving the Kolmogorov-Arnold network. Specifically, the Gaussian radial basis function and the Lorentz function are mixed to obtain a mixed basis function, and the mixed basis function replaces the Gaussian radial basis function of the Kolmogorov-Arnold network; the mixed basis function is: Where g is the core center; δ is the scale parameter; α∈[0,1] is the learnable weight, and α=1 is set at the beginning of training and then gradually adjusted adaptively; t Each component of is applied with a mixed basis function, and through a linear mapping based on a spline function, a deep nonlinear mapping is formed after stacking L layers of MKAN, and the output is an enhanced hidden state in, represents the conversion operation of the l-th layer MKAN; (5c) The enhanced hidden state generated by the stacked MKAN layer The final six-degree-of-freedom pose estimation is obtained by mapping through the fully connected regression layer in, and are the weight matrix and bias parameter of the regression layer, d m is the dimension of MKAN output.

7. An electronic device comprising: processor; as well as A memory, in which computer program instructions are stored, and when the computer program instructions are executed by the processor, the processor executes the six-degree-of-freedom visual-inertial odometer method based on dual attention and mixed basis functions as described in any one of claims 1-6.

8. A computer-readable storage medium having computer program instructions stored thereon, wherein when the computer program instructions are executed by a processor, the processor executes the six-degree-of-freedom visual-inertial odometer method based on dual attention and mixed basis functions as described in any one of claims 1-6.

Citation Information

Cited By

  • Radar, inertia and vision fusion navigation method of inertia-guided attention mechanism

    CN121252769A

  • inertial radar, inertial, vision fusion navigation method

    CN121252769B