Visual inertia mileage pose positioning method based on cross attention and dynamic weight
Through the visual inertial odometer positioning method of cross attention and dynamic weights, the problem of insufficient utilization of modal features in the visual inertial odometer is solved, and higher positioning accuracy and stability are achieved, adapting to complex scenarios and maintaining good performance when data is damaged.
Patent Information
- Application Number
- CN202510507849.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-25
AI Technical Summary
The existing visual inertial odometry method fails to fully utilize the complementarity between visual image features and inertial measurement features, and cannot adaptively adjust the application weights of different modal features, resulting in low positioning accuracy and insufficient stability and reliability.
The visual inertial mileage positioning method based on cross attention and dynamic weights is adopted to establish a cross-modal correlation between visual optical flow characteristics and inertial variable characteristics through the cross-attention mechanism, and dynamic balance between different modal features is achieved through the dynamic weight fusion mechanism, improving the positioning accuracy and stability of positioning.
It significantly improves the accuracy and stability of positioning of visual inertial mileage pose information, can adapt to complex scenarios, show good robustness and stability, especially in data corruption scenarios.
Smart Images

Figure CN120368966A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision, and particularly relates to a visual-inertial odometry pose positioning method based on cross-attention and dynamic weights. Background Art
[0002] Visual Odometry (VO) is an important technology in the fields of computer vision and robotics, and is used to estimate the motion trajectory of a camera. As a positioning technology based on a monocular or binocular camera, VO realizes pose estimation through feature matching and motion estimation between consecutive images. Due to the lack of scale information and being restricted by environmental factors such as illumination, texture, and occlusion, pure vision systems have significant limitations in complex scenarios. The Inertial Measurement Unit (IMU) can provide high-frequency angular velocity and acceleration measurements through three-axis gyroscopes and accelerometers, has a stronger ability to obtain information on fast motion states, and its sensing ability is not easily affected by external environmental factors.
[0003] Visual-Inertial Odometry (VIO) is a technology that combines visual and inertial sensor data to estimate the position and orientation of a device or vehicle in real time. It is commonly used in fields such as robotics, unmanned aerial vehicles, and autonomous vehicles. Visual-inertial odometry uses a camera to capture images of the environment, and an Inertial Measurement Unit (IMU) to measure acceleration and rotation. By fusing these two sets of data, the system can better estimate pose information such as the position and orientation of the device or vehicle in three-dimensional space. In VIO, by utilizing the complementarity of visual image information and inertial measurement data, visual image information is used to provide long-term pose correction, reducing the divergence and cumulative measurement errors caused by the zero bias of the IMU; the inertial measurement data, based on the gravity vector, provides physical scale information for the system. However, currently, VIO based on deep learning generally realizes feature fusion by directly splicing visual and inertial features, and fails to fully utilize the implicit information in these two modalities. It not only cannot fully exert the complementarity of the two types of information, but also cannot adaptively and dynamically adjust the application weights of the two different modality features according to different scenarios. These factors have all resulted in the low pose positioning accuracy of current VIO, and insufficient stability and reliability in multiple scenarios. Summary of the Invention
[0004] Aiming at the deficiencies of the above-mentioned prior art, the present invention provides a visual-inertial odometry pose positioning method based on cross-attention and dynamic weights, which enhances the cross-strengthening and dynamic coupling between visual image features and inertial measurement features, thereby further improving the pose positioning accuracy of the visual-inertial odometer and enhancing the pose positioning stability and reliability for different scenarios.
[0005] To solve the above technical problems, the present invention adopts the following technical solutions:
[0006] A visual-inertial odometry pose localization method based on cross-attention and dynamic weights, which acquires visual image data and inertial measurement data of a moving device and inputs them into a pre-trained visual-inertial odometry pose prediction model to predict the pose information of the moving device;
[0007] The visual-inertial odometry pose prediction model includes a feature extraction layer, a cross-attention feature fusion layer, and a pose estimation layer; the feature extraction layer is used to extract features from the input visual image data and inertial measurement data respectively to obtain visual optical flow features and inertial variable features; the cross-attention feature fusion layer is used to perform two-way modality interaction enhancement processing on the visual optical flow features and inertial variable features through a cross-attention mechanism and perform dynamic weight fusion to obtain visual-inertial enhanced fusion features; the pose estimation layer is used to perform pose regression processing based on the visual-inertial enhanced fusion features to obtain a pose information prediction result, which is used as the output of the visual-inertial odometry pose prediction model.
[0008] As a preferred solution, the visual image data includes two adjacent frames of visual images, and the inertial measurement data includes inertial time-series data corresponding to the time period of the two adjacent frames of visual images.
[0009] As a preferred solution, in the visual-inertial odometry pose prediction model, the feature extraction layer includes a visual feature encoder and an inertial feature encoder;
[0010] The visual feature encoder is used to stack and splice the input two adjacent frames of visual images in the channel dimension, and then perform multi-scale downsampling feature extraction through multiple convolutional layers, and then output to obtain visual optical flow features;
[0011] The inertial feature encoder is used to perform high-dimensional feature encoding processing on the inertial time-series data corresponding to the time period of the two adjacent frames of visual images input through multiple inertial encoding units, and then output to obtain inertial variable features; each inertial encoding unit includes a one-dimensional convolutional module, a batch normalization module, a LeakyReLU activation function, and a Dropout module connected in series.
[0012] As a preferred solution, in the visual-inertial odometry pose prediction model, the cross-attention feature fusion layer includes a bi-directional cross-attention module and a gated fusion module connected in series;
[0013] The cross-attention module is used to perform two-way modality interaction enhancement processing on the visual optical flow features and inertial variable features through a cross-attention mechanism to obtain visual enhancement features and inertial enhancement features respectively;
[0014] The gating fusion module is used to dynamically select fusion weights through a gating mechanism, and perform dynamic weight fusion on the visual enhancement feature and the inertial enhancement feature to obtain a visual-inertial enhancement fusion feature.
[0015] As a preferred solution, the bidirectional cross-attention module includes a visual attention branch and an inertial attention branch; the visual attention branch first decomposes the visual optical flow feature X v into a visual key vector K v , a visual value vector V v , and a visual query vector Q v , and cross-transmits the visual query vector Q v to the inertial attention branch; the inertial attention branch first decomposes the inertial variable feature X i into an inertial key vector K i , an inertial value vector V i , and an inertial query vector Q i , and cross-transmits the inertial query vector Q i to the visual attention branch; then, in the visual attention branch, the visual key vector K v , the visual value vector V v , and the inertial query vector Q i are subjected to attention feature extraction, linearly transformed, then added to the input visual optical flow feature X v , and after feature dimension conversion through a multi-layer perceptron model MLP, the visual enhancement feature as the output is obtained; in the inertial attention branch, the inertial key vector K i , the inertial value vector V i , and the visual query vector Q v are subjected to attention feature extraction, linearly transformed, then added to the input inertial variable feature X i , and after feature dimension conversion through a multi-layer perceptron model MLP, the inertial enhancement feature as the output is obtained;
[0016] The expressions of the visual enhancement feature and the inertial enhancement feature are:
[0017]
[0018] Q v = W Q ·X v , K v = W K ·X v , V v = W V ·X v ;
[0019] Qi = W Q ·X i , K i = W K ·X i , V i = W V ·X i ;
[0020] Among them, and respectively represent the output visual enhancement feature and inertial enhancement feature; MLP(·) represents the processing of the multi-layer perceptron MLP; Attention(·) represents the attention feature extraction operation; represents the concatenation addition operation; Softmax(·) represents the Softmax function operation; is the attention ratio factor; T represents the transpose symbol; W Q , W K , W V , W a are all linear weight parameters of the bidirectional cross-attention module.
[0021] As a preferred solution, the gating fusion module divides the visual enhancement feature and the inertial enhancement feature into the same number of groups of visual enhancement feature components and groups of inertial enhancement feature components respectively, and then passes each visual enhancement feature component and inertial enhancement feature component through a convolutional layer to obtain their respective initial fusion weights, multiplies them weighted with their respective feature components, then passes through a Softmax layer for normalization processing to obtain their respective normalized fusion weights, multiplies them weighted with their respective feature components again to obtain their respective corresponding weighted feature components, and finally concatenates and fuses the groups of visual enhancement weighted feature components and the groups of inertial enhancement weighted feature components to obtain the visual-inertial enhancement fusion feature as the output;
[0022] The expression of the visual-inertial enhancement fusion feature is:
[0023] F fusion = Concat(F v , F i );
[0024]
[0025]
[0026] Among them, F fusion represents the obtained visual-inertial enhancement fusion feature, F v represents the visual enhancement weighted feature composed of the connection of groups of visual enhancement weighted feature components, F idenotes the inertia-enhanced weighted feature formed by connecting each group of inertia-enhanced weighted feature components; Concat(·) represents the feature concatenation operation; denotes the vision-enhanced feature the nth vision-enhanced feature component after being divided, denotes the inertia-enhanced feature the nth inertia-enhanced feature component after being divided, n = 1, 2, …, N, where N represents the total number of divided feature components; R v,n 、S v,n respectively denote the initial fusion weight and the normalized fusion weight of the nth vision-enhanced feature component ; R i,n 、S i,n respectively denote the initial fusion weight and the normalized fusion weight of the nth inertia-enhanced feature component ; Softmax(·) represents the Softmax function operation; e is the natural constant; denotes the feature concatenation operation for feature components n = 1, 2, …, N.
[0027] As a preferred solution, in the visual-inertial odometry pose prediction model, the processing process of the pose estimation layer is as follows:
[0028] Input the visual-inertial enhanced fusion feature into a two-layer LSTM network to extract the temporal information in the visual-inertial enhanced fusion feature. Among them, after each layer of LSTM processing, a multi-layer perceptron MLP is passed through for feature dimension conversion, and the output of the LSTM network is then regressively predicted through a fully connected layer to obtain the pose information prediction result; the pose information prediction result is a 6-dimensional vector, including a 3-dimensional translation vector and a 3-dimensional rotation vector
[0029] As a preferred solution, the visual-inertial odometry pose prediction model is trained in the following manner:
[0030] S101: Prepare a sample data set containing visual image data and inertial measurement data, which is divided into a training data set and a test data set. The visual image data and inertial measurement data in the sample data set are both pre-marked with the true values of the translation vector and the rotation vector;
[0031] S102: Input the training data set into the visual-inertial odometry pose prediction model for training, and optimize the parameters of the visual-inertial odometry pose prediction model with the goal of minimizing the loss function until the visual-inertial odometry pose prediction model converges, obtaining the trained visual-inertial odometry pose prediction model;
[0032] S103: Test the visual-inertial odometry pose prediction model using a test data set to confirm the pose information prediction and positioning performance of the trained visual-inertial odometry pose prediction model; if the pose information prediction and positioning performance meets the requirements, end the training of the visual-inertial odometry pose prediction model; otherwise, return to step S102.
[0033] As a preferred solution, the loss function used for training the visual-inertial odometry pose prediction model is the mean squared error loss function, and the mean squared error loss function L pose has the following expression:
[0034]
[0035] where T v is the total number of frames of the training image sequence; v t respectively represent the predicted value and the true value of the translation vector; φ t respectively represent the predicted value and the true value of the rotation vector; ||·||2 represents the L2 norm operation; α is the loss weight parameter.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] 1. The method of the present invention uses a visual-inertial odometry pose prediction model to perform positioning prediction of visual-inertial odometry pose information. In this visual-inertial odometry pose prediction model, a cross-attention mechanism is used to establish a cross-modal feature correlation between visual optical flow features and inertial variable features, thereby enhancing the attention to the spatial regions in the visual optical flow features that are highly correlated with inertial variable information; at the same time, with the help of a dynamic weight fusion mechanism, a dynamic balance between different modal features is achieved through a learnable weight allocation strategy, helping to avoid getting stuck in the local optimal solution problem, and significantly improving the accuracy of visual-inertial odometry pose information positioning prediction and the stability and reliability for different scenarios.
[0038] 2. In the method of the present invention, the introduction of the cross-attention mechanism enables the visual-inertial odometry pose prediction model network to flexibly adjust the degree of attention to image regions according to the dynamic state of the moving device, thereby improving the adaptability of the visual-inertial odometry pose prediction model to complex scenarios. At the same time, the introduction of the dynamic weight fusion mechanism can dynamically adjust the weights of visual and inertial features according to the current state of the moving device, thereby achieving feature balance in different scenarios. Therefore, the method of the present invention can more accurately reflect the motion trajectory of the moving device, can better estimate the pose of the current state when facing curved driving and high-altitude driving, and has better adaptability to different dynamic scenarios.
[0039] 3. The method of the present invention has good performance in data corruption scenarios. Even when there is a certain degree of data corruption, the drift phenomenon of the method of the present invention remains at a low level, and it can still maintain good pose estimation performance, showing stronger robustness and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to make the objectives, technical solutions, and advantages of the invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings, where:
[0041] Figure 1 is a schematic diagram of the architecture of the visual-inertial odometry pose prediction model in the method of the present invention;
[0042] Figure 2 is a schematic diagram of the structure of the FlowNetS network;
[0043] Figure 3 is a schematic diagram of the structure of the inertial feature encoder;
[0044] Figure 4 is a schematic diagram of the structure of the bidirectional cross-attention module;
[0045] Figure 5 is a schematic diagram of the structure of the gated fusion module;
[0046] Figure 6 is a schematic diagram of the architecture of the KITTI dataset acquisition platform in the embodiment;
[0047] Figure 7 is the visualization result of the odometry of the KITTI dataset sequence 10 in the embodiment;
[0048] Figure 8 is the visualization result of the odometry of the KITTI dataset sequence 09 in the embodiment;
[0049] Figure 9 is the ablation experiment comparison diagram of the KITTI dataset sequence 09 in the embodiment;
[0050] Figure 10 is the attention map of the right turn (right) and left turn (left) scenarios in the embodiment;
[0051] Figure 11 is the speed map (left) and feature weight map (right) of the KITTI dataset sequence 09 in the embodiment;
[0052] Figure 12 is the speed map (left) and feature weight map (right) of the KITTI dataset sequence 10 in the embodiment;
[0053] Figure 13 is the experimental result of the data corruption scenario of the KITTI dataset sequences 09 and 10 in the embodiment. Detailed implementation manners
[0054] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the accompanying drawings herein may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0055] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0056] The present invention provides a visual-inertial odometry pose positioning method based on cross-attention and dynamic weights. The method obtains visual image data and inertial measurement data of a moving device and inputs them into a pre-trained visual-inertial odometry pose prediction model to predict the pose information of the moving device.
[0057] The visual-inertial odometry pose prediction model framework proposed by the present invention mainly consists of three parts, namely: (1) a feature extraction layer for extracting visual features and inertial features; (2) a cross-attention feature fusion layer for enhancing and fusing the two features; and (3) a pose estimation layer for pose regression processing.
[0058] Figure 1 shows the framework structure of the visual-inertial odometry pose prediction model proposed by the present invention. As Figure 1As shown in the figure, when performing visual-inertial odometry pose positioning, the obtained visual image data of the moving device includes two adjacent frames of visual images, and the obtained inertial measurement data of the moving device includes the inertial time-series data corresponding to the time period of the two adjacent frames of visual images. The visual-inertial odometry pose prediction model proposed by the present invention includes a feature extraction layer, a cross-attention feature fusion layer, and a pose estimation layer. Among them, the feature extraction layer is used to perform feature extraction on the input visual image data and inertial measurement data respectively to obtain visual optical flow features and inertial variable features. The cross-attention feature fusion layer is used to perform two-way modality interaction enhancement processing on the visual optical flow features and inertial variable features through a cross-attention mechanism, and perform dynamic weight fusion to obtain visual-inertial enhanced fusion features. The pose estimation layer is used to perform pose regression processing based on the visual-inertial enhanced fusion features to obtain a pose information prediction result, which is used as the output of the visual-inertial odometry pose prediction model.
[0059] The method of the present invention uses a visual-inertial odometry pose prediction model to perform positioning prediction of visual-inertial odometry pose information. In this visual-inertial odometry pose prediction model, a cross-attention mechanism is used to establish a cross-modal feature correlation between visual optical flow features and inertial variable features, thereby enhancing the attention to the spatial regions in the visual optical flow features that are highly correlated with inertial variable information. At the same time, with the help of a dynamic weight fusion mechanism, a dynamic balance between different modality features is achieved through a learnable weight allocation strategy, which helps to avoid falling into the local optimal solution problem, and significantly improves the accuracy of visual-inertial odometry pose information positioning prediction and the stability and reliability for different scenarios.
[0060] The visual-inertial odometry pose positioning method based on cross-attention and dynamic weight of the present invention and the used visual-inertial odometry pose prediction model will be described in more detail below.
[0061] 1 Feature extraction layer
[0062] In the visual-inertial odometry pose prediction model, the feature extraction layer includes a visual feature encoder and an inertial feature encoder, which are respectively used to perform feature extraction on visual image data and inertial measurement data.
[0063] (1.1) Visual feature encoder
[0064] In the odometry task, feature extraction is one of the core links to achieve accurate motion estimation. The visual feature encoder needs to estimate the motion pattern of pixel points in the image sequence, that is, calculate the optical flow. By analyzing the pixel displacement between consecutive image frames, key motion information can be provided for the odometry. In the present invention, the visual feature encoder is used to stack and splice the input two adjacent frames of visual images in the channel dimension, and then perform multi-scale downsampling feature extraction through multiple convolutional layers, and finally output the visual optical flow features.
[0065] As a specific implementation, FlowNetS (Simple) can be selected as the feature extraction network. The network structure of FlowNetS is as Figure 2 shown. In the input processing stage, two adjacent frames of visual images are directly concatenated along the channel dimension to form a 6-channel input (with a size of H×W×6, where H and W are the pixel height and width of the image respectively). The entire FlowNetS encoder consists of 9 convolutional layers, labeled as conv1, conv2, conv3, conv3_1, conv4, conv4_1, conv5, and conv5_1, which are used to gradually extract multi-scale features; among them, the first 3 convolutional layers use larger convolutional kernels (7×7 and 5×5) to capture global motion patterns, and the subsequent convolutional layers use a convolutional kernel size of 7×7; among them, the convolutional strides of conv3_1, conv4_1, and conv5_1 are 1, and the strides of other convolutional layers are all 2; then, optical flow features are generated in the output layer, and the resolution of the resulting visual optical flow features is 1 / 32 of the original video image. Optical flow is a vector that describes the pixel motion between adjacent image frames, and the goal is to estimate pixel-level motion information from two adjacent frames of visual images, that is, to predict the translational motion information contained in the motion vector (corresponding pixels in the image) between two adjacent frames of visual images.
[0066] The main advantage of FlowNetS lies in its ability to directly process image pairs, thus avoiding the complex dual-branch feature extraction and correlation calculation layers in FlowNet-C. This structural simplification not only reduces the number of parameters of FlowNetS by about 30%, but also significantly improves its performance in high dynamic range and large displacement scenarios. In addition, FlowNetS can directly learn the mapping relationship from raw pixels to the optical flow field, getting rid of the dependence on manually designed features. This feature makes it show higher stability in complex lighting changes and weak texture areas (such as walls and skies). Even when only trained with synthetic datasets (such as Flying Chairs), FlowNetS can still be effectively transferred to real scenes (such as the KITTI dataset), demonstrating good generalization ability. Therefore, the present invention selects this network structure as the visual feature encoder.
[0067] (1.2) Inertial Feature Encoder
[0068] In the VIO system, the feature extraction of IMU data for inertial sensors is crucial for improving the accuracy and robustness of the system. Since inertial navigation is essentially a non - linear problem, deep learning methods have significant advantages in dealing with such problems. Deep learning models can automatically learn complex non - linear relationships, while traditional methods usually rely on complex geometric and kinematic models, which are often difficult to fully match the real situation in practical applications. Deep learning methods achieve high - precision navigation through adaptive training and are particularly excellent in dealing with non - linear problems.
[0069] In the present invention, the inertial feature encoder is used to perform high - dimensional feature encoding processing on the inertial time - series data corresponding to the adjacent two frames of visual images in sequence through a plurality of inertial encoding units, and then output the inertial variable features; each inertial encoding unit includes a one - dimensional convolution module, a batch normalization module, a LeakyReLU activation function, and a Dropout module cascaded in sequence. Figure 3 An example structure of the inertial feature encoder used in the present invention is shown, which includes three cascaded inertial encoding units for encoding the input inertial time - series data into high - dimensional features. Each feature dimension of the input inertial time - series data corresponds to 3 accelerometer channels and 3 gyroscope channels of the IMU. However, generally, the acquisition frequency of the IMU and the acquisition frequency of the visual image data frames are not necessarily the same. For example, the acquisition frequency of the visual image data is 10Hz, and the acquisition frequency of the inertial data of the IMU may be 100Hz. Therefore, it is necessary to align the timing of the inertial data with the timing of the adjacent two frames of visual image data, find the inertial time - series data corresponding to the adjacent two frames of visual images, perform inertial feature encoding through the inertial sensor, calculate the difference change amount between the inertial data in the corresponding period, and obtain the inertial variable features.
[0070] 2 Cross - attention Feature Fusion Layer
[0071] In the visual - inertial odometry pose prediction model, the cross - attention feature fusion layer includes a bidirectional cross - attention module and a gated fusion module connected in series; the cross - attention module is used to perform two - way modal interaction enhancement processing on the visual optical flow features and the inertial variable features through the cross - attention mechanism, and respectively obtain visual enhanced features and inertial enhanced features; the gated fusion module is used to dynamically select fusion weights through the gated mechanism and perform dynamic weight fusion on the visual enhanced features and the inertial enhanced features to obtain visual - inertial enhanced fusion features.
[0072] (2.1) Bidirectional Cross - attention Module
[0073] Compared with visual optical flow features, inertial variable features have more intuitive prior information. Different from visual odometry which needs to extract features first and then judge the motion state, the data from inertial sensors can directly judge the approximate current motion state of the vehicle. Therefore, the prior information contained in the inertial variable features can be used to assist in the extraction of deep information in the visual optical flow features. In the present invention, the bidirectional cross-attention module is designed for this purpose, and the feature representation is enhanced by calculating the correlation between the visual optical flow features and the inertial variable features.
[0074] As Figure 4 shown, the bidirectional cross-attention module (BCAM, Bidirectional CrossAttention Module) proposed in the present invention includes two branches, namely the visual attention branch and the inertial attention branch. In the two branches, the visual optical flow features and the inertial variable features will respectively generate the key vector K, the value vector V, and the query vector Q through the linear layer. The query vector Q is then cross-transmitted to the other branch, and then the attention-weighted feature representation is obtained through calculation and added to the original feature, and then converted to the final feature dimension through the MLP. The specific processing process of the bidirectional cross-attention module is as follows:
[0075] The visual attention branch first decomposes the visual optical flow feature X v through the linear layer to generate the visual key vector K v , the visual value vector V v , and the visual query vector Q v , and cross-transmits the visual query vector Q v to the inertial attention branch; the inertial attention branch first decomposes the inertial variable feature X i through the linear layer to generate the inertial key vector K i , the inertial value vector V i , and the inertial query vector Q i , and cross-transmits the inertial query vector Q i to the visual attention branch; then, in the visual attention branch, the visual key vector K v , the visual value vector V v , and the inertial query vector Q i are subjected to attention feature extraction and then linearly transformed, and then added to the input visual optical flow feature X v , and after feature dimension conversion through the multi-layer perceptron model MLP, the visual enhanced feature as the output is obtained; in the inertial attention branch, the inertial key vector K i , the inertial value vector V i , and the visual query vector Q v are subjected to attention feature extraction and then linearly transformed, and then added to the input inertial variable feature X iAfter adding them up and performing feature dimension conversion through a multi-layer perceptron model MLP, inertial enhanced features are obtained as the output.
[0076] The expressions for the visual enhanced features and the inertial enhanced features are:
[0077]
[0078] Q v =W Q ·X v ,K v =W K ·X v ,V v =W V ·X v ;
[0079] Q i =W Q ·X i ,K i =W K ·X i ,V i =W V ·X i ;
[0080] Among them, and respectively represent the output visual enhanced features and inertial enhanced features; MLP(·) represents the processing by the multi-layer perceptron MLP; Attention(·) represents the attention feature extraction operation; represents the concatenation and addition operation; Softmax(·) represents the Softmax function operation; is the attention ratio factor; T represents the transpose symbol; W Q 、W K 、W V 、W a are all linear weight parameters of the bidirectional cross-attention module, and the values of these linear weight parameters are all optimized and determined during the training process of the visual-inertial odometry pose prediction model.
[0081] Through bidirectional modal interaction, the bidirectional cross-attention module realizes the deep complementarity between vision and inertia. Visual features help enhance the motion sensitivity of inertial data, and inertial features help improve the robustness of visual features, thereby improving the odometry performance in complex motion environments.
[0082] (2.2) Gated Fusion Module
[0083] In actual scenarios, the network should focus on different features under different circumstances. Inspired by the multi-gate mixture-of-experts model in multi-task learning, the present invention designs a gating fusion module, which selects the weight outputs of different modalities through a gating mechanism to enhance the adaptability to inertial features and visual features. The traditional single-gate mixture-of-experts model increases the flexibility of the model and allows each task to select the most suitable expert network according to its own needs. However, the single gate limits the possible complex dependencies between tasks. Therefore, the gating mechanism designed by the present invention preserves the dependencies between features by determining the input weights of different modalities used to process the current task.
[0084] As Figure 5 shown, in order to make full use of the complementarity of the two features and adjust their contributions to the final pose estimation, the gating fusion module divides the visually enhanced features and the inertial enhanced features into the same number of groups of components, enabling independent calculation of the fusion weights within each group, while reducing the computational amount to keep the model lightweight. A learnable probability gate is designed for each group of components, and the normalized fusion weights corresponding to each group of components are obtained through normalization processing by a convolutional layer and a Softmax layer. Then, they are multiplied and weighted with their respective feature components respectively. Finally, component concatenation and cross-dimensional concatenation operations are performed to obtain the visually inertial enhanced fusion feature. The specific processing process of the gating fusion module is as follows:
[0085] The gating fusion module divides the visually enhanced features and the inertial enhanced features into the same number of groups of visually enhanced feature components and groups of inertial enhanced feature components respectively. Then, each visually enhanced feature component and inertial enhanced feature component are respectively passed through a convolutional layer to obtain their respective initial fusion weights, which are then multiplied and weighted with their respective feature components. After normalization processing through a Softmax layer to obtain their respective normalized fusion weights, they are multiplied and weighted with their respective feature components again to obtain their respective weighted feature components. Finally, the groups of visually enhanced weighted feature components and the groups of inertial enhanced weighted feature components are concatenated and fused to obtain the visually inertial enhanced fusion feature as the output.
[0086] The expression of the visually inertial enhanced fusion feature is:
[0087] F fusion = Concat(F v , F i );
[0088]
[0089]
[0090] Among them, F fusion represents the obtained visually inertial enhanced fusion feature, and F vDenote the visually enhanced weighted feature formed by connecting each group of visually enhanced weighted feature components as F i Denote the inertia enhanced weighted feature formed by connecting each group of inertia enhanced weighted feature components; Concat(·) represents the feature concatenation operation; Denote the visually enhanced feature The nth visually enhanced feature component after division, Denote the inertia enhanced feature The nth inertia enhanced feature component after division, n = 1, 2, …, N, where N represents the total number of divided feature components; R v,n 、S v,n Denote the initial fusion weight and the normalized fusion weight of the nth visually enhanced feature component respectively; R i,n 、S i,n Denote the initial fusion weight and the normalized fusion weight of the nth inertia enhanced feature component respectively; Softmax(·) represents the Softmax function operation; e is the natural constant; Denote the feature concatenation operation for feature components n = 1, 2, …, N.
[0091] 3 Pose estimation layer
[0092] In the visual-inertial odometry pose prediction model, the processing process of the pose estimation layer is as follows: input the visually-inertially enhanced fusion feature into a two-layer LSTM network to extract the temporal information in the visually-inertially enhanced fusion feature. Among them, after each layer of LSTM processing, a multi-layer perceptron MLP is passed through for feature dimension conversion, and the output of the LSTM network is then regressively predicted through a fully connected layer to obtain the pose information prediction result; the pose information prediction result is a 6-dimensional vector, including a 3-dimensional translation vector and a 3-dimensional rotation vector
[0093] 4 Training of the visual-inertial odometry pose prediction model
[0094] The visual-inertial odometry pose prediction model of the present invention is trained in the following manner:
[0095] S101: Prepare a sample data set containing visual image data and inertial measurement data, which is divided into a training data set and a test data set. The visual image data and inertial measurement data in the sample data set are both pre-marked with the true values of the translation vector and the rotation vector;
[0096] S102: Input the training data set into the visual-inertial odometry pose prediction model for training, and optimize the parameters of the visual-inertial odometry pose prediction model with the goal of minimizing the loss function until the visual-inertial odometry pose prediction model converges, obtaining the trained visual-inertial odometry pose prediction model;
[0097] S103: Test the visual-inertial odometry pose prediction model with the test data set to confirm the pose information prediction and positioning performance of the trained visual-inertial odometry pose prediction model; if the pose information prediction and positioning performance meets the requirements, end the training of the visual-inertial odometry pose prediction model; otherwise, return to step S102.
[0098] During the training process, the method of the present invention calculates the loss using the mean square error. The mean square error loss function L pose of attitude estimation is expressed as follows:
[0099]
[0100] where, T v is the total number of frames of the training image sequence; v t respectively represent the predicted value and the true value of the translation vector; φ t respectively represent the predicted value and the true value of the rotation vector; ||·||2 represents the L2 norm operation; α is the loss weight parameter, which is a preset value. This loss function is used to optimize the estimation of the translation vector v t and the rotation vector φ t .
[0101] 5 Verification Example
[0102] For those skilled in the art to better understand the effect of this method, the following verification experiment is carried out to further illustrate the present invention.
[0103] (5.1) Public Data Set
[0104] This experiment uses the publicly available KITTI dataset. As a widely recognized computer vision algorithm evaluation tool in the field of autonomous driving, this dataset was jointly established by the Karlsruhe Institute of Technology (KIT) in Germany and the Toyota Technological Institute at Chicago (TTIC) in the United States. This dataset has received great attention from the international academic and industrial communities due to its applications in various computer vision tasks, including stereo vision, optical flow computation, visual odometry, 3D object detection, and 3D tracking, etc.
[0105] As Figure 6 shown, the data collection platform of this dataset is equipped with advanced sensor systems, including binocular camera devices, 64-line lidar scanners (Velodyne HDL-64E), and integrated GPS and inertial measurement unit (IMU) navigation systems. The collaborative work of these high-precision sensors ensures the extensiveness and accuracy of data collection. The dataset covers various traffic scenarios such as urban streets, rural roads, and highways, providing rich image materials, where each frame of image may contain up to 15 cars and 30 pedestrians, involving different degrees of occlusion and truncation situations.
[0106] In this experiment, the KITTI Odometry dataset was selected to verify the method of the present invention. This dataset contains 22 binocular video sequences, among which sequences 00 to 10 provide accurate ground truth trajectories for training and verification; while sequences 11 to 22 do not contain ground truth for independent evaluation. Since sequence 03 does not contain inertial data, we selected sequences 00, 01, 02, 04, 05, 06, 07, 08 for model training, and used sequences 09 and 10 for testing. The resolution of the images is 1226×370, and the IMU data is the three-axis angular velocity and three-axis acceleration. The images and pose ground truth are recorded at a frequency of 10Hz, while the IMU data is collected at a frequency of 100Hz. Since there is no strict synchronization between the IMU data and the images, the present invention performs temporal synchronization processing on the original IMU data through an interpolation method to match the sampling frequency of the images and pose ground truth. In this embodiment, mainly monocular images captured by the left camera in the KITTI dataset are used.
[0107] (5.2) Experimental Environment and Parameter Settings
[0108] In the experimental design of the present invention, the Ubuntu 22.04 operating system is used as the basic environment of the experimental platform, and the PyTorch framework is adopted as the core tool for algorithm development. At the same time, the relevant programming work is completed with the help of the Python language. In terms of hardware, the experimental platform is equipped with an Intel i9-14900KF processor and a GTX 3080Super GPU graphics card to meet the computational resource requirements of the experiment.
[0109] In the model training stage, the sizes of all input images are uniformly adjusted to 512×256 pixels. For the image feature extraction task, we select the FlowNetS optical flow feature encoder with the last layer removed as the visual feature encoder, and this encoder has been pre-trained on the Flying Chairs dataset. During the training process, a strategy with a batch size of 14 is adopted, and the maximum number of iterations is limited to 80 times. To efficiently adjust the model parameters, the Adam optimization algorithm is used. The training process is divided into two stages: the initial training stage and the joint training stage. In the initial training stage, the learning rate is set to 5×10 -4 , and 40 iterations of training are carried out, and the gating fusion module is disabled, and feature stitching is directly performed. The main purpose of this stage is to enable the network to fully learn visual and inertial features. In the joint training stage, the gating fusion module is re-enabled, and the learning rate is adjusted to 5×10 -5 , and 40 iterations of training are carried out. In addition, in the subsequent 20 iterations, the learning rate is further reduced to 1×10 -6 for fine-tuning. In the construction of the loss function, we set the weight of the α parameter to 100 to balance the rotation and translation losses.
[0110] (5.3) Global pose processing
[0111] What is obtained after the model output is the relative displacement and relative rotation from the previous frame to the current frame. However, in order to accurately evaluate the performance of the VIO method and describe the motion trajectory, these relative displacements and rotations need to be restored to the global pose, that is, the coordinates of the camera in the global coordinate system. In this experiment, a method based on the accumulation of displacement increments and rotation matrices is adopted to restore the global pose of the camera.
[0112] The relative rotation predicted by the VIO method of the model of the present invention is represented by Euler angles, and the Euler angles need to be converted into rotation matrices for subsequent conversion operations. Among them, the Euler angles yaw, roll, and pitch are the angles of rotation around the X, Y, and Z axes predicted by the VIO algorithm respectively. The conversion from Euler angles to rotation matrices is as follows:
[0113] R = R x (roll)·R y (pitch)·Rz (yaw);
[0114]
[0115]
[0116] The specific operation of the global pose recovery algorithm is as follows: At the initial state of the sequence, the starting point of the target pose is initialized as the origin of the three-dimensional coordinate system XYZ, and the coordinate system at this moment is defined as the global coordinate system. At this time, the global pose of the camera is the unit matrix I. For each frame, according to the relative displacement t rel and the relative rotation matrix R rel obtained by estimating with the VIO algorithm, the global pose is calculated.
[0117]
[0118] Among them, and represent the global rotation matrix and the global position vector of the previous frame respectively, and represent the global rotation matrix and the global position vector of the current frame respectively. Finally, the global pose is output, and the global pose of each frame is represented as a transformation matrix This matrix is composed of the global rotation matrix and the global position vector. Finally, the six-dimensional vector of the relative displacement and relative rotation output by the VIO algorithm is converted into the global pose, so as to realize the accurate description and performance evaluation of the motion trajectory.
[0119] (5.4) Evaluation Index
[0120] In this embodiment, the core of evaluating the performance of VIO lies in its navigation accuracy. The evaluation of the KITTI dataset in this experiment includes two aspects: qualitative analysis and quantitative analysis. In qualitative analysis, in order to visually compare, odometer trajectories generated by various different methods are drawn for direct comparison. In terms of quantitative analysis, the absolute trajectory error t abs , the relative translation error t ref and the relative rotation error r ref are used as indicators to evaluate the odometer trajectory accuracy.
[0121] The absolute trajectory error (ATE) is a key indicator to measure the accuracy of the odometer, and it is evaluated by comparing the deviation between the trajectory predicted by the algorithm and the actual trajectory. The quantization of ATE is carried out using the root mean square error (RMSE). The absolute translation error t abs is defined as follows:
[0122]
[0123] Among them, n represents the total number of image sequences, and t i respectively represent the translation vector and the true translation vector predicted by the model at time i.
[0124] The Relative Pose Error (RPE) is used to measure the local accuracy of a fixed time difference Δ in the motion trajectory. By comparing the true pose value with the estimated value in real time, the drift of the system can be estimated. The relative translation error t ref and the relative rotation error r ref are defined as follows:
[0125]
[0126] Among them, P1,..., P n ∈ SE(3) represents the pose estimated by the algorithm, Q1,..., Q n ∈ SE(3) represents the true pose, E i represents the relative pose error RPE of the i-th frame, trans(E i ) and rot(E i ) respectively represent taking the translation part and the rotation part in the relative pose error.
[0127] (5.5) Comparative experiments
[0128] (1) Comparative experiments based on the KITTI dataset
[0129] To verify the performance of the VIO method proposed in the present invention, experiments were carried out on the KITTI dataset in this embodiment, and quantitative and qualitative comparisons were made with some current classic visual odometry schemes and deep learning-based odometry schemes.
[0130] Table 1 Experimental results of odometry for sequence 09 of the KITTI dataset (M represents monocular, S represents binocular)
[0131]
[0132] Table 2 Experimental results of odometry for sequence 10 of the KITTI dataset (M represents monocular, S represents binocular)
[0133]
[0134] In this embodiment, three metrics, namely the absolute translation error, relative rotation error, and relative translation error, are used to evaluate the performance of the VIO method proposed in the present invention. Tables 1 and 2 show the quantitative comparison results of several selected comparison methods and the method of the present invention on sequences 09 and 10 of the KITTI dataset. These comparison methods include traditional algorithms and learning-based methods. Among these evaluation metrics, the smaller the three errors, the smaller the error in odometer prediction and the better the positioning effect. It can be seen from the data in the table that the method proposed in the present invention has achieved competitive results in the above three evaluation metrics, which fully demonstrates the effectiveness of the method of the present invention.
[0135] Figure 7 and Figure 8 show the trajectory generation results of the relevant algorithms in sequences 09 and 10 of the KITTI dataset. In sequence 09, which is relatively static and highly structured, ORB-SLAM3 can make full use of its advantages in feature point extraction and matching to effectively perform pose estimation and trajectory accumulation, thus achieving relatively accurate odometer trajectory generation. The generated trajectory highly coincides with the ground truth trajectory, and the trajectory accuracy is at the leading level among all comparison algorithms. However, in sequence 10, which contains more dynamic elements, the performance of ORB-SLAM3 significantly deteriorates. Its feature point extraction and tracking mechanism are vulnerable to interference in dynamic scenarios, resulting in tracking failure and ultimately unable to generate a complete trajectory. This reflects the limitations of ORB-SLAM3 in the face of dynamic scenarios. In contrast, the method proposed in the present invention can more accurately reflect the vehicle's motion trajectory in both sequences 09 and 10, showing better adaptability to dynamic scenarios.
[0136] In the experimental evaluation of the KITTI dataset, there are significant differences in the performance of traditional VIO algorithms in dynamic scenarios. Among them, ORB-SLAM3 shows an average displacement error of 0.013 meters in sequence 09, thanks to its robust matching mechanism based on ORB features and the global bundle adjustment optimization strategy, performing relatively well. However, in sequence 10 where the proportion of dynamic elements is as high as 35%, ORB-SLAM3 experiences trajectory tracking failure. The fundamental reason is that the feature extraction mechanism of ORB-SLAM3 is highly sensitive to dynamic interference. When there are moving pedestrians or vehicles in the scene, the feature points generated on the surface of dynamic objects are easily misjudged as static landmarks, resulting in a high false matching rate during the optical flow tracking process, and further causing cumulative errors in pose estimation. Especially when the dynamic area occupies more than 40% of the image field of view (such as the occlusion of the dynamic truck in sequence 10), the trajectory deviation can reach 3 times that of the static scene. In addition, the sensor configuration characteristics of the KITTI dataset also exacerbate the degradation of the algorithm performance. Its IMU sampling rate (100Hz) and image frame rate (10Hz) are significantly lower than those of typical indoor datasets. The time alignment error caused by the asynchrony of this cross-modal data stream increases the pre-integration standard deviation of ORB-SLAM3, thereby triggering the drift of motion estimation. VINS, which relies on strict initialization conditions, is more vulnerable to this influence. In contrast, the VIO method based on cross-attention and dynamic weights in the present invention reduces the sensitivity to time series through end-to-end motion modeling.
[0137] (5.6) Ablation experiment
[0138] To illustrate the impact of each module and loss function in the method of the present invention on the result improvement, the present invention conducts ablation experiments. To fully verify the effectiveness of the modules included in the method of the present invention, sequence 09 with more straight roads and curves is selected for testing.
[0139] Table 3 Results of ablation experiments on sequence 09 of the KITTI dataset
[0140]
[0141] As can be seen from Table 3, the performance of the basic network without adding any modules is the worst. The performance improves when adding one module in turn, and is optimal when all modules are included. It can be seen that the cross-attention module and dynamic weights can effectively improve the fusion of visual features and inertial features. And it can also be clearly seen from Figure 9 that the basic network has large trajectory errors both in the case of curves and high-speed straight roads in sequence 09, and only performs well on low-speed straight roads. While the complete network framework has the best effect and achieves the smallest error in the case of curves. It is proved that the method of the present invention can better estimate the pose of the current state when facing curve driving and high-speed driving.
[0142] (5.7) Expansion Experiment
[0143] (1) Bend Scenario Analysis Experiment
[0144] As Figure 10 shown, by visualizing the attention map of the cross - attention layer and aligning the attention map with the input image, the regions that the model focuses on when processing the image are displayed. In addition, the attention map is normalized for intuitive display, so as to more clearly show the distribution of the model's focus points. The experimental results show that when the vehicle is in different driving states, the model can focus on the significant regions in the image. For example, when the vehicle executes a left - turn or right - turn action, due to the large changes in the image from the outer - bend perspective, and even some regions may gradually disappear, the model can, through the cross - attention mechanism and according to the IMU data, pay attention to the regions on the outer side of the bend in the image. This shows that the introduction of the cross - attention mechanism enables the network to flexibly adjust the degree of attention to image regions according to the dynamic state of the vehicle, thereby improving the model's adaptability to complex scenarios.
[0145] (2) Feature Weight Analysis Experiment
[0146] As Figure 11 and Figure 12 shown, the feature weights of the model are analyzed on sequences 09 and 10 of the KITTI dataset. The left - hand figure shows the vehicle speed during data collection, and the right - hand figure shows the local feature weights at each moment. It can be intuitively seen that there is a significant correlation between the model's feature fusion strategy and the vehicle motion state. When the vehicle moves in a straight line at high speed, the weight network tends to increase the weight ratio of visual features, and at this time the system dominates the attitude estimation by enhancing visual features; while in the low - speed turning scenario, the weight network will reduce the weight ratio of visual features and increase the weight ratio of inertial features, giving play to the advantages of inertial features in bends. The experimental results prove that the feature weight module of the present invention can dynamically adjust the weights of visual and inertial features according to the current vehicle state, thereby achieving feature balance in different scenarios.
[0147] (3) Data Corruption Scenario Analysis Experiment
[0148] To deeply evaluate the pose estimation performance of the method of the present invention in the face of data corruption, based on the 09 and 10 sequences of the KITTI dataset, the present invention constructs three different types of data corruption scenarios to simulate the abnormal sensor data that may occur in actual applications. These three data corruption scenarios are: IMU data loss, image data loss, and foreign object occlusion. The simulation conditions for each scenario are as follows: IMU data loss scenario: In the two test sequences, 5% of the IMU data is randomly selected, and the selected IMU data is replaced with all-zero data to simulate the situation of IMU data loss; Image data loss scenario: In the two test sequences, 5% of the total amount of image data is randomly selected, and the selected image data is replaced with pure black pictures to simulate the scenario of image data loss; Foreign object occlusion scenario: In the two test sequences, 10% of the total amount of image data is randomly selected, and a pixel coordinate is randomly selected from the selected image data. Centered on this pixel coordinate, a random noise area of 100×100 pixels is placed to block the texture information of the original image data.
[0149] Figure 13 The following shows the comparison experiment results of the baseline algorithm and the method of the present invention in the data corruption scenarios. It can be seen that in the data corruption scenarios, the baseline algorithm is much worse than the method of the present invention in both the 09 and 10 sequences. The baseline algorithm only has better effects in the starting parts of the 09 and 10 sequences. As the trajectory length increases, a large drift occurs in the curved road area. With the accumulation of the drift, the gap between the trajectory and the true trajectory gradually increases, and it is difficult to meet the requirements of positioning and navigation. While the method of the present invention has good performance in the data corruption scenarios. Even when there is a certain degree of data corruption, the drift phenomenon of the method of the present invention remains at a low level, and it can still maintain good pose estimation performance, showing stronger robustness.
[0150] 6 Summary
[0151] In summary, the present invention proposes a visual-inertial odometry pose localization method based on cross-attention and dynamic weights, which is used to enhance the cross-reinforcement and dynamic coupling between visual image features and inertial measurement features, thereby further improving the pose localization accuracy of the visual-inertial odometer, enhancing the stability and reliability of pose localization for different scenarios, and focusing on solving the problems of relatively single fusion method of visual and inertial features, poor pose localization accuracy, and insufficient stability and reliability in multi-scenarios in traditional deep learning VIO algorithms. The method of the present invention uses a visual-inertial odometry pose prediction model to perform localization prediction of visual-inertial odometry pose information. In this visual-inertial odometry pose prediction model, a cross-attention mechanism is used to establish a cross-modal feature association between visual optical flow features and inertial variable features, thereby enhancing the attention to the spatial regions in the visual optical flow features that are highly correlated with inertial variable information; at the same time, with the help of a dynamic weight fusion mechanism, a dynamic balance between different modal features is achieved through a learnable weight allocation strategy, which helps to avoid getting stuck in the local optimal solution problem, significantly improving the accuracy of visual-inertial odometry pose information localization prediction and the stability and reliability for different scenarios. The introduction of the cross-attention mechanism enables the visual-inertial odometry pose prediction model network to flexibly adjust the degree of attention to the image region according to the dynamic state of the moving device, thereby improving the adaptability of the visual-inertial odometry pose prediction model to complex scenarios. At the same time, the introduction of the dynamic weight fusion mechanism can dynamically adjust the weights of visual and inertial features according to the current state of the moving device, thereby achieving feature balance in different scenarios. Therefore, the method of the present invention can more accurately reflect the motion trajectory of the moving device, can better estimate the pose of the current state when driving on a curve and at a high altitude, and has better adaptability to different dynamic scenarios. The method of the present invention has a good performance in the data corruption scenario. Even if there is a certain degree of data corruption, the drift phenomenon of the method of the present invention remains at a low level, and it can still maintain good pose estimation performance, showing stronger robustness and stability.
[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art should understand that any modifications or equivalent replacements to the technical solutions of the present invention without departing from the purpose and scope of the present technical solution should be covered by the scope of the claims of the present invention.
Claims
1. A visual-inertial odometry pose localization method based on cross-attention and dynamic weights, characterized in that Obtain the visual image data and inertial measurement data of the motion device, and input them into the pre-trained visual-inertial odometry pose prediction model to predict the pose information of the motion device; The visual-inertial odometry pose prediction model includes a feature extraction layer, a cross-attention feature fusion layer, and a pose estimation layer; the feature extraction layer is used to respectively extract features from the input visual image data and inertial measurement data to obtain visual optical flow features and inertial variable features; the cross-attention feature fusion layer is used to perform bidirectional modality interaction enhancement processing on the visual optical flow features and inertial variable features through a cross-attention mechanism, and perform dynamic weight fusion to obtain visual-inertial enhanced fusion features; the pose estimation layer is used to perform pose regression processing based on the visual-inertial enhanced fusion features to obtain a pose information prediction result, which is used as the output of the visual-inertial odometry pose prediction model.
2. The visual-inertial odometry pose positioning method based on cross-attention and dynamic weights according to claim 1, characterized in that The visual image data includes two adjacent frames of visual images, and the inertial measurement data includes the inertial time series data corresponding to the time period of the two adjacent frames of visual images.
3. The visual-inertial odometry pose positioning method based on cross-attention and dynamic weights according to claim 2, wherein In the visual-inertial odometry pose prediction model, the feature extraction layer includes a visual feature encoder and an inertial feature encoder; The visual feature encoder is used to stack and splice the input two adjacent frames of visual images in the channel dimension, and then perform multi-scale downsampling feature extraction through multiple convolutional layers, and then output to obtain visual optical flow features; The inertial feature encoder is used to perform high-dimensional feature encoding processing on the inertial time series data corresponding to the time period of the two adjacent frames of visual images input through multiple inertial coding units in sequence, and then output to obtain inertial variable features; each inertial coding unit includes a one-dimensional convolutional module, a batch normalization module, a LeakyReLU activation function, and a Dropout module connected in series in sequence.
4. The visual-inertial odometry pose positioning method based on cross-attention and dynamic weights according to claim 2, wherein In the visual-inertial odometry pose prediction model, the cross-attention feature fusion layer includes a bidirectional cross-attention module and a gated fusion module connected in series; The cross-attention module is used to perform bidirectional modality interaction enhancement processing on the visual optical flow features and inertial variable features through a cross-attention mechanism to respectively obtain visual enhanced features and inertial enhanced features; The gated fusion module is used to dynamically select fusion weights through a gated mechanism to perform dynamic weight fusion on the visual enhanced features and inertial enhanced features to obtain visual-inertial enhanced fusion features.
5. The visual-inertial odometry pose localization method based on cross-attention and dynamic weights according to claim 4, characterized in that The bidirectional cross-attention module includes a visual attention branch and an inertial attention branch; The visual attention branch first passes the visual optical flow feature X through a linear layer v to decompose and generate a visual key vector K v , a visual value vector V v and a visual query vector Q v , and cross-transmits the visual query vector Q v to the inertial attention branch; The inertial attention branch first passes the inertial variable feature X through a linear layer i to decompose and generate an inertial key vector K i , an inertial value vector V i and an inertial query vector Q i , and cross-transfers the inertial query vector Q i to the visual attention branch; then, in the visual attention branch, the visual key vector K v , the visual value vector V v and the inertial query vector Q i are subjected to attention feature extraction and then linearly transformed, and then added to the input visual optical flow feature X v . After passing through a multi-layer perceptron model MLP for feature dimension transformation, the visual enhanced feature as the output is obtained; in the inertial attention branch, the inertial key vector K i , the inertial value vector V i and the visual query vector Q v are subjected to attention feature extraction and then linearly transformed, and then added to the input inertial variable feature X i . After passing through a multi-layer perceptron model MLP for feature dimension transformation, the inertial enhanced feature as the output is obtained; The expressions of the visual enhanced feature and the inertial enhanced feature are: Q v = W Q · X v , K v = W K · X v , V v = W V · X v ; Q i = W Q ·X i ,K i = W K ·X i ,V i = W V ·X i ; Among them, and respectively represent the output visual enhancement feature and inertial enhancement feature; MLP(·) represents the processing of the multi-layer perceptron MLP; Attention(·) represents the operation of attention feature extraction; represents the concatenation and addition operation; Softmax(·) represents the operation of the Softmax function; is the attention scale factor; T represents the transpose symbol; W Q 、W K 、W V 、W a are all linear weight parameters of the bidirectional cross-attention module.
6. The visual inertial odometry pose localization method based on cross-attention and dynamic weights according to claim 4, characterized in that, The gating fusion module divides the visual enhanced features and the inertial enhanced features into the same number of groups of visual enhanced feature components and groups of inertial enhanced feature components respectively. Then, each visual enhanced feature component and inertial enhanced feature component are respectively passed through a convolutional layer to obtain their respective initial fusion weights, and then weighted multiplication is performed with their respective feature components. After that, normalization processing is carried out through a Softmax layer to obtain their respective normalized fusion weights, and then weighted multiplication is performed with their respective feature components again to obtain their respective corresponding weighted feature components. Finally, the groups of visual enhanced weighted feature components and the groups of inertial enhanced weighted feature components are concatenated and fused to obtain the visual-inertial enhanced fusion feature as the output; The expression of the visual-inertial enhanced fusion feature is: F fusion = Concat(F v , F i ); Among them, F fusion represents the obtained visual-inertial enhanced fusion feature, and F v represents the visually enhanced weighted feature formed by connecting each group of visually enhanced weighted feature components, and F i represents the inertia-enhanced weighted feature formed by connecting each group of inertia-enhanced weighted feature components; Concat(·) represents the feature concatenation operation; represents the visually enhanced feature The nth visually enhanced feature component after being divided, represents the inertia-enhanced feature The nth inertia-enhanced feature component after being divided, where n = 1, 2, …, N, and N represents the total number of divided feature components; R v,n and S v,n respectively represent the initial fusion weight and the normalized fusion weight of the nth visually enhanced feature component ; R i,n and S i,n respectively represent the initial fusion weight and the normalized fusion weight of the nth inertia-enhanced feature component ; Softmax(·) represents the Softmax function operation; e is the natural constant; represents the feature concatenation operation for feature components n = 1, 2, …, N.
7. The visual-inertial odometry pose positioning method based on cross-attention and dynamic weights according to claim 2, characterized in that In the visual-inertial odometry pose prediction model, the processing process of the pose estimation layer is: The visually inertial enhanced fusion features are input into a two-layer LSTM network to extract the temporal information in the visually inertial enhanced fusion features. Among them, after each layer of LSTM processing, a multi-layer perceptron (MLP) is used for feature dimension transformation. The output of the LSTM network is then subjected to regression prediction through a fully connected layer to obtain the predicted result of the pose information. The predicted result of the pose information is a 6-dimensional vector, including a 3-dimensional translation vector and a 3-dimensional rotation vector 8. The visual-inertial odometry pose localization method based on cross-attention and dynamic weights according to claim 1, wherein The visual-inertial odometry pose prediction model is trained in the following manner: S101: Prepare a sample data set containing visual image data and inertial measurement data, which is divided into a training data set and a test data set. The visual image data and inertial measurement data in the sample data set are both pre-marked with the true values of the translation vector and the rotation vector; S102: Input the training data set into the visual-inertial odometry pose prediction model for training, and optimize the parameters of the visual-inertial odometry pose prediction model with the goal of minimizing the loss function until the visual-inertial odometry pose prediction model converges, obtaining the trained visual-inertial odometry pose prediction model; S103: Test the visual-inertial odometry pose prediction model through the test data set to confirm the pose information prediction and positioning performance of the trained visual-inertial odometry pose prediction model; if the pose information prediction and positioning performance meets the requirements, end the training of the visual-inertial odometry pose prediction model; otherwise, return to step S102.
9. The visual inertial odometry pose positioning method based on cross-attention and dynamic weights according to claim 8, characterized in that, The loss function used for training the visual-inertial odometry pose prediction model is the mean squared error loss function, and the mean squared error loss function L pose has the following expression: Among them, T v is the total number of frames of the training image sequence; respectively represent the predicted value and the true value of the translation vector; respectively represent the predicted value and the true value of the rotation vector; ||·||2 represents the L2 norm operation; α is the loss weight parameter.
Citation Information
Cited By
Human body posture estimation method and device based on visual inertia fusion
CN121170848A
Unmanned aerial vehicle pose positioning method and device based on visual inertial odometer
CN121453036A
A method and device for UAV pose localization based on visual inertial odometry.
CN121453036B
Multi-sensor fusion positioning method and system based on feature enhancement and dynamic weighting
CN121612275A
A multi-sensor fusion positioning method and system based on feature enhancement and dynamic weighting
CN121612275B