Visual odometry method based on RGB-d dual-mode mutual guidance

CN117974785BActive Publication Date: 2026-10-09BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410135321.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-31
Publication Date
2026-10-09
Estimated Expiration
2044-01-31

AI Technical Summary

Technical Problem

这导致位姿的计算精度还有待进一步提升

Benefits of technology

[0006] Compared with the prior art, the beneficial effects of the present invention are as follows: It proposes a visual odometry method based on RGB-D dual-modal mutual guidance. According to the different characteristics of RGB information and depth information, an RGB-guided depth detail enhancement module and a depth-guided RGB semantic enhancement module are designed to perform bidirectional guidance, fully explore the matching information between multimodalities, and improve the accuracy of visual odometry calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117974785B_ABST
    Figure CN117974785B_ABST
Patent Text Reader

Abstract

The application discloses a visual odometer method based on RGB-D dual-mode mutual guidance. The method adopts a specially designed convolutional neural network model, which can fully mine the relationship between the RGB and depth modes, and uses a pose decoder to calculate the camera pose. The model comprises a pose estimation network and a depth estimation network. The depth estimation network is used for generating a depth map from a single image. In the pose network encoder part, two branch networks process the channel splicing data of adjacent RGB images and depth images respectively. Through an RGB-guided depth detail enhancement module and a depth-guided RGB semantic enhancement module, dual-mode mutual guidance between the RGB and depth data is realized, and the complementary information between the multi-modal data is effectively mined. Finally, the depth features and the RGB features accurately calculate the camera pose through the pose decoder. The application has a significant improvement in feature expression capability, and effectively improves the accuracy of the visual odometer method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of digital image processing and computer vision technology, and relates to a visual odometry method based on RGB-D dual-modal mutual guidance. Background Technology

[0002] Visual odometry (VOM) is a technique that estimates the position and pose of a camera when each frame is captured by analyzing the correspondence between frames. It takes a series of consecutive images as input and outputs a motion trajectory consisting of the position and pose corresponding to each image. As a key component of visual SLAM (Simultaneous Localization and Mapping) systems, VOM has been widely used in various fields, including virtual reality, augmented reality, autonomous driving, and robot navigation.

[0003] Visual odometry (VO) methods are mainly divided into two categories: multi-view geometry-based methods and deep learning-based methods. Multi-view geometry-based VO estimates camera motion by analyzing the geometric relationships between images. While practical, this method heavily relies on manually designed feature descriptors, making it difficult to capture deeper semantic information. With the rapid development of artificial intelligence and deep learning technologies, deep learning-based VO has gradually become a research hotspot. These methods primarily use RGB images for pose estimation, but their accuracy is limited under conditions of drastic changes in lighting, shadows, or environmental variations. Therefore, some research has begun to explore incorporating depth modal information to improve robustness in complex environments. Nevertheless, current VO methods based on the interaction of depth and RGB information have shortcomings in feature encoding. They either treat these two modalities indiscriminately or use overly simplistic interaction methods that fail to fully exploit the rich information between modalities. This means that the accuracy of pose calculation still needs further improvement. Summary of the Invention

[0004] This invention relates to an improved visual odometry method, termed the RGB-D bimodal guided visual odometry method, which aims to enhance feature representation capabilities during the encoding stage, thereby improving the accuracy of pose estimation. This invention innovatively designs two core modules: an RGB-guided depth detail enhancement module and a depth-guided RGB semantic enhancement module. The application of these two modules enables this invention to more effectively mine and integrate the complementarity between RGB images and depth information. Through this deep fusion of bimodal information, the proposed method significantly improves feature representation capabilities, thereby achieving more accurate pose estimation in visual odometry applications and mitigating, to some extent, the trajectory drift problem caused by long-term error accumulation.

[0005] To achieve this goal, the technical solution of this invention is as follows: In the shallow layers of the encoder, the pose estimation network primarily performs detail matching, while deep features ignore complex texture information, resulting in weak feature representation capabilities in the shallow layers. Therefore, an RGB-guided depth detail enhancement module is designed to provide color and detail information to the deep features. In the higher layers of the encoder, the features learned by the network contain more semantic information; deep features provide rich geometric structures and exhibit strong internal consistency. Therefore, a depth-guided RGB semantic enhancement module is designed to provide rich geometric structures for the RGB branch features, strengthening the internal consistency of the RGB features. Finally, the pose is calculated from the deep features and RGB features through a feature decoding module.

[0006] Compared with the prior art, the beneficial effects of the present invention are as follows: It proposes a visual odometry method based on RGB-D dual-modal mutual guidance. According to the different characteristics of RGB information and depth information, an RGB-guided depth detail enhancement module and a depth-guided RGB semantic enhancement module are designed to perform bidirectional guidance, fully explore the matching information between multimodalities, and improve the accuracy of visual odometry calculation. Attached Figure Description

[0007] Figure 1 This is a flowchart of the method involved in the present invention;

[0008] Figure 2 This is a schematic diagram of the network framework of the method of the present invention;

[0009] Figure 3 This is a schematic diagram of the depth estimation network structure of the present invention;

[0010] Figure 4 This is a schematic diagram of the RGB guided depth detail enhancement module proposed in this invention;

[0011] Figure 5 This is a schematic diagram of the deep-guided RGB semantic enhancement module proposed in this invention;

[0012] Figure 6 This is a visualization result of the trajectory of sequences 09 and 10 in the KITTI dataset according to the present invention. Detailed Implementation

[0013] To make the technical solutions and advantages of the present invention easier to understand, the present invention will be further described in detail below with reference to the accompanying drawings.

[0014] The flowchart and overall network framework of this invention are as follows: Figure 1 , Figure 2 As shown, the specific steps include:

[0015] S101: The network input is two consecutive frame images I t with It+1 The method used in this invention first inputs the original images into a depth estimation network (F...). depth In the depth estimation network, the original image input is first processed by a ResNet18 network pre-trained on the ImageNet dataset with the fully connected layers removed for feature extraction. Then the data flows through 11 convolutional layers, each using a 3×3 convolutional kernel. The first 10 layers are followed by the ELU non-linear activation function to enhance expressive power. In addition, after layers 1, 3, 5, 7, and 9, nearest neighbor upsampling is performed to enlarge the feature map size to twice its original size. Finally, after the 11th convolutional layer (with 1 channel), the depth map is normalized by the Sigmoid function to output the final depth map. The specific operation is as follows:

[0016] D t =F depth (I t ), D t+1 =F depth (I t+1 (1)

[0017] Where F depth D represents a deep estimation network. t D t+1 For continuous frame images I t with I t+1 The corresponding depth map.

[0018] S102: The RGB image pairs and depth image pairs are concatenated along the channel dimension and input into the pose estimation network. The pose estimation network encoder is divided into an RGB encoding branch and a depth encoding branch. Both encoders use a ResNet18 network pre-trained on the ImageNet dataset with the fully connected layers removed, using this as the backbone to extract corresponding multi-level feature representations. To enable ResNet18 to handle 6-channel input from the concatenated image, the first 7×7 convolutional layer's input dimension is changed from 3 to 6, while other structures remain unchanged. The RGB and depth features extracted from the shallow Enc_l layer of this network are represented as follows: The features extracted by the deep Enc_h network are represented as follows: l∈{1,2} represents the shallow feature layer of the encoder, h∈{3,4,5} represents the shallow feature layer of the encoder, and rgb and d represent the RGB and depth branches, respectively. This step describes the process of using RGB to guide the depth detail enhancement module to extract features from layers 1-2 of the encoder in the interactive module.

[0019] In the RGB-guided depth detail enhancement module, the upper part first enhances the depth features. and RGB features The features are concatenated along the channels. Next, the concatenated features are input into convolutional module 1, which outputs residual features. Convolution module 1 consists of one 1x1 convolution and one 3x3 convolution plus an activation function. The specific operation is shown in formula (2).

[0020]

[0021] In formula (2), [·, ·] represents the concatenation operation along the channel, Conv_1 represents convolution module 1, and l∈{1, 2} represents the shallow encoder features.

[0022] In the lower half of this module, depth features Pooling is performed along the channel dimension, and then the result is fed into convolutional module 2 to output a weight map w. d Convolutional module 2 consists of two 7×7 convolutions, a ReLU activation function, and a Sigmoid function. (The last part, "w", appears to be a typo and can be left as is.) d With residual characteristics Multiply and with depth features Add them together to output refined depth features. The specific operation is as follows: Formula (3):

[0023]

[0024] In formula (3), maxpool represents max pooling along the channel dimension. * represents element-wise multiplication. Conv_2 represents convolutional module 2. l∈{1,2} represents the shallow encoder feature layer. This indicates the output of the RGB Guided Depth Detail Enhancement (RDDE) module, which will serve as the input to the next layer in the depth branch.

[0025] S103: Input the features obtained in S102 into layers 3 to 5 of the encoder network, which uses the deep-guided RGB semantic enhancement module as the interaction module. The deep-guided RGB semantic enhancement module operates at two main levels: the attention level and the feature level, thereby introducing depth information. At the attention level, the depth features are first processed... Perform max pooling along the channel dimension, and use these pooled features as weights to optimize the features in the RGB branch. The features are re-represented along the spatial dimension to obtain the re-represented features. Then, a weight map obtained through global average pooling is multiplied with the spatially re-represented features for channel-level re-representation, outputting the features after spatial and channel re-representation. The specific operation is as follows: Formula (4)

[0026]

[0027] In formula (4), maxpool represents pooling along the channel dimension, and GAP represents global average pooling operation. h∈{3,4,5} represents the feature layer of the high-level encoder.

[0028] In terms of features, deep features are input into an attention module for feature re-representation to suppress noise in the deep features. The attention module uses the CBAM attention mechanism. For details, see the 2018 ECCV paper "Cbam: Convolutional block attention module" by Sanghyun Woo et al. Subsequently, the features re-represented by the attention module are compared with those re-represented by spatial and channel approaches. Add them together to output refined RGB features. This process can be represented as:

[0029]

[0030] In formula (5), attention represents the attention module. h∈{3,4,5} represents the feature layer of the high-level encoder.

[0031] This represents the output of the Deeply Guided RGB Semantic Enhancement (DRSE) module, which will serve as the input to the next layer of the RGB branch.

[0032] S104: Features and The input is concatenated along the channel dimension and flattened before being fed into a decoder module consisting of fully connected layers (FC layers) to perform pose calculation. The first fully connected layer has a size of 122880×1024, the second 1024×512, the third 512×256, and the fourth 256×6. After the first, second, and third fully connected layers, a ReLU activation function is used to enhance the model's non-linear expressive power.

[0033] The experimental setup is briefly described below. Through comparative analysis of the experiments, the effectiveness of the present invention is demonstrated.

[0034] Experimental environment and hyperparameter configuration: The hardware test platform of this invention uses a 13th Gen CPU. Core TMThe processor is an i9-13900K, and the GPU is an NVIDIA GeForce RTX 4090. The software testing platform for this invention is Ubuntu 22.04.2 LTS. The training and prediction processes are completed using Python, and the deep learning framework used is PyTorch. The Adam optimization function is employed, with a learning rate of 0.001, weights multiplied by 0.5 every 15 epochs, a batch size of 12, and 40 epochs.

[0035] The method disclosed in this invention is evaluated on the KITTIOdometry benchmark, a widely used VO / SLAM evaluation benchmark. This benchmark contains 22 urban and highway driving sequences. Only 11 of these sequences have real trajectory labels from GPS / IMU readings. This invention selects sequences 00-08 as the training set. Quantitative and qualitative evaluations are performed on sequences 09 and 10.

[0036] 1. Comparison of quantitative results

[0037] This invention uses absolute trajectory error (ATE) as the quantification metric, and the results are shown in Table 1. The MLF_VO method is the "Self-Supervised Ego-Motion Estimation Based on Multi-Layer Fusion of RGB and Inferred Depth" method published by Jiang et al. at ICRA 2022. Experimental results show that the average error of the method proposed in this invention is lower than that of the MLF_VO method on sequences 09 and 09, 10, proving that the method of this invention improves the accuracy of visual odometry.

[0038] Table 1. ATE results for sequences 09 and 10 in the KITTI dataset.

[0039]

[0040] 2. Comparison of Qualitative Results

[0041] This method visualizes the trajectories of sequences 09 and 10 as follows: Figure 5 As shown in the figure. The dotted line represents the truth trajectory, the dashed line represents the trajectory of the MLF_VO method, and the solid line represents the method of this invention. The results show that the method of this invention can maintain the global trajectory and reduce trajectory drift. On sequence 09, the trajectory of this invention is significantly improved compared to MLF_VO, and it achieves similar results compared to sequence 10.

[0042] The function and effect of this method

[0043] This invention discloses a visual odometry method based on RGB-D bimodal mutual guidance. It mainly comprises two modules: a depth estimation module and a pose estimation module. This invention improves the pose estimation module by designing an RGB-guided depth detail enhancement module between the two-branch encoders to re-represent depth features using RGB features, thereby enhancing the network's feature representation capability in the shallow layers of the encoder. Simultaneously, a depth-guided RGB semantic enhancement module is designed to re-represent RGB features using depth features, enhancing the network's feature representation capability in the higher layers of the encoder. These two designs can fully exploit intermodal information. Comparative experiments on the KITTI dataset demonstrate that the proposed method has better feature representation capability, more accurate camera pose estimation, and can further mitigate trajectory drift.

Claims

1. A visual odometry method based on RGB-D dual-modal mutual guidance, characterized in that: S101: The network input is two consecutive frame images I t with I t+1 First, the original images are input into the depth estimation network F. depth middle; In the depth estimation network, the original image input is first processed by a ResNet18 network pre-trained on the ImageNet dataset with fully connected layers removed for feature extraction. Then, the data flows through 11 convolutional layers, each using a 3×3 convolutional kernel, and the ELU non-linear activation function is applied after the first 10 layers. Nearest neighbor upsampling is performed after layers 1, 3, 5, 7, and 9 to double the size of the feature map. Finally, after the 11th convolutional layer with 1 channel, the feature map is normalized using the Sigmoid function to output the final depth map. The specific operation is as follows: (1) D t =F depth (I t ),D t+1 =F depth (I t+1 ) (1) Where F depth D represents a deep estimation network. t D t+1 For continuous frame images I t with I t+1 The corresponding depth map; S102: Concatenate the RGB image pairs and depth image pairs along the channel dimension and input them into the pose estimation network; the pose estimation network encoder is divided into an RGB encoding branch and a depth encoding branch; both encoders use a ResNet18 network pre-trained on the ImageNet dataset with fully connected layers removed, and use it as the backbone to extract corresponding multi-level feature representations; The input dimension of its first 7×7 convolutional layer was changed from 3 to 6, while other structures remained unchanged; the RGB and depth features extracted from the shallow Enc_l layer of this network are represented as follows: The features extracted by the deep Enc_h network are represented as follows: The encoder shallow feature level is represented by h∈{3,4,5}, and rgb and d represent the RGB and depth branches, respectively. In the RGB-guided depth detail enhancement module, the upper part first enhances the depth features. and RGB features The features are concatenated along the channels; then, the concatenated features are input into convolutional module 1 to output residual features. Convolution module 1 consists of a 1x1 convolution and a 3x3 convolution plus an activation function; the specific operation is as shown in formula (2). In formula (2), [·,·] represents the concatenation operation along the channel, Conv_1 represents convolution module 1, and l∈{1,2} represents the shallow encoder features; In the lower half of this module, depth features Pooling is performed along the channel dimension, and then the result is fed into convolutional module 2 to output a weight map w. d Convolutional module 2 consists of two 7×7 convolutions, a ReLU activation function, and a Sigmoid function; w d With residual characteristics Multiply and with depth features Add them together to output refined depth features. The specific operation is the formula. (3): In formula (3), maxpool represents max pooling along the channel dimension; * represents element-wise multiplication; Conv_2 represents convolution module 2; l∈{1,2} represents the shallow encoder feature layer; This indicates the output of the RGB Guided Depth Detail Enhancement (RDDE) module, which will serve as the input to the next layer in the depth branch; S103: Input the features obtained in S102 into layers 3 to 5 of the encoder network with the deep-guided RGB semantic enhancement module as the interaction module; the deep-guided RGB semantic enhancement module works at two main levels: the attention level and the feature level. At the attention level, the deep features are first processed... Perform max pooling along the channel dimension, and use these pooled features as weights to optimize the features in the RGB branch. The features are re-represented in the spatial dimension to obtain the re-represented features. Then, the weight map obtained by global average pooling is multiplied with the re-represented features in the spatial dimension to perform channel-level re-representation, outputting the features after spatial and channel re-representation. The specific operation is as follows: Formula (4) In formula (4), maxpool represents pooling along the channel dimension, GAP represents global average pooling operation; h∈{3,4,5} represents the feature layer of the high-level encoder; In terms of features, deep features are input into the attention module for feature re-representation to suppress noise in the deep features; The attention module employs the CBAM attention mechanism; features re-represented through the attention module differ from those re-represented through spatial and channel approaches. Add them together to output refined RGB features. This process is represented as: In formula (5), attention represents the attention module; h∈{3,4,5} represents the feature layer of the high-level encoder; This represents the output of the deep-guided RGB semantic enhancement module, h∈{3,4,5}, which will be used as the input to the next layer of the RGB branch; S104: Features and The input is concatenated along the channel dimension and flattened before being fed into a decoder module consisting of fully connected layers (FC layers) to perform pose calculation. The size of the first fully connected layer is 122880×1024, the size of the second fully connected layer is 1024×512, the size of the third fully connected layer is 512×256, and the size of the fourth fully connected layer is 256×6. After the first, second, and third fully connected layers, the ReLU activation function is used.

2. The visual odometry method based on RGB-D dual-modal mutual guidance according to claim 1, characterized in that: The hardware testing platform uses a 13th Gen CPU. Core TM The system used an i9-13900K processor and an NVIDIA GeForce RTX 4090 GPU. The software testing platform was Ubuntu 22.04.2 LTS. The training and prediction process was completed using Python and the deep learning framework PyTorch was used. The Adam optimization function was employed, with a learning rate of 0.001, weights multiplied by 0.5 every 15 epochs, a batch size of 12, and 40 epochs.

Citation Information

Patent Citations

  • Monocular video depth estimation method based on deep convolutional network

    CN113570658A

  • Dynamic environmental sense speedometer method of RGB-D camera based on edge information

    CN113837243A