Monocular visual odometer positioning method based on end-to-end deep learning

Through the monocular visual odometer method of end-to-end deep learning, the monocular visual odometer is integrated with time and depth information using the cosine-cosine-modulation function and wavelet convolution technology, and the problems of fuzzy scale and high computing resources of the monocular visual odometer are solved, and stable positioning and efficient calculations are achieved in complex scenarios.

CN120259619AActive Publication Date: 2025-07-04NORTHWESTERN POLYTECHNICAL UNIV

Patent Information

Application Number
CN202510753781.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing monocular visual odometer system based on deep learning has problems such as blurred scale, insufficient time series utilization and high computing resource requirements, resulting in instability in positioning and poor real-time performance.

Method used

The end-to-end deep learning method is adopted to encode the time stamps positionally through the cosine-cosine-modulation function, and the global depth information is extracted in combination with wavelet convolution and convolution layer. The fusion module is used to integrate initial pose estimation, global depth information and time pose information to achieve 6-DOF pose estimation.

Benefits of technology

It significantly improves the accuracy of the translation and rotation estimation of monocular visual odometers, ensures stable and reliable positioning performance in complex scenarios, reduces the computational complexity, and is suitable for fields such as unmanned driving, robot navigation and augmented reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259619A_ABST
    Figure CN120259619A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision and robot navigation. The invention provides a monocular vision odometer positioning method based on end-to-end deep learning. According to the embodiment of the invention, an explicit scale correction method is adopted, deep modeling is carried out on a complete time sequence, and fine motion changes between continuous frames can be accurately captured, so that the accuracy of the monocular visual odometer in translation and rotation estimation is remarkably improved, and the problems of error accumulation and scale blur in a traditional method are effectively solved. The fusion module efficiently integrates initial pose estimation, global depth information and time pose information, ensures that the system can maintain stable and reliable positioning performance in various complex scenes such as cities, villages and high-speed driving, and shows excellent robustness and wide applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical fields of computer vision and robot navigation, and in particular, to a monocular visual odometry localization method based on end-to-end deep learning. Background Art

[0002] In the forefront of robotics and intelligent perception, visual odometry (VO) is gradually becoming an important enabling technology for robot autonomous navigation and environmental perception. Traditional positioning systems often rely on multi-sensor fusion and pre-built maps, while visual odometry uses consecutive video frames captured by a camera to directly infer the motion trajectory and pose information of a robot in space through real-time image processing and motion estimation. The core challenge of visual odometry lies in how to extract stable and reliable motion features from raw visual data and accurately calculate the 6-degree-of-freedom (6-DOF) pose transformation between adjacent frames of a robot or a camera device, that is, the information including both translation and rotation. Especially in monocular visual odometry, since the system only relies on the image sequence captured by a single camera and lacks direct depth information, this makes the pose estimation process more complex and unstable. In recent years, with the continuous progress of image processing algorithms, feature matching techniques, and deep learning methods, monocular visual odometry has achieved remarkable improvements in accuracy, robustness, and real-time performance, and has been gradually applied to multiple fields such as unmanned driving, robot navigation, and augmented reality.

[0003] Traditional visual odometry mainly relies on geometric methods, which are generally divided into two categories: indirect methods and direct methods: (1) Indirect methods, also known as feature point methods, mainly first detect a set of representative feature points (such as SIFT, ORB, SURF, etc.) in an image, and then achieve image matching by extracting descriptors of these feature points. After the matching is completed, geometric constraints (such as the essential matrix or the fundamental matrix) and robust estimation methods such as RANSAC are used to eliminate incorrect matches, so as to calculate the relative motion between adjacent frames. Indirect methods usually perform pose estimation based on sparse feature points, so they are relatively efficient computationally and have good robustness to illumination changes. However, it depends on sufficient texture information in the image and is easily affected by motion blur and occlusion during the feature point detection and matching process, and may accumulate large errors during long-term operation.

[0004] (2) The direct method estimates camera motion by directly utilizing the image pixel intensity information without explicitly extracting and matching feature points. This method assumes that the image satisfies brightness invariance between adjacent frames, that is, the same physical points have similar gray values in consecutive frames. The direct method solves for the relative pose of the camera by minimizing the photometric error between consecutive frames and using iterative optimization (such as the Gauss-Newton or Levenberg-Marquardt algorithm). This method can utilize all pixel information in the image, so relatively stable estimates can also be obtained in low-texture regions. However, the direct method is highly sensitive to illumination changes, motion blur, and large displacements, and is prone to falling into local optima when the initialization is poor, with a large computational amount, and sometimes it is difficult to ensure real-time performance.

[0005] In recent years, deep learning technology has made significant progress in the field of visual odometry. The end-to-end network model can directly learn the pose mapping from the image sequence, reducing the engineering complexity and improving the robustness to a certain extent.

[0006] However, the existing end-to-end visual odometry systems based on deep learning have the following deficiencies: (1) Scale ambiguity problem: In a monocular system, due to the lack of absolute depth information, scale uncertainty often occurs in translational estimation; (2) Insufficient utilization of time series: Some models only use local time information and fail to fully capture the time-dependent relationships in the complete sequence; (3) High computational resource requirements: Although some hybrid methods improve the accuracy, steps such as backend bundle optimization bring a large computational overhead, which is not conducive to real-time applications.

[0007] Therefore, it is necessary to improve one or more problems existing in the above related technical solutions.

[0008] It should be noted that this part aims to provide background or context for the technical solutions of the present disclosure stated in the claims. The descriptions herein are not admitted to be prior art merely because they are included in this part. Summary of the Invention

[0009] The purpose of the embodiments of the present disclosure is to provide a monocular visual odometry positioning method based on end-to-end deep learning, thereby at least overcoming one or more problems caused by the limitations and defects of the related technologies to a certain extent.

[0010] According to the embodiments of the present disclosure, a monocular visual odometry positioning method based on end-to-end deep learning is provided, and the method includes: Obtain consecutive video frames, synchronously record the timestamp of each video frame, and preprocess each video frame to obtain an input image; Perform position encoding on the timestamp using sine-cosine harmonic functions to obtain a high-dimensional mapping; Perform implicit transformation on the high-dimensional mapping to obtain time pose information; wherein, the time pose information includes a translation vector and a rotation vector; Initialize the input image to obtain an initial pose estimate; wherein, the initial pose estimate includes an initial translation estimate and an initial rotation estimate; Extract features from the input image to obtain the depth map of the input image; Use wavelet convolution operations and convolutional layers to extract the global attention of the depth map to obtain global depth information; Perform multi-modal fusion on the initial pose estimate, global depth information, and time pose information to obtain a 6-DOF pose to complete positioning.

[0011] Furthermore, the method further includes: Construct a monocular visual odometer, which includes a preprocessing module, a pose initialization module, a time pose network, a depth estimation network, a depth information extraction module, and a fusion module.

[0012] Furthermore, in the step of obtaining continuous video frames, synchronously recording the timestamp of each video frame, and preprocessing each video frame to obtain the input image, it includes: Continuously capture video frames at a predetermined frame rate and synchronously record the timestamp of each video frame; Input the video frames into the preprocessing module for size normalization, color correction, and noise suppression processing to obtain the input image to improve image quality and data consistency.

[0013] Furthermore, in the step of initializing the input image to obtain an initial pose estimate, it includes: Input the input image into the pose initialization module for initialization to obtain an initial translation estimate and an initial rotation estimate.

[0014] Furthermore, in the step of performing position encoding on the timestamp using sine-cosine harmonic functions to obtain a high-dimensional mapping, it includes: Input the timestamp into the time pose network, use sine-cosine harmonic functions to generate 2L sine values and cosine values with different frequencies according to the timestamp, and embed the timestamp into the high-dimensional space to obtain the corresponding high-dimensional mapping; Wherein, the expression of the high-dimensional mapping is:

[0015] In the formula, L is the sequence length, is the sine-cosine harmonic function, is the timestamp.

[0016] Further, in the step of implicitly transforming the high-dimensional mapping to obtain the time pose information, it includes: Input the high-dimensional mapping into the T-Net and R-Net in the time pose network respectively. The T-Net implicitly transforms the high-dimensional mapping, uses the MLP combined with the ReLU activation function and the tanh activation function, adjusts the output to 3 dimensions, and limits its value in the interval [-1, 1] to output the translation vector. The R-Net implicitly transforms the high-dimensional mapping, adjusts the output to 4 dimensions, represents the rotation in the form of quaternion [qw, qx, qy, qz], performs the tanh activation on the [qx, qy, qz] components, performs the 1 - tanh transformation on the [qw] component, and performs the L2 normalization overall to construct the unit quaternion to output the rotation vector.

[0017] Further, in the step of extracting features from the input image to obtain the depth map of the input image, it includes: Input the input image into the pre-trained depth estimation network for feature extraction to generate the depth map corresponding to the input image.

[0018] Further, in the step of using the wavelet convolution operation and the convolutional layer to extract the global attention of the depth map to obtain the global depth information, it includes: Input the depth map into the depth information extraction module, use the wavelet convolution operation to downsample the depth map, adjust the size of the depth map to H / 2×W / 2, and add all frequency bands along the channel dimension to make the number of channels become 4 times the original to obtain the low-frequency subband, the first high-frequency subband, the second high-frequency subband, and the third high-frequency subband; where H is the height of the depth map and W is the width of the depth map. Use the convolutional layer to extract and reconstruct the low-frequency subband, the first high-frequency subband, the second high-frequency subband, and the third high-frequency subband respectively to obtain the depth map feature with the size of H / 2×W / 2×D; where D is the number of channels. Construct the Query, Key, and Value of the attention module, calculate the attention score in combination with the depth map feature to obtain the global depth information. Among them, the expression of the global depth information is:

[0019] Among them, is the global depth information, is the first weight matrix, is the second weight matrix, is the third weight matrix, represents the wavelet convolution operation, represents the Sigmoid activation, is the depth map feature.

[0020] Further, in the step of performing multi-modal fusion on the initial pose estimation, global depth information, and temporal pose information to obtain the 6-DOF pose, it includes: Using a fusion module to perform linear weighted fusion on the translational initial estimation, global depth information, and translational vector to obtain a translational fusion feature; Using a fusion module to perform linear weighted fusion on the rotational initial estimation and rotational vector to obtain a rotational fusion feature; Using a pose regression multi-layer perceptron preset in the fusion module to perform non-linear mapping on the translational fusion feature and rotational fusion feature to obtain the 6-DOF pose.

[0021] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: In the embodiments of the present disclosure, through the above-mentioned monocular visual odometry positioning method based on end-to-end deep learning, on the one hand, by adopting an explicit scale correction method and performing depth modeling on the complete time series, it can accurately capture the subtle motion changes between consecutive frames, thereby significantly improving the accuracy of the monocular visual odometry in translational and rotational estimation and effectively solving the problems of error accumulation and scale ambiguity in traditional methods. On the other hand, the fusion module efficiently integrates the initial pose estimation, global depth information, and temporal pose information, ensuring that the system can maintain stable and reliable positioning performance in various complex scenarios such as urban, rural, and high-speed driving, demonstrating excellent robustness and wide applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0023] Figure 1 A step diagram showing a monocular visual odometry positioning method based on end-to-end deep learning in an exemplary embodiment of the present disclosure; Figure 2 A network structure diagram showing a monocular visual odometry in an exemplary embodiment of the present disclosure; Figure 3 A schematic diagram showing the principle of wavelet convolution operation in an exemplary embodiment of the present disclosure; Figure 4 A test result diagram showing sequence 09 in an exemplary embodiment of the present disclosure; Figure 5 A test result diagram showing sequence 10 in an exemplary embodiment of the present disclosure; Figure 6 Shows the GPU memory occupancy results in the exemplary embodiments of the present disclosure. Detailed implementation manners

[0024] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0025] In addition, the accompanying drawings are only schematic illustrations of the embodiments of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.

[0026] A monocular visual odometry positioning method based on end-to-end deep learning is provided in the present example embodiment. Referring to Figure 1 as shown in, the monocular visual odometry positioning method based on end-to-end deep learning may include: Step S101: Obtain consecutive video frames, synchronously record the timestamps of each video frame, and preprocess each video frame to obtain an input image; Step S102: Perform position encoding on the timestamps using a sine-cosine harmonic function to obtain a high-dimensional mapping; Step S103: Perform an implicit transformation on the high-dimensional mapping to obtain time pose information; wherein, the time pose information includes a translation vector and a rotation vector; Step S104: Initialize the input image to obtain an initial pose estimate; wherein, the initial pose estimate includes an initial translation estimate and an initial rotation estimate; Step S105: Extract features from the input image to obtain a depth map of the input image; Step S106: Use wavelet convolution operations and convolutional layers to extract the global attention of the depth map to obtain global depth information; Step S107: Perform multi-modal fusion on the initial pose estimate, the global depth information, and the time pose information to obtain a 6-DOF pose to complete the positioning.

[0027] Through the above monocular visual odometry localization method based on end-to-end deep learning, on the one hand, by adopting an explicit scale correction method and performing deep modeling on the complete time series, it can accurately capture the subtle motion changes between consecutive frames, thereby significantly improving the accuracy of monocular visual odometry in translation and rotation estimation, and effectively solving the problems of error accumulation and scale ambiguity in traditional methods. On the other hand, the fusion module efficiently integrates the initial pose estimation, global depth information, and temporal pose information, ensuring that the system can maintain stable and reliable localization performance in various complex scenarios such as urban, rural, and high-speed driving, demonstrating excellent robustness and wide applicability.

[0028] Next, reference will be made to Figures 1 to 6 to describe each step of the above monocular visual odometry localization method based on end-to-end deep learning in the present exemplary embodiment in more detail.

[0029] In step S101, consecutive video frames are acquired, the timestamp of each video frame is synchronously recorded, and each video frame is preprocessed to obtain an input image.

[0030] Specifically, in the initial stage, the monocular camera continuously captures video frames at a predetermined frame rate (10Hz), and each video frame is accompanied by an accurate timestamp record to ensure the accuracy of the temporal information in the subsequent processing. The acquired video frames first pass through a preprocessing module, which performs size normalization, color correction, and noise suppression on each frame of the image, aiming to improve the image quality and data consistency, and thus ensure the high reliability of the input data for the depth estimation network.

[0031] In steps S102 and S103, the timestamp is position-encoded using a sine-cosine harmonic function to obtain a high-dimensional mapping; the high-dimensional mapping is implicitly transformed to obtain temporal pose information; where the temporal pose information includes a translation vector and a rotation vector.

[0032] Specifically, for each timestamp t, first, a sine-cosine harmonic function is used for position encoding to obtain a high-dimensional mapping t*. By using the sin and cos functions to perform different-frequency transformations on the timestamp, it is mapped into a d-dimensional vector containing multi-level frequency-domain information. This method can not only capture the periodic changes and smooth continuity in time but also reflect the subtle differences between adjacent frames, providing rich and stable temporal features for the subsequent network, thereby achieving an accurate continuous mapping from time to pose.

[0033] Input t* into the temporal pose network TimePoseNet (including T_Net and R_Net), T_Net outputs a translation vector v ∈ R³, and R_Net outputs a rotation quaternion (i.e., a rotation vector) q ∈ R 4 (satisfying ).

[0034] Among them, T-Net consists of 8 layers of MLP. The output dimension of the last layer is fixed at 3, corresponding to the translational vector of three degrees of freedom. The ReLU activation function is used in the middle layer, and the tanh activation function is used in the last layer to limit the translation value in the interval [-1, 1].

[0035] R-Net also consists of 8 layers of MLP. The output dimension of the last layer is fixed at 4, corresponding to the quaternion [qw, qx, qy, qz], which is used to represent rotation. The ReLU activation function is used in the middle layer. For the [qx, qy, qz] components in the last layer, the tanh activation is performed, and for the qw component, the 1 - tanh transformation is performed. Then, the entire quaternion is L2-normalized to ensure that the output is a unit quaternion.

[0036] In step S104, the input image is initialized to obtain an initial pose estimate; among them, the initial pose estimate includes an initial translation estimate and an initial rotation estimate. Obtaining the initial pose estimate is a prior art and will not be elaborated in detail here.

[0037] In step S105 and step S106, the input image is feature-extracted to obtain the depth map of the input image; the global attention of the depth map is extracted by using wavelet convolution operations and convolutional layers to obtain global depth information.

[0038] Specifically, the preprocessed input image is then fed into a pre-trained depth estimation network, which uses an advanced convolutional neural network structure to quickly generate the depth map corresponding to the current frame, providing a solid foundation for subsequent scale correction and feature extraction.

[0039] The wavelet convolution operation is used to perform lossless downsampling on the depth map. By using the simple and efficient characteristics of the Haar wavelet, the depth map is decomposed into 4 sub-bands (i.e., the low-frequency sub-band, the first high-frequency sub-band, the second high-frequency sub-band, and the third high-frequency sub-band), and through an appropriate combination of convolutional kernels, the extraction and reconstruction of the low-frequency information of the image are realized. The size of the downsampled depth map is reduced to 1 / 2 of the original image, and the number of channels d = 64, thereby reducing the computational amount while providing stable and accurate global depth information for subsequent feature extraction and attention mechanisms.

[0040] The depth features are extracted through subsequent convolutional layers, and the Query, Key, and Value of the attention module are constructed. After calculating the attention scores, the global depth feature vector F_d (i.e., global depth information) is obtained. This vector fully reflects the overall depth distribution of the image and provides reliable data information and guidance for subsequent translational scale correction.

[0041] In step S107, the initial pose estimate (obtained from the initial pose network) and the time pose information obtained by TimePoseNet are respectively extended to the same dimension as F_d through fully connected layers to ensure effective docking between different modality features. Using the fusion module, first, the initial translation feature is linearly weighted and fused with F_d to obtain the fused feature F_fuse (i.e., the translation fused feature); at the same time, a similar process is performed on the rotation part to obtain the rotation fused feature. Finally, the translation fused feature and the rotation fused feature after fusing the high-dimensional features are respectively input into a preset pose regression multi-layer perceptron, and the final 6-degree-of-freedom pose P_final is output, realizing the conversion from multi-source information to accurate pose estimation.

[0042] In a specific embodiment, as Figure 2 shown, the network structure diagram of the monocular visual odometer, and its specific processing flow is as follows: Preprocess the input adjacent frame images, crop them to an appropriate size and convert them into the corresponding format, and obtain the initial pose estimate through the pose initialization module. Among them, the initial pose estimate includes the initial translation estimate and the initial rotation estimate.

[0043] To learn the time-to-pose mapping, an MLP network called TimePoseNet (i.e., the time pose information) is designed to achieve an accurate continuous mapping from time to pose, so as to obtain the time pose information fθ(t). Among them, TimePoseNet contains two independent sub-networks T_Net and R_Net, which are respectively used to learn the mapping from time input to translation and the mapping from time input to rotation. Each network consists of 8 layers of MLP, including the ReLU activation function and 256-dimensional hidden units.

[0044] Before inputting into the network, first use the sine-cosine harmonic function to perform position encoding on the timestamp t. For each timestamp t, the γ function (i.e., the sine-cosine harmonic function) generates 2L values, which are the sine and cosine values of different frequencies (determined by ), where i ranges from 0 to L - 1), so as to embed the time variable t into the high-dimensional space and obtain its high-dimensional mapping , and the formula is as follows: (1) where L represents the sequence length.

[0045] Input into the time pose network (including T_Net and R_Net). T_Net outputs the translation vector , and R_Net outputs the rotation quaternion (i.e., the rotation vector) (satisfying ).

[0046] Input the current frame image (i.e., video frame) into the depth estimation network to obtain the depth map of the image. Combine wavelet transform to extract global attention and obtain global depth information. Input wavelet convolution to downsample it, adjust its shape to H / 2×W / 2, and then sum all frequency bands along the channel dimension to make the number of channels become 4 times the original. Then, use the convolutional layer to extract features while increasing its dimension. Finally, obtain the depth map features with the shape of H / 2×W / 2×D , where the number of channels D = 64 is set.

[0047] Input the depth map features after wavelet convolution into the linear layer, construct Query, Key and Value, perform attention calculation according to formula (2) to obtain global depth information, and finally convert it into a depth feature vector suitable for subsequent fusion .

[0048] (2) where, is the first weight matrix, is the second weight matrix, is the third weight matrix, represents the wavelet convolution operation, represents the Sigmoid activation, is the depth map feature.

[0049] Finally, after obtaining the global depth information and the temporal pose information fθ(t), use the fusion module to perform effective multi-modal fusion on the initial pose, depth map information and temporal information, and finally map it into the required 6-DOF pose through the fully connected layer.

[0050] In a specific embodiment, as Figure 3 shown, it is the schematic diagram of the wavelet convolution operation. The image is divided into one low-frequency part and three high-frequency parts through wavelet transform (i.e., wavelet convolution operation), and they are connected together on the channel for convolution, so that a larger receptive field can be obtained only by using a small convolution kernel, saving computing resources while improving the effect of feature extraction, and also providing additional frequency domain information for subsequent features, which is beneficial to better focus on image details. At the same time, based on the characteristics of wavelet transform, the image is also downsampled during the frequency domain transformation (one wavelet transform can change the feature map of H×W×D into H / 2×W / 2×4D). Since all frequency band data are used subsequently, it can be considered that the wavelet transform realizes lossless downsampling operation, reducing information loss in the whole feature extraction process.

[0051] In a specific embodiment, this application is implemented on a system with a 12th Gen Intel(R) Core(TM) i5-12400F 2.50 GHz CPU, 64G of memory, an NVIDIA RTX 3090 graphics card, and the Ubuntu20.04 operating system, based on the Pytorch1.7.1 and Python3.9 language environments.

[0052] In terms of the dataset, the KITTI odometry dataset is a widely used benchmark in autonomous driving and computer vision research, jointly released by KIT and TTI Chicago. It contains 22 stereo sequences saved in the lossless PNG format: 11 sequences (00-10) contain ground truth trajectories, and 10 sequences (11-21) do not contain ground truth trajectories and are commonly used for testing.

[0053] The KITTI odometry evaluation criteria examine subsequences ranging from 100 to 800 meters in length and report the average translational error and rotational error. The Absolute Trajectory Error (ATE) measures the root mean square error between the predicted camera poses and the ground truth, while the Relative Pose Error (RPE) captures the relative pose error between frames. Lower metric values indicate that the results are closer to the ground truth.

[0054] 1. Training Settings During the system training phase, this application uses the 00-08 sequences of the KITTI dataset as supervised data and employs the AdamW optimizer to train the entire network. During the training process, the system uses the Mean Squared Error (MSE) as the loss function for both the translational and rotational parts, and applies an additional weight (multiplied by 100) to the rotational loss to ensure that the rotational error is fully considered. The entire network is trained within 30 epochs, with 32 frames of images processed in each batch, and a linear learning rate decay strategy is adopted to ensure the smooth convergence of the training process.

[0055] 2. Experimental Results First, the method proposed in this application and the existing monocular visual odometry methods are tested and compared on sequences 09 and 10 that were not seen during training. As shown in Table 1, the experimental results prove that the method proposed in this application reaches the best level in the vast majority of metrics on the two unknown sequences. This indicates that this application has high accuracy and stability overall. The best results are shown in bold, and the sub-optimal results are underlined.

[0056] Table 1 Sequence Test Results

[0057] At the same time, a visual comparison of each comparison method and the system of this application with the ground truth was also conducted, and the results are asFigure 4 and Figure 5 as shown. Among them, Figure 4 is the test result graph of Sequence 09; Figure 5 is the test result graph of Sequence 10.

[0058] From Figure 4 and Figure 5 it can be seen that the method proposed in this application is the closest to the ground truth results on both sequences and is more suitable for monocular visual odometry tasks in a wide range of scenarios.

[0059] Finally, the GPU memory occupancy during runtime was tested for the overall system of this application, as Figure 6 shown.

[0060] From Figure 6 it can be seen that during runtime, this application only requires 2.3G of memory occupancy, which is lower than many other existing end-to-end visual odometry methods, demonstrating that this application not only has an advantage in positioning accuracy but also meets the actual requirements in real-time applications and resource-constrained environments.

[0061] Through the above monocular visual odometry positioning method based on end-to-end deep learning, on the one hand, by adopting an explicit scale correction method and performing deep modeling on the complete time series, it can accurately capture the subtle motion changes between consecutive frames, thereby significantly improving the accuracy of monocular visual odometry in translational and rotational estimation and effectively solving the problems of error accumulation and scale ambiguity in traditional methods. On the other hand, the fusion module efficiently integrates the initial pose estimation, global depth information, and temporal pose information to ensure that the system can maintain stable and reliable positioning performance in various complex scenarios such as urban, rural, and high-speed driving, demonstrating excellent robustness and wide applicability. At the same time, the end-to-end network architecture effectively simplifies the system design, avoids the heavy backend bundle optimization steps in traditional methods, thereby reducing the computational complexity and implementation difficulty and facilitating popularization and deployment in practical applications.

[0062] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of this disclosure, "a plurality" means two or more unless otherwise specifically defined.

[0063] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0064] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.

Claims

1. A monocular visual odometry localization method based on end-to-end deep learning, characterized in that The method includes: Obtain consecutive video frames, synchronously record the timestamp of each video frame, and preprocess each video frame to obtain an input image; Perform position encoding on the timestamp using a sine-cosine harmonic function to obtain a high-dimensional mapping; Perform implicit transformation on the high-dimensional mapping to obtain time pose information; wherein, the time pose information includes a translation vector and a rotation vector; Initialize the input image to obtain an initial pose estimate; wherein, the initial pose estimate includes an initial translation estimate and an initial rotation estimate; Extract features from the input image to obtain the depth map of the input image; Use wavelet convolution operations and convolutional layers to extract the global attention of the depth map to obtain global depth information; Perform multi-modal fusion on the initial pose estimate, global depth information, and time pose information to obtain a 6-DOF pose to complete positioning.

2. The monocular visual odometry localization method based on end-to-end deep learning according to claim 1, wherein The method further includes: Construct a monocular visual odometer, which includes a preprocessing module, a pose initialization module, a time pose network, a depth estimation network, a depth information extraction module, and a fusion module.

3. The monocular visual odometry localization method based on end-to-end deep learning according to claim 2, wherein In the step of obtaining consecutive video frames, synchronously recording the timestamp of each video frame, and preprocessing each video frame to obtain an input image, it includes: Continuously capture video frames at a predetermined frame rate and synchronously record the timestamp of each video frame; Input the video frames into the preprocessing module for size normalization, color correction, and noise suppression processing to obtain an input image to improve image quality and data consistency.

4. The monocular visual odometry localization method based on end-to-end deep learning according to claim 3, wherein In the step of initializing the input image to obtain an initial pose estimate, it includes: Input the input image into the pose initialization module for initialization to obtain an initial translation estimate and an initial rotation estimate.

5. The monocular visual odometry positioning method based on end-to-end deep learning according to claim 4, characterized in that, In the step of performing position encoding on the timestamp using a sine-cosine harmonic function to obtain a high-dimensional mapping, it includes: Input the timestamp into the time pose network, use the sine-cosine harmonic function to generate 2L sine values and cosine values with different frequencies according to the timestamp, and embed the timestamp into a high-dimensional space to obtain the corresponding high-dimensional mapping; Wherein, the expression of the high-dimensional mapping is: where L is the sequence length, is the sine-cosine harmonic function, is the timestamp.

6. The monocular visual odometry localization method based on end-to-end deep learning according to claim 5, characterized in that, In the step of performing implicit transformation on the high-dimensional mapping to obtain time pose information, it includes: Input the high-dimensional mapping into the T-Net and R-Net in the time pose network respectively. T-Net performs implicit transformation on the high-dimensional mapping, uses an MLP combined with ReLU activation function and tanh activation function, adjusts the output to 3 dimensions, and limits its value in the interval [-1,1] to output a translation vector; R-Net performs implicit transformation on the high-dimensional mapping, adjusts the output to 4 dimensions, represents the rotation in the form of a quaternion [qw, qx, qy, qz], performs tanh activation on the [qx, qy, qz] components, performs a 1 - tanh transformation on the [qw] component, and performs L2 normalization as a whole to construct a unit quaternion to output a rotation vector.

7. The monocular visual odometry localization method based on end-to-end deep learning according to claim 6, wherein In the step of extracting features from the input image to obtain the depth map of the input image, it includes: Input the input image into a pre-trained depth estimation network for feature extraction to generate the depth map corresponding to the input image.

8. The monocular visual odometry positioning method based on end-to-end deep learning according to claim 7, characterized in that, In the step of extracting the global attention of the depth map by using wavelet convolution operation and convolutional layer to obtain the global depth information, it includes: Input the depth map into the depth information extraction module, and use wavelet convolution operation to downsample the depth map, adjust the size of the depth map to H / 2×W / 2, and add all frequency bands along the channel dimension to make the number of channels become 4 times the original, so as to obtain a low-frequency subband, a first high-frequency subband, a second high-frequency subband, and a third high-frequency subband; where H is the height of the depth map and W is the width of the depth map; Use the convolutional layer to extract and reconstruct the low-frequency subband, the first high-frequency subband, the second high-frequency subband, and the third high-frequency subband respectively to obtain a depth map feature with a size of H / 2×W / 2×D; where D is the number of channels; Construct Query, Key, and Value of the attention module, and calculate the attention score in combination with the depth map feature to obtain the global depth information; Among them, the expression of the global depth information is: Among them, is the global depth information, is the first weight matrix, is the second weight matrix, is the third weight matrix, represents the wavelet convolution operation, represents the Sigmoid activation, is the depth map feature.

9. The monocular visual odometry positioning method based on end-to-end deep learning according to claim 8, characterized in that, In the step of performing multi-modal fusion on the initial pose estimation, the global depth information, and the temporal pose information to obtain the 6-DOF pose, it includes: Use the fusion module to perform linear weighted fusion on the translational initial estimation, the global depth information, and the translational vector to obtain a translational fusion feature; Use the fusion module to perform linear weighted fusion on the rotational initial estimation and the rotational vector to obtain a rotational fusion feature; Use the preset pose regression multi-layer perceptron in the fusion module to perform non-linear mapping on the translational fusion feature and the rotational fusion feature to obtain the 6-DOF pose.

Citation Information

Patent Citations

  • A monocular vision odometer method adopting deep learning and mixed pose estimation

    CN111899280A

  • Monocular vision odometer method fusing deep learning and geometric reasoning

    CN112906766A

  • Monocular visual odometer method based on Kalman pose estimation network

    CN114663496A

  • Visual odometer method and system based on image depth prediction and monocular geometry

    CN118736009A

  • End-to-end monocular visual odometer method fusing space-time semantic information

    CN120088332A

Cited By

  • Self-supervised monocular visual odometer method based on optical flow guidance and dynamic mask

    CN121437629A