Self-supervised monocular visual odometer method based on brightness consistency and homography constraint

By introducing brightness consistency and homography constraints in the self-supervised monocular visual odometer method, the local minimum value and scale blur problems caused by photometric error are solved, and a higher accuracy and stable visual odometer is achieved.

CN119991782APending Publication Date: 2025-05-13BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510058354.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing self-supervised monocular visual odometry method is prone to fall into local minimum values ​​when using photometric errors as a supervision signal, and there are problems such as blurring in scale and inconsistent in front and back frame brightness.

Method used

A self-supervised monocular visual odometry method based on brightness consistency and homography constraints is proposed. The depth and relative camera attitude of image pairs are estimated by the depth estimation network and the posture estimation network, and the brightness alignment and homography estimation are performed. Combining structural similarity, smoothness and geometric consistency loss functions, network prediction is optimized.

Benefits of technology

Through brightness alignment and homography constraints, the error caused by photometric error is reduced, local minimum values ​​are avoided, the accuracy and stability of the visual odometer are improved, and error indicators such as average absolute trajectory error are significantly reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991782A_ABST
    Figure CN119991782A_ABST
Patent Text Reader

Abstract

The invention provides a brightness consistency and homography constraint-based self-supervised monocular vision odometer method. The method comprises the following steps of: training a network by using a luminosity error as a supervised signal in the self-supervised monocular vision odometer method; according to the mode, the system is likely to fall into the local minimum value, and the problem of falling into the local minimum value is solved by using a new homography loss. Meanwhile, the self-supervised monocular visual odometer has the problem of scale fuzziness, and the overall consistent scale is kept by using geometric consistency loss. Since brightness levels in adjacent frames are different due to continuous change of a visual angle, change of ambient lighting or movement during camera exposure adjustment, brightness alignment is carried out through image blocking and image block brightness alignment, and the image after brightness alignment is applied to training of the self-supervised monocular visual odometer, so that the self-supervised monocular visual odometer training efficiency is improved, and the self-supervised monocular visual odometer training efficiency is improved. And finally, a result similar to the best method at present is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visual odometers, and in particular to a self-supervised monocular visual odometer method based on brightness consistency and homography constraints. Background Art

[0002] Visual odometry (VO) is a technique that estimates camera motion (including displacement and rotation) by analyzing the changes between consecutive image frames. It is used in many practical applications, such as autonomous driving, augmented reality, robot navigation, visual SLAM systems, etc. Many studies have proposed visual odometry using self-supervised methods. Self-supervised methods can use only monocular image sequences or video frames to synthesize target images by predicting depth information and camera self-motion, and use photometric errors as supervisory signals to train the network.

[0003] However, the use of photometric error as a supervisory signal is likely to cause the visual odometer system to fall into a local minimum. Weak texture areas or feature-consistent areas in the image will cause the photometric error to deviate. At the same time, the depth predicted by the monocular self-supervision method is not the true depth, and has an unknown scaling for the real world, that is, there is a scale ambiguity problem. And due to the constant change of viewing angle, changes in ambient lighting, or movement during camera exposure adjustment, these factors may cause different brightness levels in adjacent frames, introducing additional noise. These problems are all technical problems to be solved by the present invention. Summary of the invention

[0004] In order to solve the problem that the photometric error may cause the training of the visual odometer system to fall into a local minimum when used as a supervisory signal, the scale of the self-supervised method is fuzzy, and the brightness of the previous and next frames is not uniform. The present invention proposes a method for a self-supervised monocular visual odometer system based on brightness consistency and homography constraints, comprising the steps of:

[0005] Step S1: Given an image pair with a relative motion relationship (I t , I t+1 ), where I t is the image at time t, I t+1 For the image at time t+1, we first use the depth estimation network and the pose estimation network to estimate their corresponding depth map D t , D t+1 and the relative camera pose

[0006] Step S2: For image I t+1 The pixel p t+1 , according to the depth value D estimated by the pixel t+1 (p t+1 ) and the camera intrinsic parameter K to restore the corresponding spatial 3D coordinates. The intrinsic parameter K is obtained according to the specific parameters of the camera:

[0007]

[0008] where f x , f y are the focal lengths of the x-axis and y-axis, c x , c y is the origin of the image coordinate system.

[0009] According to the following equation, we can get pixel p t+1 in I t The corresponding projected pixel p′ on t+1 , K -1 is the inverse matrix of the internal parameter K:

[0010]

[0011] Step S3: I t+1 (p t+1 ) after reprojection is I′ t+1 (p′ t+1 ), since the target pixel may have non-integer coordinates after reprojection, a differentiable bilinear interpolation algorithm is used to convert I t The pixel value of I′ is interpolated to t+1 The non-integer coordinates (p′ t+1 )middle:

[0012] I t (p′ t+1 )=I t [p t ]

[0013] Among them [p t ] indicates that I t (p t ) pixel values ​​around the pixel are bilinearly interpolated to obtain the non-integer coordinates (p′ t+1 ) corresponds to the pixel value I t (p′ t+1 ).

[0014] Step S4: I′ t+1 and I t Divide into 20 small blocks, each of which has the same size. Align the brightness of each block. First, calculate the average brightness of each block, and then calculate the brightness adjustment coefficient α corresponding to each block. i :

[0015]

[0016] Among them I t [i] represents image I t The i-th image block in t+1[i] represents image I′ t+1 The i-th image in the image, ∈ is a small constant, which we set to 10 -8 , used to avoid division by zero. Adjust the brightness of the target image block and set I′ t+1 Multiply each image by the corresponding brightness adjustment coefficient to obtain the final adjusted image I″ t+1 :

[0017] I″ t+1 [i] = I′ t+1 [i] α i

[0018] Step S5: Define the synthetic image I″ using structural similarity (SSIM) and L1 norm t+1 and reference image I t The luminosity loss between:

[0019]

[0020] Where μ0=0.15, μ1=0.85, |v| is the number of valid points of successful projection. The valid points of successful projection are defined as whether the average value of each pixel in the color channel is greater than 1e-3. A valid point is defined as one greater than 1e-3.

[0021]

[0022] Among them, μ x , μ y denote the means of x and y, σ x and σ y denote the variance of x and y, σ xy Represents the covariance of x and y. C1 and C2 are constants with default values ​​of 0.0001 and 0.0009 to avoid the case where the denominator is 0.

[0023] Step S6: Calculate edge-aware smoothing loss L S :

[0024]

[0025] where Δ is the first-order derivative along the spatial direction, is an exponential term, where I t The gradient at pixel p is Indicates D t The gradient at pixel p.

[0026] Step S7: Use geometric consistency loss to constrain the network to predict scale-consistent depth. First, the depth map D is calculated by a differentiable depth inconsistency operation. tWith D t+1 The pixel difference between D t+1 Reprojection to get D′ t+1 Since the target pixel may have non-integer coordinates after reprojection, a differentiable bilinear interpolation algorithm is used to calculate D t Interpolation is performed, and the interpolated D t Use ΔD t Denotes the depth difference D diff The calculation formula is summarized as:

[0027]

[0028] Step S71: Set the geometric consistency L G Defined as:

[0029]

[0030] |v| is the set of valid points for successful projection.

[0031] Step S8: Using the depth difference D diff Define the mask M s For removing moving objects:

[0032] M s =1-D diff

[0033] Mask-weighted photometric loss for:

[0034]

[0035] M s (p) represents the mask value at pixel p, L P (p) represents the photometric loss value at pixel p, and |v| is the valid point set of successful projection.

[0036] Step S9: I t and I″ t+1 After passing through the feature extraction module f(·), the feature map is generated and Mask pair feature map generated by mask prediction module and To perform weighting:

[0037]

[0038]

[0039] Get the weighted feature map and Then the weighted feature map and As input to the homography estimation network, four 2D offset vectors (8 values) are generated as output, and finally a homography matrix with 8 DOF is solved through direct linear transformation The homography loss is calculated as:

[0040]

[0041] in is the homography matrix obtained by the homography estimation module, and I is the unit matrix.

[0042] Step S10: The final loss function is:

[0043]

[0044] in Represents passing through M s Weighted L P The luminosity loss after S represents the smoothness loss, L G is the geometric consistency loss, L H is the homography loss. [α, β, γ, θ] are the loss weighting terms, and α=1, β=0.1, γ=0.1, θ=0.01 are determined through experiments.

[0045] The present invention adjusts the image brightness first, unifies the brightness of the image pairs that need to calculate the loss, and reduces the error caused by inconsistent brightness. The homography matrix between the image pairs is estimated by using a homography estimation network, and the similarity of the two images is judged by calculating whether the homography matrix is ​​close to the unit matrix. The homography estimated by the homography estimation network fully considers most of the areas that best represent the image motion, avoiding errors caused by occlusion and other problems. At the same time, the homography constraint provides another image similarity evaluation, helping the self-supervised visual odometer to get rid of local minima. The present invention realizes a more accurate self-supervised visual odometer. The test results on the KITTI dataset show that the best effect is shown in error indicators such as the mean absolute trajectory error. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is the core flow chart of the present invention;

[0047] Figure 2 Flowchart of the homography estimation network;

[0048] Figure 3 Schematic diagram of brightness difference;

[0049] Figure 4 Qualitative analysis of the KITTI dataset 09 sequence images;

[0050] Figure 5 Qualitative analysis of 10 image sequences for the KITTI dataset. DETAILED DESCRIPTION

[0051] The technical solution of the present invention is further described below in conjunction with the accompanying drawings, but is not limited thereto. Any modification or equivalent replacement of the technical solution of the present invention without departing from the spirit and scope of the technical solution of the present invention should be included in the protection scope of the present invention.

[0052] The present invention is based on the traditional self-supervised visual odometer calculation method and provides a self-supervised monocular visual odometer based on brightness consistency and homography constraints. The method includes the following contents:

[0053] 1. Luminosity loss

[0054] In the absence of labeled data, the self-supervised method is designed as a view synthesis problem. The synthesis of views involves the mutual transformation between 2D pixels and 3D points, and between 3D points and 3D points. This process requires corresponding depth information and self-motion data. The loss function is usually designed as the difference between the target image and the synthesized image. Specifically, for a given image pair with a relative motion relationship (I t , I t+1 ), where I t is the image at time t, I t+1 For the image at time t+1, we first use the depth estimation network and the pose estimation network to estimate their corresponding depth map D t , D t+1 and the relative camera pose The depth estimation network uses a ResNet50 encoder to extract features. The decoder is DispNet. The pose estimation network uses a ResNet18 encoder to extract features and modifies the first layer to six channels. The features are then decoded into 6-DoF parameters through four convolutional layers.

[0055] For image I t+1 The pixel p t+1 , according to the depth value D estimated by the pixel t+1 (p t+1 ) and the camera intrinsic parameter K to restore the corresponding spatial 3D coordinates. The intrinsic parameter K is obtained according to the specific parameters of the camera:

[0056]

[0057] where f x , f y are the focal lengths of the x-axis and y-axis, c x , c y is the origin of the image coordinate system.

[0058] According to the following equation, we can get pixel p t+1 in I t The corresponding projected pixel p′ on t+1 , K -1 is the inverse matrix of the internal parameter K:

[0059]

[0060] I t+1 (p t+1 ) after reprojection is I′ t+1 (p′ t+1 ), in the process of reprojection, non-integer coordinates of the target pixel may be obtained. Since the pixel grid of the image is discrete, this non-integer coordinate cannot be directly mapped to an actual pixel position. The differentiable bilinear interpolation algorithm can be used to map the I t The pixel value of I′ is interpolated to t+1 The non-integer coordinates (p′ t+1 )middle:

[0061] I t (p′ t+1 )=I t [p t ]

[0062] Among them [p t ] indicates that I t (p t ) pixel values ​​around the pixel are bilinearly interpolated to obtain the non-integer coordinates (p′ t+1 ) corresponds to the pixel value I t (p′ t+1 ).

[0063] The photometric loss needs to satisfy the grayscale invariance assumption. In real outdoor scenes, the ambient lighting conditions are constantly changing, and changes in camera exposure will cause changes in image brightness, which will violate this assumption. We also found that the brightness differences between different regions in two adjacent frames are not the same. Figure 3 Shown are image pairs from the KITTI dataset that exhibit different brightness differences.

[0064] Before calculating the luminosity loss, I′ t+1 Perform brightness alignment and set I′ t+1 and I t Divide into 20 small blocks, each of which has the same size. Align the brightness of each block. First, calculate the average brightness of each block, and then calculate the brightness adjustment coefficient α corresponding to each block. i :

[0065]

[0066] Among them I t [i] represents image I t The i-th image block in t+1 [i] represents image I′ t+1 The first block image in ∈ is a small constant, which we set to 10 -8 , used to avoid division by zero. Adjust the brightness of the target image block and set I′ t+1 Multiply each image by the corresponding brightness adjustment coefficient to obtain the final adjusted image I″ t+1 :

[0067] I″ t+1 [i] = I′ t+1 [i] α i

[0068] Finally, the structural similarity (SSIM) and L1 norm are used to define the synthetic image I″ t+1 and the real image I t+1 The luminosity loss between:

[0069]

[0070] where μ0 and μ1 are manually set hyperparameters and V is the set of valid points that are successfully projected and interpolated.

[0071]

[0072] Among them, μ x , μ y denote the means of x and y, σ x and σ y denote the variance of x and y, σ xy Represents the covariance of x and y. C1 and C2 are constants with default values ​​of 0.0001 and 0.0009 to avoid the situation where the denominator is 0.

[0073] Calculate the edge-aware smoothing loss L S :

[0074]

[0075] in is the first-order derivative along the spatial direction, is an exponential term, where I t The gradient at pixel p is Indicates D t The gradient at pixel p.

[0076] 2. Geometric consistency loss and masking

[0077] Since the monocular system does not have the ability to directly obtain depth information, the depth information is generally predicted by the depth estimation network. However, this method may generate inconsistent scale predictions on different image frames, resulting in scale ambiguity. For continuous image frames or videos, maintaining scale consistency is crucial to the prediction accuracy of the self-supervised visual odometer system. The present invention uses a geometric consistency loss to constrain the network to predict the scale-consistent depth. Specifically, the predicted two-frame depth map D is calculated by a differentiable depth inconsistency operation. t With D t+1 The pixel difference between them is consistent with the synthetic image method of calculating the photometric loss, D t+1 After reprojection, we get D′ t+1 , since the pixel position after reprojection is not strictly located in D t In the grid, D t Perform differentiable bilinear interpolation and obtain the same value as D′ t+1 The pixel value at the same pixel position, the interpolated D t Use ΔD t Depth difference D diff The calculation formula is summarized as:

[0078]

[0079] Normalizing the depth difference can make the output between 0 and 1, which is more conducive to training. The geometric consistency is defined as:

[0080]

[0081] By minimizing the depth inconsistency of images in a batch, we achieve depth consistency of the batch, and apply this method to the training of the entire image sequence, and finally achieve depth consistency of the entire sequence, so that the image frames of a sequence have the same scale prediction.

[0082] Using the depth difference D diff Define the mask M s For removing moving objects:

[0083] M s =1-D diff

[0084] The mask-weighted photometric loss is:

[0085]

[0086] M s (p) represents the mask value at pixel p, L P (p) represents the luminosity loss value at pixel p.

[0087] 3. Homography Loss

[0088] Due to the object occlusion problem and inconsistent light brightness as the viewing angle moves, the convergence values ​​of different frames are different when calculating the photometric error. In addition, the weak texture areas or feature-consistent areas in the image will cause the photometric error to deviate. Using only the photometric error as a supervisory signal is likely to cause the VO system to fall into a local minimum.

[0089] The existence of these problems inspires us to rethink how to evaluate the similarity between the reprojected synthetic image and the target frame image. We should ignore the existence of these small interference items and focus on the internal points that can express the motion, that is, the content that best represents the image motion.

[0090] Based on the above considerations, the present invention uses the homography corresponding to two frames of images as the loss function of image similarity, ignores the interference caused by moving objects and occlusions, and only uses the inner points that best represent the perspective movement to measure the similarity.

[0091] The homography estimation network is as follows Figure 2 As shown, there are two image blocks I t and I″ t+1 As input, it will be able to characterize I t and I″ t+1 The homography matrix of the corresponding relationship as output. The entire structure can be divided into three modules: a feature extractor module f(·), a mask prediction module m(·) and a homography estimation module h(·). f(·) and m(·) are fully convolutional networks. The feature extraction module is a 3-layer fully convolutional network with a convolution kernel of 3*3, a step size of 1, an initial number of channels of 4, and a final number of channels of 1. The mask prediction module is a 5-layer fully convolutional network with a convolution kernel of 3*3, a step size of 1, an initial number of channels of 4, and a final number of channels of 1. h(·) uses ResNet-34 as the backbone. The convolution kernel of the first layer is 7*7, and the remaining layers are 3*3. The initial number of channels is 64, and the final output dimension is 8.

[0092] I t and I″ t+1 After passing through the feature extraction module f(·) the feature map is generated and As the input of subsequent modules, features are usually more robust than pixel intensity and can overcome the impact of brightness changes.

[0093] A subnetwork is also used to handle the problem of moving objects. This subnetwork is used to estimate the mask to highlight the content that is valuable for homography estimation and cover the moving objects. The mask prediction module m(·) is a fully convolutional network, and the size of the mask generated is consistent with the output content of the feature extractor. The mask generated by the mask prediction module is used to predict the feature map. and To perform weighting:

[0094]

[0095] Get the weighted feature map and As the input of the homography estimation module h(·), the mask here is equivalent to an attention map, highlighting the important content and removing the distracting content.

[0096] The weighted feature map and As input to the homography estimation network, four 2D offset vectors (8 values) are generated as output, and finally a homography matrix with 8 DOF is solved through direct linear transformation

[0097] The homography matrix is ​​a 3*3 matrix used to describe the geometric transformation relationship between images at different viewing angles. If two images are similar enough, the homography matrix is ​​closer to the unit matrix, so the present invention uses the norm of the homography and the unit matrix as the result of measuring the similarity of images. The homography loss function can be described as:

[0098]

[0099] in is the homography matrix obtained by the homography estimation module, and I is the unit matrix.

[0100] 4. Total loss function

[0101] The total loss function is designed as:

[0102]

[0103] in Represents passing through M s Weighted L P The luminosity loss after S represents the smoothness loss, L G is the geometric consistency loss, L H is the homography loss. [α, β, γ, θ] is the loss weighting term. The hyperparameters are set to α=1, β=0.1, γ=0.1, θ=0.01.

[0104] 5. Experimental Procedure

[0105] The present invention uses the processed KITTI dataset provided by SC-Depth for training, wherein the 09-10 sequence images are used for testing. All experiments in the present invention use the PyTorch library and are performed on a high-performance server. The server is equipped with an AMD EPYC 9654 central processing unit (CPU) with 60GB of memory. In addition, the server is also equipped with an RTX 4090 graphics card with 24GB of video memory.

[0106] During the training process, three consecutive image frames are used as training samples, the projection and loss from the second frame to the other frames are calculated, and they are inverted again to maximize data usage. During the training process, the images are enhanced by random scaling, cropping, and horizontal flipping. The present invention uses the ADAM optimizer and sets the learning rate to 10 -4 .

[0107] Table 1 shows the quantitative analysis of the present invention and other visual odometry methods, all of which are performed in a self-supervised manner, where the best results are highlighted in bold, the suboptimal results are underlined in italics, and "OS" indicates whether it is open source work. The results show that the present invention achieves the best results in mean absolute trajectory error (ATE) and err and the relative rotation error r err In both items, good accuracy was also achieved, close to the current best results.

[0108] Figure 4 The qualitative comparison between the proposed method and the baseline method is shown, and the trajectory diagrams of two sequences of KITTI datasets 09 and 10 are plotted respectively. The blue curve is the trajectory diagram drawn by the true value, the green curve is the trajectory diagram of the pose estimated by the baseline algorithm, and the orange curve is the trajectory diagram of the pose estimated by the proposed method. It can be seen from the visualized trajectory that compared with the baseline method, the proposed method has better performance, which verifies the effectiveness of the proposed method.

[0109] In order to further prove the effectiveness of each component in the present invention, some ablation studies were conducted, as shown in Table 2. The first row is the result of the baseline algorithm, which was obtained by training with the open source model provided on GitHub, and the best reproduced result was selected. In order to verify the present invention, the homography prediction network was first added for training, and the results are shown in the second row of Table 2. The results show that the errors in each dimension are reduced after adding the homography loss, which also shows that compared with the baseline algorithm, the local minimum is eliminated and the algorithm accuracy is improved.

[0110] We also tested the results after aligning the photometric values, as shown in the third row of Table 2. The results show that after eliminating the impact of inconsistent photometric values, the errors in all dimensions have been improved.

[0111] The last row shows the result after adding homography loss and performing brightness alignment, which is better than the previous results in all dimensions. The results prove the effectiveness of each part of the present invention.

[0112] In order to verify the generalization of the strategy of using homography to evaluate image similarity in the self-supervised visual odometry method, the present invention also tests this strategy in the open source algorithm MotionHint.

[0113] Table 3 shows the quantitative test results. Since MotionHint has different training modes, we give the results of all modes here. We test the method based on the Unpaired training mode. The last line shows the results of adding homography loss to the method. The results show that after adding homography loss, the errors in all dimensions of the 09 and 10 sequences are reduced. The experiment proves that homography loss improves the self-supervised visual odometer method.

[0114] Table 1: Quantitative analysis with existing self-supervised pose estimation methods. “OS” indicates whether it is an open source algorithm. The best performance is shown in bold, and the suboptimal ones are shown in italics and underlined.

[0115]

[0116] Table 2: Quantitative results of ablation experiments

[0117]

[0118] Table 3 Hnet generalization analysis

[0119]

Claims

1. A self-supervised monocular visual odometry method based on brightness consistency and homography constraints, characterized in that: Includes steps: Step S1: Given an image pair with a relative motion relationship (I t , I t+1 ), where I t is the image at time t, I t+1 For the image at time t+1, we first use the depth estimation network and the pose estimation network to estimate their corresponding depth map D t , D t+1 and the relative camera pose Among them, the depth estimation network uses ResNet50 encoder to extract features, and the decoder uses DispNet; the pose estimation network uses ResNet18 encoder to extract features, and the first layer is modified to have six channels, and then the features are decoded into 6-DoF parameters through four convolutional layers; Step S2: Reprojection using the estimated depth information. t+1 The pixel p t+1 , according to the depth value D estimated by the pixel t+1 (p t+1 ) and the camera intrinsic parameter K to restore the corresponding spatial 3D coordinates. The intrinsic parameter K is obtained according to the specific parameters of the camera: where f x , f y are the focal lengths of the x-axis and y-axis, c x , c y is the origin of the image coordinate system; According to the following equation, we can get pixel p t+1 in I t The corresponding projected pixel p′ on t+1 , K -1 is the inverse matrix of the internal parameter K: Step S3: I t+1 (p t+1 ) after reprojection is I′ t+1 (p′ t+1 ), since the target pixel may have non-integer coordinates after reprojection, a differentiable bilinear interpolation algorithm is used to convert I t The pixel value of I′ is interpolated to t+1 The non-integer coordinates (p′ t+1 )middle: I t (p′ t+1 )=I t [p t ] Among them [p t ] indicates that I t (p t ) pixel values ​​around the pixel are bilinearly interpolated to obtain the non-integer coordinates (p′ t+1 ) corresponds to the pixel value I t (p′ t+1 ); Step S4: I′ t+1 and I t Divide into 20 small blocks, each of which has the same size. Align the brightness of each small block. First, calculate the average brightness of each small block, and then calculate the brightness adjustment coefficient α corresponding to each small block. i : Among them I t [i] represents image I t The i-th image block in t+1 [i] represents image I′ t+1 The i-th image block in ∈ is set to 10 -8 , used to avoid division by zero; adjust the brightness of the target image block and change I′ t+1 Multiply each image by the corresponding brightness adjustment coefficient to obtain the final adjusted image I″ t+1 : I″ t+1 [i]=I′ t+1 [i]·α i Step S5: Define the synthetic image I″ using structural similarity (SSIM) and L1 norm t+1 and reference image I t The luminosity loss between: Where μ0=0.15, μ1=0.85, |v| is the number of valid points of successful projection. The valid points of successful projection are defined as whether the average value of each pixel in the color channel is greater than 1e-3. If it is greater than 1e-3, it is defined as a valid point. Among them, μ x , μ y denote the means of x and y, σ x and σ y denote the variance of x and y, σ xy Represents the covariance of x and y, C1 and C2 are constants, with default values ​​of 0.0001 and 0.0009, which are used to avoid the situation where the denominator is 0; Step S6: Calculate edge-aware smoothing loss L S : in is the first-order derivative along the spatial direction, is an exponential term, where Indicates I t The gradient at pixel p is Indicates D t The gradient at pixel p; Step S7: Use geometric consistency loss to constrain the network to predict scale-consistent depth; first, the depth map D is calculated through a differentiable depth inconsistency operation t With D t+1 The pixel difference between D t+1 Reprojection to get D′ t+1 Since the target pixel may have non-integer coordinates after reprojection, a differentiable bilinear interpolation algorithm is used to calculate D t Interpolation is performed, and the interpolated D t Use ΔD t Denotes the depth difference D diff The calculation formula is summarized as: Step S71: Set the geometric consistency L G Defined as: |v| is the set of valid points for successful projection; Step S8: Using the depth difference D diff Define the mask M s For removing moving objects: M s =1-D diff Mask-weighted photometric loss for: M s (p) represents the mask value at pixel p, L P (p) represents the luminosity loss value at pixel p; |v| is the valid point set of successful projection; Step S9: The entire structure of the homography estimation network is divided into three modules: a feature extractor module f(·), a mask prediction module m(·) and a homography estimation module h(·); f(·) and m(·) are full convolutional networks, the feature extraction module is a 3-layer full convolutional network, the convolution kernel is 3*3, the step size is 1, the initial number of channels is 4, and the final number of channels is 1; the mask prediction module is a 5-layer full convolutional network, the convolution kernel is 3*3, the step size is 1, the initial number of channels is 4, and the final number of channels is 1; h(·) uses ResNet-34 as the backbone, the convolution kernel of the first layer is 7*7, the remaining layers are 3*3, the initial number of channels is 64, and the final output dimension is 8; I t and I″ t+1 After passing through the feature extraction module f(·) the feature map is generated and Mask pair feature map generated by mask prediction module and To perform weighting: Get the weighted feature map and Then the weighted feature map and As the input of the homography estimation network, four 2D offset vectors are generated as output, and finally a homography matrix with 8 DOF is solved through direct linear transformation The homography loss is calculated as: in is the homography matrix obtained by the homography estimation module, and I is the unit matrix; Step S10: The final loss function is: in Represents passing through M s Weighted L P The luminosity loss after S represents the smoothness loss, L G is the geometric consistency loss, L H is the homography loss; [α, β, γ, θ] is the loss weighting term; The final loss function is used to train the pose estimation network and the depth estimation network, and the learning rate is set to 10 -4 ; The trained pose estimation network takes two adjacent frame images as input and the relative pose as output. The accumulated relative pose is used as the trajectory of the image sequence for odometer applications.