Depth geometric consistency adaptive optimization method

By employing a depth geometric consistency adaptive optimization method, image reprojection is performed using camera pose and intrinsic parameters, and depth estimation is optimized by combining a loss function. This solves the scale inconsistency problem of depth prediction results in monocular videos and improves the accuracy and stability of the visual SLAM system.

CN120953339APending Publication Date: 2025-11-14JIANGSU UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511132351.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing monocular video self-supervised depth estimation methods lack camera motion information, resulting in scale ambiguity in depth prediction results and an inability to maintain consistency in video applications, especially affecting the accuracy of camera tracking in visual SLAM systems.

Method used

By designing an adaptive optimization method for depth geometric consistency, image reprojection is performed using camera pose and intrinsic parameters. By combining SmoothL1 and SSIM loss functions, a spatial geometric consistency mask is calculated, a spatial geometric consistency loss function is constructed, and the depth estimation results are optimized.

Benefits of technology

It effectively solves the scale inconsistency problem in monocular depth estimation systems, improves the performance and system stability of depth estimation models in video applications, and is particularly suitable for camera tracking in visual SLAM systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953339A_ABST
    Figure CN120953339A_ABST
Patent Text Reader

Abstract

The invention discloses a depth geometric consistency adaptive optimization method. The method comprises the following steps: selecting a source image, a target image and camera internal parameters; calculating an initial depth result of the given image and calculating a relative camera pose between the source image and the target image; according to the camera pose and the internal reference, back-projecting the initial depth result of the source image to a target image space to obtain a back-projection depth result, and interpolating the back-projection depth result to obtain a conversion depth result aligned with the target image depth result; calculating a space geometric consistency mask; determining a motion trend of visible pixels between the target image and the source image; and based on the space geometric consistency mask, the back projection depth result and the transformation depth result, constructing a space geometric consistency loss function for adaptive optimization of depth geometric consistency. According to the method, the geometric consistency in video depth estimation is enhanced on the premise of avoiding extra information, and the problem of scale inconsistency in a prediction result of a monocular depth estimation system is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a deep geometric consistency adaptive optimization method. Background Technology

[0002] In the field of self-supervised monocular depth estimation, the loss function is designed by combining the depth predicted by the depth model with the six-DOF pose predicted by the camera pose model, using interpolation techniques to deform and generate another frame, and then supervising model training by comparing the photometric reprojection loss between the real first frame and the transformed first frame. However, compared to scene depth estimation methods based on stereo matching, self-supervised depth estimation techniques based on monocular videos lack prior camera motion information, leading to a common problem of scale ambiguity in the predicted depth results—that is, the predicted depth is unknown in scale to the real world. Due to scale ambiguity across different frames, most existing monocular video-based methods typically produce scale-inconsistent predictions. While these predictions do not affect single-image-based visual tasks, they are crucial for some video-based applications, such as camera tracking in visual SLAM systems. Summary of the Invention

[0003] This invention provides a depth geometric consistency adaptive optimization method to address the problems existing in the prior art.

[0004] The technical solutions adopted in this invention are as follows:

[0005] A deep geometry consistency adaptive optimization method includes the following steps:

[0006] S1: Select source image, target image, and camera intrinsic parameters

[0007] S2: Calculate the initial depth results of the source image and the target image using the depth estimation model, and calculate the relative camera pose between the source image and the target image using the camera pose estimation model;

[0008] S3: Based on the camera pose and intrinsic parameters, backproject the initial depth result of the source image to the target image space to obtain the backprojected depth result, and interpolate the backprojected depth result using a differentiable bilinear interpolation algorithm to obtain a transformed depth result aligned with the depth result of the target image.

[0009] S4: Compute a spatially geometrically consistent mask, including:

[0010] The loss function is calculated by minimizing the photometric reprojection loss pixel by pixel using SmoothL1 and SSIM.

[0011] A static mask is obtained through an automatic masking method to detect static pixels that do not have relative motion between the target image and the source image;

[0012] Define an occlusion mask to determine the motion trend of visible pixels between the target image and the source image;

[0013] S5: Based on the spatial geometric consistency mask, back-projection depth results, and transformed depth results, construct a spatial geometric consistency loss function for adaptive optimization of depth geometric consistency.

[0014] Furthermore, S3 specifically refers to:

[0015] Using camera pose Depth of the target image Projected into 3D space, and then projected onto the source image via the camera's intrinsic parameter K. The image plane is used to obtain the backprojection depth results. ;

[0016] For the transformation depth result Using a differentiable bilinear interpolation algorithm, from Interpolation Its formula is expressed as:

[0017] ,

[0018] ,

[0019] in This represents locally differentiable bilinear sampling. Indicates the position of the camera The two-dimensional coordinates projected onto the camera's intrinsic parameter K. This indicates pixel-by-pixel multiplication, where each pixel... For pixels in a two-dimensional image.

[0020] Furthermore, the specific process of calculating the spatial geometric consistency mask in S4 is as follows:

[0021] A spatially geometrically consistent mask is calculated using the target image, the source image, the corresponding initial depth results, and the camera pose.

[0022] First, define The loss function, which is a pixel-wise minimization of photometric reprojection loss using SmoothL1 and SSIM, is formulated as follows:

[0023] ,

[0024] in, It is a constant parameter.

[0025] Furthermore, the static mask μ solution function is:

[0026] ,

[0027] in, This represents Iverson brackets, which take the value 1 when the condition inside the brackets is met, and 0 otherwise.

[0028] The source view that is transformed during the i-th update is defined as:

[0029] ,

[0030] Where K represents the camera intrinsic parameter matrix, This represents locally differentiable bilinear sampling. express When updating the projection depth for the i-th time... 2D coordinates.

[0031] Furthermore, define the occlusion mask O. t−1 and O t+1 Determine the target image I t With source image I t-1 and I t+1 The movement trend of visible pixels:

[0032] ,

[0033] .

[0034] Furthermore, in S5, the spatial geometric consistency loss function L gc The calculation formula is:

[0035] ,

[0036] ,

[0037] in, This indicates pixel-by-pixel multiplication;

[0038] Diff t′ (p): Represents the result of the transformation depth. With backprojection depth results The pixel-by-pixel depth difference between the two is used to measure the inconsistency between them at pixel p; O t′ It indicates a mask or shield.

[0039] The present invention has the following beneficial effects:

[0040] 1. By designing a spatial geometric consistency loss function, the depth estimation results are ensured to maintain geometric consistency in video sequences, effectively solving the common scale inconsistency problem in the prediction results of monocular depth estimation systems. This enables more accurate depth estimation and improves the performance of the depth estimation model in video applications.

[0041] 2. Particularly suitable for video-based applications, such as camera tracking in visual SLAM systems, it can improve the stability and accuracy of the system.

[0042] 3. A new loss function construction method is introduced, which effectively captures and corrects geometric inconsistencies in depth estimation by comparing the differences between the transformed depth results and the back-projected depth results.

[0043] 4. Improves the geometric consistency of depth estimation without using additional information (such as stereo image pairs or depth sensor data), and reduces the complexity and cost of the system. Attached Figure Description

[0044] Figure 1 This is a schematic diagram illustrating the construction of the spatial geometric consistency loss function. Detailed Implementation

[0045] The invention will now be further described with reference to the accompanying drawings.

[0046] like Figure 1 This invention discloses a depth geometric consistency adaptive optimization method, the goal of which is to perform image reprojection based on the estimated depth and the relative camera pose between image frames, i.e., given a pixel... All observations should be mapped to a single common 3D point in the world coordinate system. And it won't drift.

[0047] Specifically, in order to achieve spatial geometric consistency modeling, a differentiable spatial depth geometric inconsistency calculation strategy is introduced to calculate the pixel-by-pixel inconsistency between two depth maps.

[0048] In the picture Indicates the source image Using depth prediction results and camera pose Depth map obtained through rigid transformation It is a depth map The depth map obtained through differentiable bilinear interpolation aims to... Alignment is performed. The entire spatial depth geometry inconsistency process can be described as follows:

[0049] First, using camera pose Depth Projected into three-dimensional space, and then projected onto... Image plane to obtain The purpose of the entire computational strategy is to compute... and The difference between them is determined using differentiable bilinear interpolation, from Interpolation And let it be with Alignment is performed, and finally, a spatial geometric consistency mask is used for comparison. and We identify depth inconsistencies to construct a spatial geometric consistency loss function. Specifically:

[0050] Step 1: Select the target image, source image, and camera intrinsic parameters, denoted as I. t I t′ K

[0051] Step 2: Calculate the initial depth results of the given target image and source image using the depth estimation model, and denot them as follows: and Furthermore, it uses a camera pose estimation model to calculate the relative camera pose between the source image and the target image. ;

[0052] Step 3: Using the camera pose and intrinsic parameters, backproject the initial depth result of the source image into the target image space to obtain the backprojected depth result and the transformed depth result from the source image's viewpoint. The process is as follows:

[0053] First, using camera pose Depth of the target image Projected into 3D space, and then projected onto the camera intrinsic parameter K. Image plane to obtain backprojection depth results .

[0054] For the transformation depth result Then, a differentiable bilinear interpolation algorithm is used to calculate the result from... Interpolation And let it be with Alignment, the above process can be formalized into the following formula:

[0055] ,

[0056] ,

[0057] in, This represents locally differentiable bilinear sampling. Indicates the position of the camera The two-dimensional coordinates projected onto the camera's intrinsic parameter K. This indicates pixel-by-pixel multiplication, where each pixel... For pixels in a two-dimensional image.

[0058] Step 4: Calculate the spatial geometric consistency mask using the target image, source image, corresponding initial depth results, and camera pose.

[0059] First, define The loss function, which is a pixel-wise minimization of photometric reprojection loss using SmoothL1 and SSIM, can be formulated as follows:

[0060] ,

[0061] in, This is a constant parameter, set to 0.85.

[0062] µ is a static mask obtained using an automatic masking method. Its function is to detect static pixels in the target image and the source image that do not have relative motion. It can be defined as:

[0063] ,

[0064] here This represents an Iverson bracket, which takes the value 1 when the condition within the bracket is met, and 0 otherwise.

[0065] This represents the source view that is transformed during the i-th update, such as:

[0066] ,

[0067] Where K represents the camera intrinsic matrix, which is shared by all input images. This represents locally differentiable bilinear sampling. express When updating the projection depth for the i-th time... 2D coordinates.

[0068] In theory, the reprojection loss of visible pixels between adjacent frames in an unoccluded region should be less than the loss of occluded pixels. Therefore, an occlusion mask is defined. Its function is to determine the motion trend of visible pixels between the target image It and the source image It-1:

[0069] ,

[0070] Similarly, to predict the trend of visible pixels between the target image It and the source image It+1, an occlusion mask is defined. as follows:

[0071] ,

[0072] The spatially geometrically consistent mask, including the occlusion mask, can be obtained through the above process. , And static mask µ.

[0073] Step 5: Construct the spatial geometric consistency loss function L using the spatial geometric consistency mask, back-projection depth results, and transformed depth results. gc The following equation is used for adaptive optimization of depth geometry consistency:

[0074] ,

[0075] ,

[0076] in, This indicates pixel-by-pixel multiplication. t′ (p): Represents the result of the transformation depth. With backprojection depth results The pixel-by-pixel depth difference between the two is used to measure the inconsistency between them at pixel p; O t′ It indicates a mask or shield.

[0077] By minimizing the spatial geometric consistency loss function Lgc for batch samples, the model can naturally propagate geometric consistency throughout the sequence, enabling the self-supervised depth estimation model to estimate depths with geometric consistency.

[0078] The above description is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several improvements without departing from the principle of the present invention, and these improvements should also be considered within the scope of protection of the present invention.

Claims

1. A depth geometric consistency adaptive optimization method, characterized in that: Includes the following steps: S1: Select source image, target image, and camera intrinsic parameters S2: Calculate the initial depth results of the source image and the target image using the depth estimation model, and calculate the relative camera pose between the source image and the target image using the camera pose estimation model; S3: Based on the camera pose and intrinsic parameters, backproject the initial depth result of the source image to the target image space to obtain the backprojected depth result, and interpolate the backprojected depth result using a differentiable bilinear interpolation algorithm to obtain a transformed depth result aligned with the depth result of the target image. S4: Compute a spatially geometrically consistent mask, including: The loss function is calculated by minimizing the photometric reprojection loss pixel by pixel using SmoothL1 and SSIM. A static mask is obtained through an automatic masking method to detect static pixels that do not have relative motion between the target image and the source image; Define an occlusion mask to determine the motion trend of visible pixels between the target image and the source image; S5: Based on the spatial geometric consistency mask, back-projection depth results, and transformed depth results, construct a spatial geometric consistency loss function for adaptive optimization of depth geometric consistency.

2. The depth geometric consistency adaptive optimization method as described in claim 1, characterized in that: S3 specifically refers to: Using camera pose Depth of the target image Projected into 3D space, and then projected onto the source image via the camera's intrinsic parameter K. The image plane is used to obtain the backprojection depth results. ; For the transformation depth result Using a differentiable bilinear interpolation algorithm, from Interpolation Its formula is expressed as: , , in This represents locally differentiable bilinear sampling. Indicates the position of the camera The two-dimensional coordinates projected onto the camera's intrinsic parameter K. This indicates pixel-by-pixel multiplication, where each pixel... For pixels in a two-dimensional image.

3. The depth geometric consistency adaptive optimization method as described in claim 1, characterized in that: The specific process of calculating the spatial geometric consistency mask in S4 is as follows: A spatially geometrically consistent mask is calculated using the target image, the source image, the corresponding initial depth results, and the camera pose. First, define The loss function, which is a pixel-wise minimization of photometric reprojection loss using SmoothL1 and SSIM, is formulated as follows: , in, It is a constant parameter.

4. The depth geometric consistency adaptive optimization method as described in claim 3, characterized in that: The static mask μ solution function is: , in, This represents Iverson brackets, which take the value 1 when the condition inside the brackets is met, and 0 otherwise. The source view that is transformed during the i-th update is defined as: , Where K represents the camera intrinsic parameter matrix, This represents locally differentiable bilinear sampling. express When updating the projection depth for the i-th time... 2D coordinates.

5. The depth geometric consistency adaptive optimization method as described in claim 3, characterized in that: Define occlusion mask O t−1 and O t+1 Determine the target image I t With source image I t-1 and I t+1 The movement trend of visible pixels: , 。 6. The depth geometric consistency adaptive optimization method as described in claim 1, characterized in that: In S5, the spatial geometric consistency loss function L gc The calculation formula is: , , in, This indicates pixel-by-pixel multiplication; Diff t′ (p): Represents the result of the transformation depth. With backprojection depth results The pixel-by-pixel depth difference between the two is used to measure the inconsistency between them at pixel p; O t′ It indicates a mask or shield.