Self-supervised monocular depth estimation method based on parameterized geometric representation
By modeling the three-dimensional scene as a plane set, using the geometric mapping relationship between plane normals and offsets, the depth discontinuity problem in the self-supervised monocular depth estimation method is solved, and the accuracy and generalization ability of depth estimation are improved, which is suitable for complex scenarios and maintains real-time and efficientness.
Patent Information
- Application Number
- CN202510416982.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-04
Smart Images

Figure CN120259398A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a self-supervised depth estimation method based on parametric geometric representation. Background Art
[0002] With the rapid development of computer vision technology, monocular depth estimation has become a crucial research direction, aiming to predict a pixel-level depth map for a given image, where each depth value represents the distance from the corresponding point in the image to the camera plane. This technology can provide powerful information for 3D scene structures and is widely used in fields such as autonomous driving, robotics, and augmented reality. In practical applications, depth estimation technology can help the system better understand and perceive the environment, and achieve accurate recognition of spatial layout and object distances. Although supervised learning methods can provide high-performance depth prediction, obtaining dense depth labels usually requires high costs. Therefore, self-supervised monocular depth estimation methods have emerged as a practical alternative. This method only relies on videos or stereo images for training, avoiding the high cost of labeled data, and has gradually become a current research hotspot.
[0003] Existing self-supervised monocular depth estimation methods mostly adopt the "point-to-depth" modeling strategy, regarding image pixels as independent points and discretely representing and independently predicting the depth values of each point. Although this paradigm simplifies model design, it ignores the continuous geometric structure characteristics of objects in the three-dimensional scene, resulting in problems such as possible jumps in depth values in the same plane area and distorted predictions in low-texture areas. For example, in the autonomous driving scenario, non-continuous fluctuations in depth values often occur in large flat areas such as roads or walls, seriously affecting the reliability of obstacle ranging and path planning. To solve this problem, some studies have tried to introduce geometric prior constraints. For example, LEGO uses the surface normals of adjacent pixels to construct a local plane hypothesis and optimizes depth smoothness through plane consistency loss; GeoNet enhances geometric continuity by jointly predicting the normal and depth maps; StructDepth combines semantic segmentation constraints to optimize depth prediction in plane areas.
[0004] These methods only rely on surface normals as plane representations, violating the plane uniqueness principle in geometric algebra - a complete plane definition requires both a normal vector (representing the orientation) and an offset (defining the spatial position). This incomplete modeling leads to the accumulation of plane reconstruction errors, and especially in complex unstructured scenes, the depth discontinuity problem remains significant. Therefore, developing a robust modeling strategy that can capture the intrinsic geometric representation has become the key to solving the depth discontinuity problem. Summary of the Invention
[0005] Aiming at the problem that existing self-supervised depth estimation methods have limited ability to capture the intrinsic geometric representation from complex unstructured scenes for solving depth discontinuities, the present invention proposes a self-supervised monocular depth estimation method based on parametric geometric representation. Through an innovative "plane-to-depth" modeling strategy, it aims to solve the inherent depth discontinuity problem in "point-to-depth" methods. The method includes the following steps:
[0006] S1: Input the target image to be predicted for depth and the adjacent image into the depth estimation network to predict the plane normal and plane offset, and derive the depth map of the target image and the depth map of the adjacent image according to their geometric mapping relationship with depth;
[0007] S2: Input the target image and the adjacent image into the pose network to predict the relative pose parameters between them;
[0008] S3: Input the predicted target depth map, relative pose parameters, and the adjacent source image into the depth discontinuity perception module. First, generate a synthetic image according to projective geometry, and then compare the difference between the synthetic image and the target image. Map the region with a smaller difference as the depth discontinuity mask as the region of depth discontinuity;
[0009] S4: Input the plane normal, plane offset, depth map, and depth discontinuity mask into the structured plane generation module. This module first uses spatio-temporal geometric cues to respectively guide the two plane parameters to recover a relatively accurate plane scene structure, and then jointly optimizes them to ensure that the coplanar points recover a unified representation.
[0010] Further, the specific content of S1 is as follows:
[0011] First, based on the depth estimation network f d predict the plane normal and plane offset of the target image and the adjacent images {I c ,I n}:
[0012]
[0013] where, {N c ,O c} are the plane normal and plane offset of the target image, and {N n ,O n} are the plane normal and plane offset of the adjacent image; then obtain the depth map based on geometric mapping; specifically, define the two-dimensional image point p = [u, v] T in the input image, where u and v are the abscissa and ordinate of the image respectively, and its homogeneous coordinate representation is Based on the plane normal N(p) and plane offset O(p) of each two-dimensional image point p, obtain the corresponding depth map through the following depth derivation formula:
[0014]
[0015] Among them, D(p) is the pixel depth of the two-dimensional image point p, and K is the internal parameter matrix of the camera; based on the pixel depth of the two-dimensional image point p of the input image, the depth maps {D c , D n} of the corresponding target image and adjacent image are obtained.
[0016] Furthermore, the specific steps of S3 are as follows:
[0017] The generation process of the synthesized image is expressed by the following formula:
[0018] I n2c = I n <proj(D c , T c2n , K)>
[0019] Among them, I n2c represents the synthesized image, <·> represents the bilinear interpolation operation, proj(·, ·, ·) represents the projection operation, and T c2n represents the relative pose parameter; then the single-pixel photometric error is used to calculate the difference between the synthesized image and the real target image, and the pixels smaller than the discontinuity threshold are determined to be in the depth discontinuity region:
[0020]
[0021] Among them, μ is the balance coefficient factor, is the depth discontinuity region mask.
[0022] Furthermore, the structured plane generation module specifically includes three sub-modules: plane normal constraint, plane offset constraint, and joint constraint;
[0023] The specific steps of the plane normal constraint are as follows:
[0024] First, the depth map of the target image is used to restore the point cloud of the target view, and the adjacent points around each point are selected to generate adjacent normals, and the average value of the adjacent normals is taken as the spatial normal clue; based on the similarity between the predicted plane normal of the target image and the spatial normal clue, the normal constraint in the spatial dimension is obtained;
[0025] At the same time, based on the plane normal of the adjacent image and the rotation matrix of the relative pose parameter, it is transformed to the target view to obtain the observed normal N′ n ; then the temporal normal clue N n2c is obtained through projective transformation:
[0026] Nn2c = N' n <proj(D c , T c2n , K)>
[0027] Among them, T c2n represents the relative pose parameter; a normal constraint in the time dimension is obtained based on the similarity between the predicted target image plane normal and the temporal normal cue;
[0028] Finally, by combining the normal constraints in the spatial dimension and the time dimension, the final plane normal constraint is obtained;
[0029] The specific plane offset constraint is as follows:
[0030] First, based on the spatial normal cue, a plane offset cue in the spatial dimension is obtained through the depth derivation formula of the target image, and the single-pixel photometric error is used to calculate the difference between the plane offset cue in the spatial dimension and the predicted plane offset, obtaining the offset constraint in the spatial dimension;
[0031] Then, given the depth map of the target image, the observed depth in the adjacent view is obtained according to the projective transformation. Based on this observed depth, the observed plane offset of the adjacent image is obtained through the depth derivation formula of the adjacent image; the sampled plane offset O″ n (p):
[0032] O″ n (p) = O n <proj(D c , T c2n , K)>
[0033] A temporal offset constraint is obtained based on the similarity between the observed plane offset and the sampled plane offset;
[0034] Finally, by combining the offset constraints in the spatial dimension and the time dimension, the final plane offset constraint is obtained;
[0035] The joint constraint is jointly optimized by penalizing the first-order gradients of the plane normal and the plane offset.
[0036] Furthermore, the discontinuity threshold is specifically:
[0037]
[0038] Among them, H and W are the height and width of the image respectively.
[0039] Furthermore, the normal constraint in the spatial dimension is defined as:
[0040]
[0041] Among them, is a downsampling function, and cos represents cosine similarity. is the spatial normal clue;
[0042] The normal constraint in the time dimension is defined as:
[0043]
[0044] Combining the normal constraints in the spatial and time dimensions, the plane normal constraint is defined as:
[0045]
[0046] The offset constraint in the spatial dimension is defined as:
[0047]
[0048] Among them is the plane offset clue;
[0049] The offset constraint in the time dimension is defined as:
[0050]
[0051] where O′ n is the observed plane offset; Combining the offset constraints in the spatial and time dimensions, the plane offset constraint is defined as:
[0052]
[0053] The joint constraint is expressed as follows:
[0054]
[0055] Among them, represents the operation of finding the first-order gradient.
[0056] Furthermore, the plane offset clue is obtained through the following formula:
[0057]
[0058] The observed plane offset is obtained through the following formula:
[0059]
[0060] Furthermore, construct a loss function to jointly optimize the structured generation module and the depth discontinuity perception module:
[0061]
[0062] Among them, is the final loss function, and γ, δ, and ω are balance coefficients.
[0063] Furthermore, the spatial normal clue is obtained in the following manner:
[0064] The point cloud of the target view is P, and eight adjacent points around each point are selected to generate an initial normal where P z is the center point, P zx , P zy are adjacent points; then four initial normals are selected at 90-degree intervals and the average value is calculated as the spatial normal clue:
[0065]
[0066] Compared with the prior art, the beneficial technical effects of the present invention are as follows:
[0067] Depth estimation accuracy and generalization ability: By innovatively designing a "plane-to-depth" modeling strategy, the present invention models a complex 3D scene as a set of planes of different sizes. The 3D points on the same plane are modeled as a unified plane representation instead of discrete points, eliminating the depth jump phenomenon in traditional point-to-depth methods and ensuring the continuity and smoothness of depth estimation. Moreover, this modeling strategy can be applied to any complex 3D scene. Therefore, the solution of the present invention not only improves the accuracy of existing depth estimation methods but also demonstrates strong generalization ability for unseen cross-domain scenes.
[0068] Real-time performance and efficiency: The modeling strategy and modules designed in the present invention can be combined with any existing lightweight depth estimation framework, and all designs basically do not introduce new model parameters. Therefore, the present invention can be implemented based on advanced lightweight depth, thus ensuring the real-time performance and efficiency of this method. Description of the Drawings
[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0070] Figure 1 is a flowchart of a self-supervised depth estimation algorithm based on parametric geometric representation provided by an embodiment of the present invention.
[0071] Figure 2 is a structural schematic diagram of a structured plane generation module provided by an embodiment of the present invention. Detailed implementation manners
[0072] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0073] A self-supervised monocular depth estimation method based on parametric geometric representation proposed in this embodiment, and the specific steps of this method are as follows:
[0074] Step S1: Construct a "plane-to-depth" paradigm; model a complex 3D scene as a set of planes of different sizes, where each plane is represented by a unique set of parameters, namely the plane normal (indicating the plane direction) and the plane offset (defining the perpendicular distance from the camera center to the plane), and explore the geometric mapping relationship between these plane parameters and the depth;
[0075] Specifically, in three-dimensional space, a plane can be uniquely determined and represented by its normal vector n and offset o. Thus, a complete three-dimensional scene can be represented as a set of multiple planes:
[0076]
[0077] where M is the number of planes in the scene and k is the serial number of the k-th plane. According to the point-normal form definition, any three-dimensional point P located on the plane π={n k ,o k} must satisfy the following constraint relationship:
[0078]
[0079] where i is the serial number of the i-th plane and T represents the transpose operation.
[0080] The present invention further proposes that through the internal parameter matrix K of the camera, the three-dimensional point p can be projected into the image space to obtain the two-dimensional image point p=[u,v] T , where u and v are the abscissa and ordinate of the image respectively, and their homogeneous coordinate representation is Then the depth d of this three-dimensional point in the image can be calculated by the following geometric formula:
[0081]
[0082] Through the above modeling, it can be found that for pixel points on the same plane, their depth changes are only related to their two-dimensional pixel positions and are consistent with the plane they are on (i.e., the normal vector and offset). Therefore, the present invention proposes to directly estimate the structured plane representation from the input image, that is, the plane normal vector N(p) and offset O(p) at each pixel position, and derive the calculation formula for the pixel depth D(p) of the two-dimensional image point p accordingly, as follows:
[0083]
[0084] where N ∈ R 3×H×W , O ∈ R 1×H×W , D ∈ R 1×H×W are the plane normal, plane offset, and depth respectively. H and W are the height and width of the image respectively. Then, based on the pixel depth of the two-dimensional image point p, the depth map of the corresponding image is obtained.
[0085] Through the above modeling process, complex 3D scenes are modeled as a set of planes of different sizes, where each plane is represented by a unique set of parameters, namely the plane normal (indicating the plane direction) and the plane offset (defining the perpendicular distance from the camera center to the plane), and the geometric mapping relationship between these plane parameters and depth is explored.
[0086] Step S2: Based on the "plane-to-depth" paradigm, construct a self-supervised depth estimation framework based on parametric geometric representation, which includes a depth estimation network, a pose network, a depth discontinuity perception module, and a structured plane generation module, as Figure 1 shown.
[0087] Step S3: Input the target image to be predicted for depth and the adjacent image into the depth estimation network to predict the plane normal and plane offset, and derive the target depth map according to their geometric mapping relationship with depth.
[0088] Specifically, given the input target image and adjacent images {I c , I n}, a depth estimation network f d first predicts their corresponding plane normal and plane offset:
[0089]
[0090] where {N c , O c} are the plane normal and plane offset of the target image, and {N n , O n} are the plane normal and plane offset of the adjacent image; then, based on the depth derivation formula in step 1, the plane parameters can be mapped to the final depth map {D c , Dn};
[0091] Step S4: Input the target image and its adjacent images into the pose network to predict the relative pose parameters between them;
[0092] Specifically, given the input target image and adjacent images {I c , I n}}, a pose network f p is used to predict the relative pose parameters between them:
[0093] T c2n = f p (I c , I n )
[0094] The pose parameter T c2n = [R c2n | t c2n can project the image in the adjacent view to the target view. R and t are the rotation matrix and translation matrix in the pose parameters respectively, and c2n is the operation flow from the target view to the adjacent view.
[0095] Step S5: Input the predicted target depth map, relative pose parameters, and adjacent source image into the depth discontinuity perception module to first generate a synthetic image, compare the difference between the synthetic image and the target image, and map the region with a smaller difference to a mask as the depth discontinuity region;
[0096] Furthermore, the working process of the depth discontinuity perception module in Step S5 is as follows: Given the predicted target depth map, relative pose parameters, and adjacent source image, the depth discontinuity perception module can transform the source image in the adjacent view to the target view according to the projection geometry (Warp), thereby generating a synthetic image corresponding to the target image:
[0097] I n2c = I n <proj(D c , T c2n , K)>
[0098] where <·> represents the bilinear interpolation operation, and proj(·, ·, ·) represents the projection operation. Generally, depth mutations mainly occur in low-texture regions, and the photometric errors in these regions are small. Therefore, the single-pixel photometric error is used to calculate the difference between the synthetic image and the real target image, and then the pixels smaller than the discontinuity threshold are determined to be in the depth discontinuity region:
[0099]
[0100] where, μ is the intensity of a balance coefficient factor controlling the threshold, is the depth discontinuity region mask, and the region with a value of 1 indicates that there is a high probability of depth jump in this region, while the region with a value of 0 indicates that the depth in this region is probably accurate.
[0101] Step S6: Input the predicted plane normal, plane offset, depth map, and depth discontinuity mask from the target image and adjacent images into the structured plane generation module. This module first uses spatio-temporal geometric cues to separately guide the two plane parameters to recover a relatively accurate plane scene structure, and then jointly optimizes them to ensure that the coplanar points recover a unified representation.
[0102] Specifically, the detailed design of the structured plane generation module is as Figure 2 shown. This module includes three sub-modules: plane normal constraint, plane offset constraint, and joint constraint. The plane normal constraint and plane offset constraint respectively exploit the geometric cues in the temporal and spatial dimensions to guide the predicted plane parameters to recover a roughly accurate plane structure representation. Subsequently, the joint constraint constructs a consistency constraint between the two plane parameters, so that the pixels on the same plane have the same plane representation.
[0103] First, the geometric cues in the spatial and temporal dimensions are exploited to constrain the predicted plane normal. Based on the back-projection geometric transformation, the depth map D of the target image is used c to recover the point cloud of the target view and eight adjacent points around each point are selected to generate adjacent initial normals where is the center point, are the adjacent points. To improve the robustness of the finally derived normal, four initial normals are selected at 90-degree intervals and their average value is calculated as the spatial normal cue:
[0104]
[0105] Using the derived normal, based on the similarity between the predicted plane normal of the target image and the spatial normal cue, the normal constraint in the spatial dimension can be defined as:
[0106]
[0107] where, is the downsampling function, which is used to eliminate the noise effect existing in the boundary pixels, and cos represents the cosine similarity.
[0108] In addition, given the plane normal N of the adjacent image n , the rotation matrix R in the relative pose parameters can be used n2c to transform it to the target view to obtain the observed normal N′ n= R n2c N n , where R n2c is the inverse matrix of R c2n . Then, the temporal normal cue under the target view can be obtained through projection transformation:
[0109] N n2c = N' n <proj(D c , T c2n , K)>
[0110] Based on the similarity between the predicted target image plane normal and the temporal normal cue, the normal constraint in the time dimension can be defined as:
[0111]
[0112] Combining the above normal constraints in the spatial and time dimensions, the complete plane normal constraint can be defined as:
[0113]
[0114] Next, geometric cues in the spatial and time dimensions are exploited to constrain the predicted plane offset. Given the derived plane normal and the estimated depth, based on the spatial normal cue, according to the "plane-to-depth" modeling formula, the plane offset cue in the spatial dimension can be defined as:
[0115]
[0116] Therefore, using the single-pixel photometric error to calculate the difference between the plane offset cue in the spatial dimension and the predicted plane offset, the offset constraint in the spatial dimension can be defined as:
[0117]
[0118] Subsequently, given the depth map D c of the target image, its observed depth D ′ n in the adjacent view can be obtained according to the projection transformation, and the observed plane offset in the adjacent view can be inferred based on this observed depth through the depth derivation formula of the adjacent image:
[0119]
[0120] where is the transposed normal of the plane normal N n of the adjacent image.
[0121] In addition, the sampled plane offset in the adjacent view can be calculated through projection transformation and bilinear interpolation:
[0122] O″ n (p)=O n <proj(D c ,T c2n ,K)>
[0123] Based on the similarity of the observation plane offset and the sampling plane offset, the offset constraint in the time dimension can be defined as:
[0124]
[0125] Combining the above offset constraints in the spatial and time dimensions, the complete plane offset constraint can be defined as:
[0126]
[0127] Finally, a joint constraint is constructed between the two plane parameters so that the pixels on the same plane have the same plane representation. By using the plane normal constraint and the plane offset constraint respectively, the approximate plane structure of the two plane properties can be restored. However, these two plane characteristics do not exist in isolation. According to the principle of unique plane representation, all points on the same plane can be represented by sharing a unique set of plane normals and plane offsets. To this end, the joint constraint jointly optimizes them by penalizing the first-order gradients of the plane normals and plane offsets to ensure that coplanar points recover a unified representation:
[0128]
[0129] where represents the operation of calculating the first-order gradient.
[0130] The present invention aims to solve the depth discontinuity problem based on "plane-to-depth" modeling. To this end, the structured generation module proposed by the present invention will be combined with the depth discontinuity perception module to optimize using the following loss function:
[0131]
[0132] where is the final loss function, γ, δ and ω are balance coefficients, which are set to 0.03, 0.01, 0.1 respectively. Using the above constraints, the method first uses the depth discontinuity perception module to identify the regions of depth jumps, and then combines the structured plane generation module to construct accurate parametric geometric representations in these regions, thereby ensuring the accuracy and continuity of the predicted depth.
[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A self-supervised monocular depth estimation method based on parametric geometric representation, characterized in that, Including the following steps: S1: Input the target image with the depth to be predicted and the adjacent images into the depth estimation network to predict the plane normal and plane offset, and derive the depth maps of the target image and the adjacent images according to their geometric mapping relationship with the depth; S2: Input the target image and the adjacent images into the pose network to predict the relative pose parameters between them; S3: Input the depth map of the predicted target image, the relative pose parameters, and the adjacent source image into the depth discontinuity perception module. First, generate a synthetic image according to the projective geometry, and then compare the difference between the synthetic image and the target image. Map the region with a smaller difference as the depth discontinuity mask as the depth discontinuous region; S4: Input the plane normal, plane offset, depth map, and depth discontinuity mask into the structured plane generation module. This module first uses spatio-temporal geometric clues to respectively guide the two plane parameters to recover a relatively accurate plane scene structure, and then jointly optimizes them to ensure that the coplanar points recover a unified representation.
2. The self-supervised monocular depth estimation method based on parametric geometric representation according to claim 1, characterized in that The specific content of S1 is as follows: First, based on the depth estimation network f d predict the planar normals and planar offsets of the target image and adjacent images {I c , I n}: Among them, {N c , O c} is the plane normal and plane offset of the target image, and {N n , O n} is the plane normal and plane offset of the adjacent image; Then a depth map is obtained based on geometric mapping; specifically, define the two-dimensional image point p = [u, v] in the input image T , where u and v are the abscissa and ordinate of the image respectively, and their homogeneous coordinate representations are Based on the plane normal N(p) and plane offset O(p) of each two-dimensional image point p, the corresponding depth map is obtained through the following depth derivation formula: where D(p) is the pixel depth of the two-dimensional image point p, and K is the internal parameter matrix of the camera; the depth maps {D c , D n} of the corresponding target image and adjacent image are obtained based on the pixel depth of the two-dimensional image point p of the input image.
3. The self-supervised monocular depth estimation method based on parametric geometric representation according to claim 2, characterized in that, The specific content of S3 is as follows: The generation process of the synthetic image is represented by the following formula: I ncc = I n <proj(D c , T c2n , K)> Among them, I n2c represents the synthesized image, <·> represents the bilinear interpolation operation, and proj(·, ·, ·) represents the projection operation. T c2n represents the relative pose parameter; then, the single-pixel photometric error is used to calculate the difference between the synthesized image and the real target image, and the pixels smaller than the discontinuity threshold are determined to be located in the depth discontinuity region: where μ is the balance coefficient factor, is the depth discontinuity region mask.
4. The self-supervised monocular depth estimation method based on parametric geometric representation according to claim 3, characterized in that The structured plane generation module specifically includes three sub-modules: plane normal constraint, plane offset constraint, and joint constraint; The specific content of the plane normal constraint is as follows: First, use the depth map of the target image to recover the point cloud of the target view, and select the surrounding adjacent points with each point as the center to generate adjacent normals. Take the average value of the adjacent normals as the spatial normal clue; Obtain the normal constraint in the spatial dimension based on the similarity between the predicted plane normal of the target image and the spatial normal clue; Based on the rotation matrix of the plane normal and relative pose parameters of adjacent images, it is transformed to the target view to obtain the observed normal N′ n ; Then, the temporal normal cue N is obtained through projective transformation n2c : N n2c = N' n <proj(D c , T c2n , K)> Among them, T c2n represents the relative pose parameter; a normal constraint in the time dimension is obtained based on the similarity between the predicted target image plane normal and the temporal normal cue. Finally, combine the normal constraints in the spatial dimension and the temporal dimension to obtain the final plane normal constraint; The specific content of the plane offset constraint is as follows: First, based on the spatial normal clue, the plane offset clue in the spatial dimension is obtained through the depth derivation formula of the target image, and the single-pixel photometric error is used Calculate the difference between the plane offset clue in the spatial dimension and the predicted plane offset to obtain the offset constraint in the spatial dimension; Then, given the depth map of the target image, its observed depth in the adjacent view is obtained according to the projection transformation. Based on this observed depth, the observed plane offset of the adjacent image is obtained through the depth derivation formula of the adjacent image; the sampling plane offset O″ of the adjacent image is obtained through the projection transformation n (p): O″ n (p) = O n <proj(D c , T c2n , K)> Obtain the offset constraint in the temporal dimension based on the similarity between the observed plane offset and the sampled plane offset; Finally, combine the offset constraints in the spatial dimension and the temporal dimension to obtain the final plane offset constraint; The joint constraint is jointly optimized by penalizing the first-order gradient of the plane normal and the plane offset.
5. A self-supervised monocular depth estimation method based on parametric geometric representation according to claim 4, characterized in that, The discontinuous threshold Specifically: Where H and W are the height and width of the image respectively.
6. The self-supervised monocular depth estimation method based on parametric geometric representation according to claim 5, wherein The normal constraint in the spatial dimension is defined as: Among them, is the downsampling function, and cos represents the cosine similarity. is the spatial normal clue; The normal constraint in the temporal dimension is defined as: Combining the normal constraints in the spatial and temporal dimensions, the plane normal constraint is defined as: The offset constraint in the spatial dimension is defined as: wherein is a planar offset clue; The offset constraint in the temporal dimension is defined as: where O n ′ is the observation plane offset; combining the offset constraints in the spatial and temporal dimensions, the plane offset constraint is defined as: The joint constraint is expressed as follows: Among them, represents the operation of finding the first-order gradient.
7. A self-supervised monocular depth estimation method based on parametric geometric representation according to claim 6, characterized in that The plane offset clue is obtained by the following formula: The observed plane offset is obtained by the following formula:
8. A self-supervised monocular depth estimation method based on parametric geometric representation according to claim 7, characterized in that Construct a loss function to jointly optimize the structured generation module and the depth discontinuity perception module: where, is the final loss function, and γ, δ, and ω are balance coefficients.
9. A self-supervised monocular depth estimation method based on parametric geometric representation according to claim 8, characterized in that, The spatial normal clue is obtained in the following way: The point cloud of the target view is P, and eight adjacent points around each point are selected to generate an initial normal vector where P z is the center point, P zx , P zy are the adjacent points; then four initial normal vectors are selected at 90-degree intervals and their average value is calculated as the spatial normal vector clue: