Surround view depth estimation system and method based on geometric consistency and base model
Patent Information
- Application Number
- CN202610548690.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-23
- Publication Date
- 2026-09-01
AI Technical Summary
本发明采用了深度估计模块,引入了几何先验引导和空间几何一致性约束模块,在视角与时序联合重建模块作了相应改进,能够有效解决现有技术中误检、漏检、深度估计不稳定等问题
1)显著提升深度估计精度与几何一致性:
Smart Images

Figure CN122675918A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and autonomous driving perception technology, specifically to a surround-view depth estimation system and method based on geometric consistency and a basic model. Background Technology
[0002] Depth estimation, a core task in computer vision, is a crucial foundation for 3D scene understanding in technologies such as autonomous driving, intelligent robots, and virtual reality (VR). It infers spatial distances between objects and sensors from image or video data, providing precise spatial geometric support for downstream core functions such as obstacle detection, path planning, and scene reconstruction. Surround-view depth estimation, leveraging the advantages of multi-camera collaborative acquisition, can achieve 360° scene coverage without blind spots. Combined with advanced algorithms to generate dense pixel-level depth maps, it achieves high-resolution, blind-spot-free environmental perception at a relatively low hardware cost. Because it provides richer texture details in the near field, it is gradually becoming a replacement for LiDAR in autonomous driving perception systems, showing great promise in the natural fusion of LiDAR point clouds and millimeter-wave radar data. Although challenges remain in extreme lighting and weather conditions, surround-view depth estimation, with its balanced performance and cost capabilities, has become an indispensable perception pillar for mass-produced intelligent vehicles to achieve advanced driver assistance functions.
[0003] Existing self-supervised look-around depth estimation methods mainly rely on photometric reconstruction constraints, using pixel matching between adjacent frames or multi-view images for model training and depth inference. However, they still face many technical bottlenecks in practical engineering applications. Most of these methods focus on cross-view constraints at the photometric level, failing to fully exploit surface normals in 3D space. This results in low depth estimation accuracy in low-texture areas. Furthermore, traditional convolutional neural networks struggle to capture the deep correlation between global semantics and local geometric features, leading to depth breaks at object edges and scene boundaries. Existing depth estimation methods generally suffer from insufficient acquisition of global contextual information and inadequate temporal dependency modeling, resulting in inaccurate depth predictions for small targets and complex structures.
[0004] The Chinese patent "3D Target Detection Method, Apparatus and Device" (application number: 202311395636.6) discloses a depth estimation network training method aimed at improving the accuracy of 3D target detection in complex traffic scenarios and solving the problem of disconnect between 3D target detection and depth estimation in multi-camera surround-view systems. The method establishes a mapping relationship between images by performing distortion correction and synchronization alignment on multi-camera surround-view images and constructing camera calibration parameters. After multi-scale feature extraction into the network, the disparity map obtained is converted into a depth map through a cost matrix. Furthermore, the model cleverly combines target detection and depth, specifically by fusing the depth map with image features and inputting them into the 3D target detection network to obtain the target's bounding box, category, and 3D coordinate information. The introduction of non-maximum suppression achieves more accurate 3D target detection. However, its limitation lies in its over-reliance on precise camera calibration parameters. When the calibration error is large or the camera installation position is slightly offset, the detection accuracy will be significantly affected.
[0005] Chinese patent "Lightweight Monocular Depth Estimation Method and Device Based on Self-Supervised Deep Learning" (application number: 202411439547.1) discloses a lightweight depth estimation method and system, aiming to solve the problems of large parameter count, high computational cost, and difficulty in deployment on mobile devices in traditional monocular depth estimation models. This method extracts the network through a lightweight architecture of depth-separable convolutions, introducing an attention mechanism to retain the ability to extract key features while maintaining a lightweight design. In the network design, a semantic information extraction module is added to force consistent depth values for regions with the same semantic category, effectively avoiding mismatches. In terms of social benefits, the temporal information of monocular videos can be used to reconstruct the loss for self-supervised training, saving a significant amount of cost associated with manually labeled depth data. In terms of economic benefits, model quantization and pruning can be tailored to the hardware characteristics of mobile devices, ensuring real-time inference on low-computing-power devices. The drawback is that while the lightweight design reduces computational cost, it sacrifices some high-order feature extraction capabilities, especially in depth estimation of distant targets and slender structures, resulting in lower accuracy compared to heavyweight models.
[0006] Existing deep learning-based depth estimation methods have achieved success in some applications, but due to the limitations of neural networks, these methods often struggle to capture global temporal information and perform poorly in small target detection. For example, they cannot effectively extract the spatial relationships between distant or elongated targets, leading to significant errors in accurately measuring target distances. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention provides a surround-view depth estimation system and method based on geometric consistency and a fundamental model. The aim is to impose geometric consistency constraints on the depth estimation results in three-dimensional space by introducing geometric prior information provided by a fundamental visual model. Furthermore, by combining a multi-view and cross-temporal joint reconstruction mechanism, it achieves high-precision and stable surround-view depth estimation under conditions where real depth annotations are unavailable. This invention employs a depth estimation module, introduces a geometric prior guidance and spatial geometric consistency constraint module, and makes corresponding improvements to the view-time and temporal joint reconstruction module. This effectively solves problems such as false detections, missed detections, and unstable depth estimation in existing technologies.
[0008] The technical solution adopted in this invention is as follows:
[0009] A look-around depth estimation system based on geometric consistency and a fundamental model, comprising: Semantic-guided multi-view depth estimation module, geometric prior guidance module, spatial geometric consistency constraint module, adaptive joint motion learning module, and view and temporal joint reconstruction module; The semantically guided multi-view depth estimation module is used to extract geometric features and high-level semantic features from the input multi-view panoramic images, obtain semantically enhanced geometric representations through cross-modal attention fusion, and output initial depth maps for each view. The geometric prior guidance module is used to generate pseudo-depth geometric priors based on the frozen depth basic model, and to combine the surface normal vector and inverse depth gradient of the initial depth map to construct scale-invariant geometric consistency constraints in order to regularize the local geometric structure of the initial depth map. The spatial geometric consistency constraint module is used to back-project the source view depth to three-dimensional space and transform it to the target view coordinate system based on the camera intrinsic and extrinsic parameters between multiple viewpoints, generate a spatially dense depth map, and construct depth consistency constraints in the overlapping areas of multiple viewpoints. The adaptive joint motion learning module is used to extract motion features from multi-view image pairs and perform adaptive weight fusion to predict the relative pose of adjacent time frames. The viewpoint and temporal joint reconstruction module is used to perform geometric projection and image reconstruction in spatial, temporal and spatiotemporal joint contexts based on the initial depth map, spatial dense depth and relative pose, to construct multi-context photometric consistency constraints, so as to jointly optimize depth and pose and output the surrounding dense depth estimation results.
[0010] The semantically guided multi-view depth estimation module includes: The frozen CLIP image encoder is used to extract high-level semantic feature tokens from input image I: ,in: This represents the input image; A depth encoder is used to extract geometric features. ,in: N represents the number of viewpoints, C is the number of feature channels; h and w represent the spatial resolution of the feature map; Cross-modal attention unit, used to perform the following operations: Convolution, projection, and scaling: ,in: To adapt to the subsequent attention mechanisms, , , Channel number and spatial resolution adapted for the attention mechanism; ,in: For querying the matrix, The key matrix, For value matrices, Let be the dimension of the key matrix; The obtained attention modules are effectively fused and then projected back to the original resolution and channels: ,in: For each channel, the modulation factor is... The Softmax function, representing the Hadamard product, outputs the probability distribution of the attention weight matrix; the Softmax function reflects the probability distribution of the attention weight matrix, φ⊙( The result of the channel-by-channel Hadamard product is input into the residual connection.
[0011] ,in: For channel recovery and projection functions; The depth encoder outputs an initial depth map. ,in: For perspective indexing, For time indexing.
[0012] The geometric prior guidance module includes: Frozen DepthAnything model used to generate pseudo-depth priors ; Depth predicted by the frozen DepthAnything V2 network; target frame image Depth values calculated by depth estimation The input is fed into the DA network, which is a normalized network that is then converted into a pseudo-depth network with an indeterminate scale. Normal calculation unit, used to calculate the predicted normal map based on the initial depth map. Calculating pseudo-normal graphs based on pseudo-depth priors And construct a 3D surface normal uniformity loss: ; in: L1 norm, superscript Indicates matrix transpose; Gradient calculation unit, used to obtain inverse depth based on initial depth map. Inverse depth is obtained based on pseudo-depth prior. And construct the inverse depth gradient consistency loss: ; in: This is the spatial gradient operator.
[0013] The spatial geometric consistency constraint module includes: Back projection unit, used to utilize the source viewpoint j Predicted depth and camera internal parameters External reference homogeneous pixel coordinates = ( , ,1) Backprojection to 3D points: ; Coordinate transformation unit, used to utilize the target viewpoint i Camera external parameters Transform 3D points to the target viewpoint coordinate system: ; The reprojection unit is used to project the transformed 3D points onto the target viewpoint image plane to obtain the reconstructed depth. ; in: This indicates taking the z-axis coordinate value of a three-dimensional point; Loss building blocks, used to construct spatial geometric consistency loss: ; in: From the perspective of the target i The predicted depth map.
[0014] The combined perspective and temporal reconstruction module includes: Geometric projection units, used for predicting depth Camera internal parameters , Using the relative pose output by the adaptive joint motion learning module, construct a geometric projection mapping: ; in: From the perspective of the target i At any moment t homogeneous pixel coordinates; To map to the source view j The corresponding pixel position at time t'; ; This means mapping a pixel in the target viewpoint i at time t to the corresponding pixel position in the source viewpoint j at time t′. This describes sampling the mapped pixel positions in the source viewpoint j and time t′ image to generate a reconstructed image. The most crucial step is mapping the pixels of the target viewpoint onto the source viewpoint image plane. The specific process is as follows: .
[0015] in: The coordinate transformation matrix takes values based on the reconstruction context as follows: During time-series reconstruction, ; During spatial reconstruction, ; hour, ; in: , The network predicts the relative motion between adjacent time frames, and calculates... This represents the result of mapping the target viewpoint pixels to the source viewpoint image plane; , This refers to the camera's external parameters.
[0016] Image reconstruction unit, used to generate reconstructed images: ; in: From the source perspective j The image at time t'; The photometric error calculation unit is used to construct the photometric consistency error function. ; in: These are the weighting coefficients. and These are the values at the same pixel position in the target image and the reconstructed image, respectively; Compare target images at the same pixel location Reconstructed image obtained by geometric projection sampling ; It measures the degree of similarity between two images in their local structure; By converting similarity into error and performing range normalization, a hyperparameter α is defined to control the proportion of SSIM, then (1 α) Control the proportion of L1.
[0017] Joint loss unit, used for separate calculations: Temporal photometric loss: ; Spatial photometric loss: ; Spacetime combined photometric loss: ; Consistency loss in multi-view reconstruction: ; And output the weighted sum of the joint photometric loss: ; in: . Represents images from the source viewpoint at the same time. This represents the source viewpoint images at adjacent times, ultimately yielding the multi-context photometric error. , Take the minimum value for each pixel.
[0018] The adaptive joint motion learning module includes: Pose encoder for extracting multi-view image pairs The motion characteristics, among which, t and t' These are adjacent time frames; An adaptive weight learning unit is used to allocate the contribution of different perspectives to motion estimation through learnable parameters and to perform weighted fusion of motion features from each perspective. The pose prediction unit is used to output the relative pose between adjacent time frames. .
[0019] The look-around depth estimation method based on geometric consistency and fundamental model includes the following steps: S1: Input the multi-view panoramic image into the system, extract the geometric features and high-level semantic features of each view, obtain the semantically enhanced geometric representation through cross-modal attention fusion, and predict the initial depth map of each view based on the geometric representation. S2: Use the frozen depth foundation model to generate pseudo-depth geometric priors for the input image, combine the initial depth map to calculate the surface normal vector and inverse depth gradient, and construct scale-invariant geometric consistency constraints to regularize the local 3D structure of the initial depth map. S3: Based on the camera's intrinsic and extrinsic parameters, backproject the predicted depth from the source viewpoint to the 3D space and transform it to the target viewpoint coordinate system to generate a spatially dense depth and construct depth consistency constraints in the multi-viewpoint overlapping region. S4: Extract motion features from multi-view image pairs and perform adaptive weight fusion to predict the relative pose between adjacent time frames; S5: Based on the initial depth map, spatially dense depth and relative pose, perform geometric projection and image reconstruction in spatial, temporal and spatiotemporal joint contexts, and construct multi-context photometric consistency constraints. S6: Jointly optimize depth and pose to output geometrically consistent and temporally stable look-around dense depth estimation results.
[0020] S1 includes: S1.1: Extract high-level semantic feature tokens from input image I using the frozen CLIP image encoder: ; S1.2: Extracting geometric features using a depth encoder ,in: N represents the number of viewpoints, C is the number of feature channels; h and w represent the spatial resolution of the feature map; S1.3: Perform cross-modal attention operation: ,in: To adapt to the subsequent attention mechanisms, , , Dimensions adapted for the attention mechanism; ,in: For querying the matrix, The key matrix, For value matrices, Let be the dimension of the key matrix; ,in: For each channel, the modulation factor is... The Softmax function, representing the Hadamard product, outputs the probability distribution of the attention weight matrix. ,in: For channel recovery and projection functions; S1.4: Output the initial depth map through the depth decoder. ,in: For perspective indexing, For time indexing.
[0021] In the cross-modal attention operation, the attention weight matrix output by the Softmax function reflects the correlation distribution between semantic features and deep features, and is used to guide the injection of semantic information into deep features.
[0022] S2 includes: S2.1: Generate pseudo-depth priors using the frozen DepthAnything model ; S2.2: Calculate the predicted normal map based on the initial depth map Calculating pseudo-normal graphs based on pseudo-depth priors Constructing 3D surface normal uniformity loss: ; in: L1 norm, superscript Indicates matrix transpose; S2.3: Used to obtain inverse depth based on the initial depth map Inverse depth is obtained based on pseudo-depth prior. And construct the inverse depth gradient consistency loss: ; in: This is the spatial gradient operator.
[0023] S2.4: Regularize the local 3D structure of the initial depth map using 3D surface normal consistency loss and inverse depth gradient consistency loss.
[0024] Step S3 specifically includes: S3.1: Utilizing the source perspective j Predicted depth and camera internal parameters External reference homogeneous pixel coordinates = ( , ,1) Backprojection to 3D points: ; S3.2: Utilizing the target perspective i Camera external parameters Transform 3D points to the target viewpoint coordinate system: ; S3.3: Project the transformed 3D points onto the target viewpoint image plane to obtain the reconstructed depth: ; in: This indicates taking the z-axis coordinate value of a three-dimensional point; S3.4: Constructing Spatial Geometric Consistency Loss: ; in: From the perspective of the target i The predicted depth map.
[0025] Step S4 specifically includes: S4.1: Pair multi-view images The input pose encoder extracts motion features, where: t and t' For adjacent time frames; S4.2: Through the adaptive weight learning unit, the contribution of different viewpoints to motion estimation is allocated using learnable parameters, and the motion features of each viewpoint are weighted and fused.
[0026] S4.3: Output the relative pose between adjacent time frames .
[0027] Step S5 specifically includes: S5.1: Based on prediction depth Camera internal parameters , Given the relative pose, construct the geometric projection mapping: pijt←t′=Πijt←t′pit pijt ← t ′=Π ijt ← t ′ pit , ; in: From the perspective of the target i At any moment t homogeneous pixel coordinates; To map to the source view j The corresponding pixel position at time t'; ; in: The coordinate transformation matrix takes values based on the reconstruction context as follows: During time-series reconstruction, ; During spatial reconstruction, ; hour, ; in: , The relative motion between adjacent time frames predicted by the network. , This refers to the camera's external parameters.
[0028] S5.2: Generate reconstructed image: ; in: From the source perspective j The image at time t'; S5.3: Constructing the photometric consistency error function: ; in: These are the weighting coefficients. and These are the values at the same pixel position in the target image and the reconstructed image, respectively; S5.4: Calculate separately: Temporal photometric loss: ; Spatial photometric loss: ; Spacetime combined photometric loss: ; Consistency loss in multi-view reconstruction: ; S5.5: Calculate the weighted sum of joint photometric losses: ; in: .
[0029] In step S6, when jointly optimizing depth and pose, a total loss function is constructed using geometric consistency constraints, spatial geometric consistency loss, and joint photometric loss. The parameters of the depth estimation network and pose network are then updated through backpropagation.
[0030] This invention provides a look-around depth estimation system and method based on geometric consistency and a fundamental model, with the following technical advantages: 1) Significantly improves depth estimation accuracy and geometric consistency: This invention introduces a semantically guided multi-view depth estimation module, which enhances geometric feature representation using high-level semantic features extracted by the CLIP model, effectively improving depth prediction quality for low-texture regions, distant targets, and object edges. Simultaneously, it incorporates the 3D surface normal consistency loss from the geometric prior guidance module (…). Consistency loss with inverse depth gradient This invention constrains the geometric rationality of local 3D structures without relying on true depth annotations. Validated on the KITTI dataset, compared to existing methods such as MonoViT and EDS-Depth, this invention reduces the AbsRel metric to 0.094 and improves the δ3 accuracy to 98.6%, demonstrating significantly better overall performance than the comparative models.
[0031] 2) Eliminate the inconsistency in depth prediction across multiple views: This invention establishes a spatial geometric consistency constraint module, which constructs a spatial geometric consistency loss by back-projecting the source view depth to three-dimensional space and then re-projecting it to the target view coordinate system. This mechanism forces depth predictions from different viewpoints to remain consistent within overlapping regions. It effectively eliminates multi-view geometric drift, improving the spatial continuity of overall 3D reconstruction and the fusion quality of cross-view depth maps.
[0032] 3) Enhance temporal stability and robustness in dynamic scenarios: This invention employs an adaptive joint motion learning module and a viewpoint-temporal joint reconstruction module to predict relative pose through adaptive weighted fusion of multi-view motion features, and constructs a multi-context photometric consistency loss under temporal, spatial, and spatiotemporal joint contexts. This scheme makes full use of multi-view and temporal information, significantly suppressing the interference caused by dynamic objects, occlusion and changes in lighting, so that the depth estimation results remain temporally stable under complex motion conditions.
[0033] 4) Reduce reliance on actual depth annotations and lower data costs: This invention employs a self-supervised training framework, requiring only a panoramic image sequence and camera calibration parameters, eliminating the need for expensive LiDAR depth ground truth or manual annotation. By leveraging the pseudo-depth prior provided by the DepthAnything model frozen in the geometric prior guidance module, combined with scale-invariant geometric consistency constraints, high-quality self-supervised deep learning is achieved, significantly reducing data acquisition and annotation costs and facilitating deployment in practical engineering scenarios such as mass-produced intelligent vehicles.
[0034] 5) Enhance the ability to perceive the depth of small targets and slender structures: The semantically guided cross-modal attention mechanism in this invention injects high-level semantic features into geometric features, enabling the network to better distinguish objects of different semantic categories and avoid incorrect matching of depth values within the same semantic region. Combined with spatial geometric consistency constraints, it effectively improves the completeness of depth estimation for small targets and slender structures (such as traffic signs and lampposts), reducing false positive and false negative rates. Attached Figure Description
[0035] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments: figure 1 This is a diagram of the overall structure of the invention.
[0036] figure 2 A flowchart for semantically guided multi-view depth estimation.
[0037] figure 3 This is a flowchart for geometric prior guidance.
[0038] figure 4 This is a flowchart of spatial geometric consistency constraints.
[0039] figure 5 This is a flowchart for adaptive joint motion learning.
[0040] figure 6 A flowchart for the joint reconstruction of perspective and time sequence.
[0041] figure 7 This is a network architecture diagram. Detailed Implementation
[0042] This embodiment provides a surround-view depth estimation system and method based on geometric consistency and a fundamental model, applicable to scenarios requiring 360° dense depth perception, such as autonomous vehicles and intelligent robots. The overall system architecture is as follows: figure 1 As shown, the system mainly includes a semantically guided multi-view depth estimation module, a geometric prior guidance module, a spatial geometric consistency constraint module, an adaptive joint motion learning module, and a viewpoint and temporal joint reconstruction module. The following sections, in conjunction with the accompanying figures, provide a detailed explanation of each module and its corresponding methodological steps.
[0043] (I): Semantic-guided multi-view depth estimation module: like figure 2 As shown, figure 2 Specifically, the process involves inputting multi-view panoramic images, which are then fed into a depth encoder and a frozen CLIP image encoder to extract geometric features and high-level semantic features. Semantic features are injected into geometric features through a cross-modal attention mechanism to obtain semantically enhanced geometric representations. Finally, the initial depth maps for each viewpoint are output by the depth decoder.
[0044] The semantically guided multi-view depth estimation module first obtains the target time. t N perspective panoramic images ,in, Indicates different camera perspectives.
[0045] 1. Feature Extraction: Images from each viewpoint are input into the depth encoder and the frozen CLIP image encoder, respectively. The CLIP image encoder extracts high-level semantic feature tokens. Provides semantic awareness to deep features. The deep encoder extracts geometric features. N represents the number of viewpoints, C represents the number of feature channels, and h and w represent the feature map space size at a resolution lower than that of the original image.
[0046] 2: Cross-modal attention fusion: To inject semantic features into geometric features, the deep features are first subjected to convolutional projection and scaling. ; in: This is a channel compression function used to adjust the number of channels to meet the dimensionality requirements of subsequent attention mechanisms. The Resize operation is used to match the spatial resolution.
[0047] Then construct cross-modal attention: ; Here, the CLIP semantic token is reshaped into a query matrix. Its transpose serves as the key matrix. Deep features as value matrix . denoted as the dimension of the key matrix.
[0048] Calculate and fuse attention weights: ; Wherein: the Softmax function outputs the probability distribution of the attention weight matrix, For each channel, the modulation factor is... This represents the Hadamard product, which is multiplied element by element; finally, the attention-enhanced features are added back to the original features through residual connections.
[0049] Finally, restore the original resolution and number of channels: ; in: This is a function for channel recovery and projection.
[0050] 3: Depth Decoding: Input the fused features into the depth decoder to obtain the initial depth map for each viewpoint: ; Through the above operations, high-level semantic information is effectively guided into geometric features, significantly improving the stability of depth prediction in weakly textured regions, distant targets, and complex scenes.
[0051] (II): Geometric Prior Guidance Module like figure 3 As shown, figure 3 Demonstrates: generating pseudo-depth geometric priors using a frozen DepthAnything model; calculating surface normals and inverse depth gradients based on the initial depth map and pseudo-depth priors respectively; and constructing a 3D surface normal consistency loss. Consistency loss with inverse depth gradient This achieves scale-invariant regularization of local three-dimensional structures.
[0052] The geometry prior guidance module uses the frozen DepthAnything V2 model to generate pseudo-depth priors and performs scale-invariant regularization on the local geometry of the initial depth map.
[0053] 1: Pseudo-depth generation: This involves generating a pseudo-depth value from the input image. Feed the frozen DepthAnything model to obtain pseudo-depth prior. The true depth and the pseudo depth satisfy an approximate relationship. ,in, This module avoids dependence on absolute scale by using a scale-invariant loss function.
[0054] 2: Surface normal vector calculation: based on predicted depth and pseudo-depth Calculate the surface normal vectors separately. For each pixel p, construct a normal map based on four pairs of orthogonal neighbors. For the j-th pair of neighboring pixels, the normal map is constructed by pixel p and its two neighboring pixels. , The cross product of the vectors formed is used to calculate the unit normal: ; Then, directional consistency fusion is performed in the four directions: ; in: The sign function ensures that the directions of the normal vectors are consistent, ultimately yielding the surface normal corresponding to pixel p. .
[0055] Using predicted depth respectively and pseudo-depth The normal diagram was calculated. and .
[0056] 3: Constructing 3D surface normal consistency loss: ; This loss constrains the difference between the two normal maps by using the L1 norm, which helps the predicted depth to align with the pseudo-depth prior on the local surface orientation.
[0057] 4: Construct inverse depth gradient consistency loss: transform depth into inverse depth and Calculate the spatial gradient: ; in: This is a spatial gradient operator (horizontal and vertical directions). This loss constraint ensures that the local rate of change of the predicted depth is consistent with the pseudo-depth prior, further regularizing the smoothness and edge sharpness of the depth map.
[0058] Through the two types of loss mentioned above, this module effectively suppresses depth noise and geometric distortion without relying on real depth annotations.
[0059] (III): Spatial Geometric Consistency Constraint Module: like figure 4 As shown, figure 4 Specifically, the process involves: backprojecting the source viewpoint pixels into 3D space using the predicted depth from the source viewpoint and camera intrinsic and extrinsic parameters; transforming the source viewpoint coordinates using extrinsic parameter transformation; projecting the reconstructed depth onto the target viewpoint image plane; and constructing a spatial geometric consistency loss. This constrains the consistency of depth predictions within overlapping regions of multiple views. The spatial geometric consistency constraint module utilizes camera extrinsics across multiple views to align depth predictions from different perspectives in 3D space.
[0060] 1: Back projection: For a certain pixel in the source viewpoint j = ( , ,1) Using its predicted depth and camera internal parameters Back projection onto three-dimensional space: ; 2: Coordinate Transformation: Utilizing the Target's Viewpoint i Camera external parameters Heyuan Perspective j Camera external parameters Transform 3D points to the target viewpoint i Camera coordinate system: ; 3: Reprojection: Project the transformed 3D points onto the image plane of the target viewpoint i to obtain the reconstructed depth. ; in: This indicates that the z-axis coordinate value of a 3D point is taken, which is the depth value.
[0061] 4: Construct spatial geometric consistency loss: ; here From the perspective of the target i The predicted depth map is obtained. This loss forces the depth prediction of the source viewpoint to be consistent with the predicted depth of the target viewpoint after reprojection, thereby eliminating geometric deviations between multiple views and improving the spatial continuity of the overall 3D reconstruction.
[0062] (iv): Adaptive Joint Motion Learning Module likefigure 5 As shown, figure 5 This demonstrates: inputting multi-view images into a pose encoder to extract motion features; weighting and fusing the motion features from each viewpoint using an adaptive weight learning mechanism; and finally predicting the relative pose between adjacent time frames. .
[0063] An adaptive joint motion learning module is used to predict the relative pose between adjacent time frames and improves the robustness of pose estimation through adaptive weight fusion.
[0064] 1) Motion feature extraction: Extracting motion features from multi-view images (in t and t' The motion features of each viewpoint are extracted by inputting adjacent time frames into the pose encoder. 2) Adaptive Weight Fusion: Using learnable attention weight parameters, the contributions of different perspectives to motion estimation are allocated. The weights are calculated as follows: ; in, For the first i Motion characteristics from a single perspective A multilayer perceptron is used, and then the features from each viewpoint are weighted and summed to obtain the fused global motion features.
[0065] 3) Pose prediction: The fused features are input into the pose prediction head, and the relative pose between adjacent time frames is output. .
[0066] (V): Joint Reconstruction Module of Perspective and Temporal Sequence: like figure 6 As shown, figure 6 Specifically, it demonstrates the following: based on predicted depth, spatially dense depth, and relative pose, geometric projection mapping is constructed in three contexts: temporal reconstruction, spatial reconstruction, and spatiotemporal joint reconstruction; reconstructed images are generated and multi-context photometric consistency loss is calculated.
[0067] The viewpoint and temporal joint reconstruction module performs geometric projection and image reconstruction in three contexts: spatial, temporal, and spatiotemporal joint, based on predicted depth, spatial dense depth, and relative pose, and constructs a multi-context photometric consistency loss.
[0068] 1) Geometric projection mapping: Mapping the pixel pitpit of the target viewpoint i at time t to the source viewpoint j at time t using depth, camera pose, and motion pose. t Image plane Mapping depth, camera pose, and motion pose to the source viewpoint j At any moment t' Image plane: ; The projection matrix is defined as: ; coordinate transformation matrix Values are determined based on the type of the reconstruction context: When reconstructing time series (same perspective, adjacent times), ; When reconstructing space (at the same time, from different perspectives), ; hour, ; in, , The relative motion between adjacent time frames predicted by the network. , This refers to the camera's external parameters.
[0069] 2) Image Reconstruction: By performing differentiable sampling on the source image, a reconstructed image is generated. ; 3) Photometric error function: An error function combining the L1 norm and SSIM (structural similarity index) is adopted. ; in, =0.85, used to balance the contributions of the two errors. These are the values at the same pixel position in the target image and the reconstructed image, respectively.
[0070] 4) Multi-context photometric loss: Calculate the photometric loss under four different contexts: Temporal photometric loss: ; Spatial photometric loss: ; Spacetime combined photometric loss: ; in: This indicates that the minimum error is selected from multiple source time frames to address occlusion and moving objects.
[0071] 5) Combined photometric loss: The weighted sum of the above four losses: ; In this embodiment, the weighting coefficients can be set as follows: , , , .
[0072] (vi): Joint optimization and deep output: During training, all the above loss functions are combined into a total loss: ; in: , , , These are the balance coefficients for each loss term. The parameters of the depth estimation network and pose network are updated end-to-end using the backpropagation algorithm.
[0073] During the inference phase, only a single-time-time multi-view panoramic image needs to be input. The semantically guided multi-view depth estimation module can directly output a high-precision, geometrically consistent, and temporally stable dense depth map without calculating any loss function.
[0074] (VII): Experimental verification: To verify the effectiveness of this invention, a quantitative evaluation was performed on the KITTI dataset. The method of this invention (GeoSurDepth) was compared with existing mainstream methods MonoViT and EDS-Depth, and the results are shown in Table 1:
[0075] Table 1 compares the quantitative performance of the method of this invention with existing mainstream depth estimation models (MonoViT, EDS-Depth), with evaluation metrics covering depth error (AbsRel, Sq Rel↓, RMSE, RMSELogn) and depth accuracy (δ1, δ2, δ3). The comparison results show that, except for the Sq Rel parameter, the method of this invention achieves the best performance in all key metrics. Specifically, AbsRel decreases from 0.099 / 0.095 to 0.094, and the δ3 accuracy improves to 98.6%, demonstrating significantly better overall performance than the comparative models. This indicates that the advantage of this invention can be explained by the introduction of a semantic guidance mechanism that injects high-level semantics into geometric features, making the depth of weak textures, reflective areas, or distant regions more stable, thereby reducing Abs Rel and RMSE. Geometric prior guidance, through normal consistency and inverse depth gradient consistency, constrains the local 3D shape without relying on absolute scale, significantly reducing RMSELogn and improving the "strict accuracy" of δ1. Spatial geometric consistency constraints force depth prediction alignment in multi-view overlapping regions, reducing cross-view geometric drift and enabling more pixels to meet the 1.25x threshold range, thereby improving δ1, δ2, and δ3. The core mechanism of adaptive joint motion learning and view-temporal joint reconstruction further enhances overall consistency by suppressing dynamic interference and occlusion errors through more robust pose estimation and multi-context photometric constraints. In summary, GeoSurDepth exhibits comprehensive advantages over comparative methods in three dimensions: lower error, more stable scale, and higher effective pixel count, making it more suitable for obtaining reliable depth results in real-world complex scenes.
[0076] The results show that the present invention achieves the best or second-best performance in absolute relative error (AbsRel), root mean square error (RMSE), logarithmic root mean square error (RMSELog), and accuracy indicators (δ1, δ2, δ3), verifying the effectiveness of the proposed module.
Claims
1. A surround depth estimation system based on geometric consistency and a fundamental model, characterized in that... The system includes: Semantic-guided multi-view depth estimation module, geometric prior guidance module, spatial geometric consistency constraint module, adaptive joint motion learning module, and view and temporal joint reconstruction module; The semantically guided multi-view depth estimation module is used to extract geometric features and high-level semantic features from the input multi-view panoramic images, obtain semantically enhanced geometric representations through cross-modal attention fusion, and output initial depth maps for each view. The geometric prior guidance module is used to generate pseudo-depth geometric priors based on the frozen depth basic model, and to combine the surface normal vector and inverse depth gradient of the initial depth map to construct scale-invariant geometric consistency constraints in order to regularize the local geometric structure of the initial depth map. The spatial geometric consistency constraint module is used to back-project the source view depth to three-dimensional space and transform it to the target view coordinate system based on the camera intrinsic and extrinsic parameters between multiple viewpoints, generate a spatially dense depth map, and construct depth consistency constraints in the overlapping areas of multiple viewpoints. The adaptive joint motion learning module is used to extract motion features from multi-view image pairs and perform adaptive weight fusion to predict the relative pose of adjacent time frames. The viewpoint and temporal joint reconstruction module is used to perform geometric projection and image reconstruction in spatial, temporal and spatiotemporal joint contexts based on the initial depth map, spatial dense depth and relative pose, to construct multi-context photometric consistency constraints, so as to jointly optimize depth and pose and output the surrounding dense depth estimation result.
2. The surround depth estimation system based on geometric consistency and fundamental model as described in claim 1, characterized in that: The semantically guided multi-view depth estimation module includes: The frozen CLIP image encoder is used to extract high-level semantic feature tokens from input image I: ,in: This represents the input image; A depth encoder is used to extract geometric features. ,in: N represents the number of viewpoints, C is the number of feature channels; h and w represent the spatial resolution of the feature map; Cross-modal attention unit, used to perform the following operations: ,in: To adapt to the subsequent attention mechanisms, , , Channel number and spatial resolution adapted for the attention mechanism; ,in: For querying the matrix, The key matrix, For value matrices, Let be the dimension of the key matrix; ,in: For each channel, the modulation factor is... The Softmax function, representing the Hadamard product, outputs the probability distribution of the attention weight matrix. ,in: For channel recovery and projection functions; The depth encoder outputs an initial depth map. ,in: For perspective indexing, For time indexing.
3. The surround depth estimation system based on geometric consistency and fundamental model as described in claim 1, characterized in that: The geometric prior guidance module includes: Frozen DepthAnything model used to generate pseudo-depth priors ; Normal calculation unit, used to calculate the predicted normal map based on the initial depth map. Calculating pseudo-normal graphs based on pseudo-depth priors And construct a 3D surface normal uniformity loss: ; in: L1 norm, superscript Indicates matrix transpose; Gradient calculation unit, used to obtain inverse depth based on initial depth map. Inverse depth is obtained based on pseudo-depth prior. And construct the inverse depth gradient consistency loss: ; in: This is the spatial gradient operator.
4. The surround depth estimation system based on geometric consistency and fundamental model as described in claim 1, characterized in that: The spatial geometric consistency constraint module includes: Back projection unit, used to utilize the source viewpoint j Predicted depth and camera internal parameters External reference homogeneous pixel coordinates = ( , ,1) Backprojection to 3D points: ; Coordinate transformation unit, used to utilize the target viewpoint i Camera external parameters Transform 3D points to the target viewpoint coordinate system: ; The reprojection unit is used to project the transformed 3D points onto the target viewpoint image plane to obtain the reconstructed depth. ; in: This indicates taking the z-axis coordinate value of a three-dimensional point; Loss building blocks, used to construct spatial geometric consistency loss: ; in: From the perspective of the target i The predicted depth map.
5. The surround depth estimation system based on geometric consistency and fundamental model as described in claim 1, characterized in that: The combined perspective and temporal reconstruction module includes: Geometric projection units, used for predicting depth Camera internal parameters , Using the relative pose output by the adaptive joint motion learning module, construct a geometric projection mapping: ; in: From the perspective of the target i At any moment t homogeneous pixel coordinates; To map to the source view j The corresponding pixel position at time t'; ; in: The coordinate transformation matrix takes values based on the reconstruction context as follows: During time-series reconstruction, ; During spatial reconstruction, ; hour, ; in: , The relative motion between adjacent time frames predicted by the network. , For camera external parameters; Image reconstruction unit, used to generate reconstructed images: ; in: From the source perspective j The image at time t'; The photometric error calculation unit is used to construct the photometric consistency error function. ; in: These are the weighting coefficients. and These are the values at the same pixel position in the target image and the reconstructed image, respectively; Joint loss unit, used for separate calculations: Temporal photometric loss: ; Spatial photometric loss: ; Spacetime combined photometric loss: ; Consistency loss in multi-view reconstruction: ; And output the weighted sum of the joint photometric loss: ; in: ; The adaptive joint motion learning module includes: Pose encoder for extracting multi-view image pairs The motion characteristics, among which, t and t' These are adjacent time frames; An adaptive weight learning unit is used to allocate the contribution of different perspectives to motion estimation through learnable parameters and to perform weighted fusion of motion features from each perspective. The pose prediction unit is used to output the relative pose between adjacent time frames. .
6. A surround depth estimation method based on geometric consistency and a fundamental model, characterized in that... Includes the following steps: S1: Input the multi-view panoramic image into the system, extract the geometric features and high-level semantic features of each view, obtain the semantically enhanced geometric representation through cross-modal attention fusion, and predict the initial depth map of each view based on the geometric representation. S2: Use the frozen depth foundation model to generate pseudo-depth geometric priors for the input image, combine the initial depth map to calculate the surface normal vector and inverse depth gradient, and construct scale-invariant geometric consistency constraints to regularize the local 3D structure of the initial depth map. S3: Based on the camera's intrinsic and extrinsic parameters, backproject the predicted depth from the source viewpoint to the 3D space and transform it to the target viewpoint coordinate system to generate a spatially dense depth and construct depth consistency constraints in the multi-viewpoint overlapping region. S4: Extract motion features from multi-view image pairs and perform adaptive weight fusion to predict the relative pose between adjacent time frames; S5: Based on the initial depth map, spatially dense depth and relative pose, perform geometric projection and image reconstruction in spatial, temporal and spatiotemporal joint contexts, and construct multi-context photometric consistency constraints. S6: Jointly optimize depth and pose to output geometrically consistent and temporally stable look-around dense depth estimation results.
7. The surround depth estimation method based on geometric consistency and fundamental model as described in claim 6, characterized in that: S1 includes: S1.1: Extract high-level semantic feature tokens from input image I using the frozen CLIP image encoder: ; S1.2: Extracting geometric features using a depth encoder ,in: N represents the number of viewpoints, C is the number of feature channels; h and w represent the spatial resolution of the feature map; S1.3: Perform cross-modal attention operation: ,in: To adapt to the subsequent attention mechanisms, , , Dimensions adapted for the attention mechanism; ,in: For querying the matrix, The key matrix, For value matrices, Let be the dimension of the key matrix; ,in: For each channel, the modulation factor is... The Softmax function, representing the Hadamard product, outputs the probability distribution of the attention weight matrix. ,in: For channel recovery and projection functions; S1.4: Output the initial depth map through the depth decoder. ,in: For perspective indexing, For time indexing.
8. The surround depth estimation method based on geometric consistency and fundamental model as described in claim 6, characterized in that: S2 includes: S2.1: Generate pseudo-depth priors using the frozen DepthAnything model ; S2.2: Calculate the predicted normal map based on the initial depth map Calculating pseudo-normal graphs based on pseudo-depth priors Constructing 3D surface normal uniformity loss: ; in: L1 norm, superscript Indicates matrix transpose; S2.3: Used to obtain inverse depth based on the initial depth map Inverse depth is obtained based on pseudo-depth prior. And construct the inverse depth gradient consistency loss: ; in: For spatial gradient operators; S2.4: Regularize the local 3D structure of the initial depth map using 3D surface normal consistency loss and inverse depth gradient consistency loss.
9. The surround depth estimation method based on geometric consistency and fundamental model as described in claim 8, characterized in that: Step S3 specifically includes: S3.1: Utilizing the source perspective j Predicted depth and camera internal parameters External reference homogeneous pixel coordinates = ( , ,1) Backprojection to 3D points: ; S3.2: Utilizing the target perspective i Camera external parameters Transform 3D points to the target viewpoint coordinate system: ; S3.3: Project the transformed 3D points onto the target viewpoint image plane to obtain the reconstructed depth: ; in: This indicates taking the z-axis coordinate value of a three-dimensional point; S3.4: Constructing Spatial Geometric Consistency Loss: ; in: From the perspective of the target i The predicted depth map.
10. The surround depth estimation method based on geometric consistency and fundamental model as described in claim 9, characterized in that: Step S4 specifically includes: S4.1: Pair multi-view images The input pose encoder extracts motion features, where: t and t' For adjacent time frames; S4.2: Through the adaptive weight learning unit, the contribution of different viewpoints to motion estimation is allocated using learnable parameters, and the motion features of each viewpoint are weighted and fused; S4.3: Output the relative pose between adjacent time frames ; Step S5 specifically includes: S5.1: Based on prediction depth Camera internal parameters , Given the relative pose, construct the geometric projection mapping: pijt←t′=Πijt←t′pit pijt ← t ′=P ijt ← t ′ pit , ; in: From the perspective of the target i At any moment t homogeneous pixel coordinates; To map to the source view j The corresponding pixel position at time t'; ; in: The coordinate transformation matrix takes values based on the reconstruction context as follows: During time-series reconstruction, ; During spatial reconstruction, ; hour, ; in: , The relative motion between adjacent time frames predicted by the network. , For camera external parameters; S5.2: Generate reconstructed image: ; in: From the source perspective j The image at time t'; S5.3: Constructing the photometric consistency error function: ; in: These are the weighting coefficients. and These are the values at the same pixel position in the target image and the reconstructed image, respectively; S5.4: Calculate separately: Temporal photometric loss: ; Spatial photometric loss: ; Spacetime combined photometric loss: ; Consistency loss in multi-view reconstruction: ; S5.5: Calculate the weighted sum of joint photometric losses: ; in: .
Citation Information
Patent Citations
Three-dimensional target detection method, device and equipment
CN117333524A
Lightweight monocular depth estimation method and device based on self-supervised deep learning
CN119494866A