Three-dimensional depth estimation method combining hierarchical prior and geometric consistency point
Through the dual pyramid structure and extended search strategy, the problem of insufficient accuracy in these areas of traditional MVS technology is solved, and high-precision and high-efficiency three-dimensional reconstruction is achieved.
Patent Information
- Application Number
- CN202510190875.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional multi-view stereo matching (MVS) technology lacks depth estimation accuracy in low-texture areas and complex scenes, easily fall into local optimal solutions, and is difficult to effectively deal with occlusion and lighting changes.
The dual pyramid structure is adopted, through the coordinated optimization of the cost pyramid and the depth pyramid, combined with extended search strategies and geometric consistency points, the full utilization of multi-scale characteristics is achieved, taking into account global consistency and local details retention.
It significantly improves the depth estimation accuracy and robustness of low-texture areas, enhances the adaptability of complex scenarios, achieves the balance between global consistency and local details, and improves the quality and efficiency of three-dimensional reconstruction.
Smart Images

Figure BDA0005279963190000032 
Figure BDA0005279963190000053 
Figure BDA0005279963190000121
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vision and image processing, and specifically refers to a three-dimensional depth estimation method combining hierarchical priors and geometrically consistent points. Background Art
[0002] Three-dimensional reconstruction (3D reconstruction) technology is used to generate three-dimensional models from two-dimensional images of multiple viewpoints, and is one of the important applications in computer vision and image processing. Multi-View Stereo (MVS) technology is one of the core methods of three-dimensional reconstruction and is widely used in fields such as cultural heritage protection, robot navigation, virtual reality, and augmented reality. MVS technology extracts disparity information from multiple images and combines depth estimation algorithms to restore three-dimensional scenes. However, three-dimensional reconstruction of low-texture regions has always been a major challenge in MVS technology, especially in scenes with complex lighting, occlusion, or scarce texture. The depth estimation accuracy of traditional methods is significantly affected.
[0003] Traditional MVS methods mainly rely on photometric consistency and geometric consistency to establish matching relationships between views. Specifically, traditional algorithms find corresponding pixels in different views and use disparity calculation to generate depth maps. This process is usually based on photometric consistency (i.e., brightness similarity between images) and geometric consistency (i.e., consistency of spatial positions). The structure of these methods usually includes the following parts:
[0004] Feature extraction: Extract feature points from multi-view images. These feature points usually rely on photometric consistency to judge the matching relationship between pixels.
[0005] Matching process: Based on the results of feature matching, calculate the geometric error between images and further optimize the matching accuracy.
[0006] Depth estimation: Calculate the depth map through the disparity information obtained by matching and generate a three-dimensional model.
[0007] However, traditional methods have significant drawbacks in the following aspects:
[0008] 1. Difficulty in matching low-texture regions
[0009] In low-texture regions (such as flat, uniform surfaces) and complex scenes, photometric consistency often fails to provide sufficient matching information because in these regions, the pixel values in the images change little, resulting in a decrease in the accuracy of feature point extraction. Therefore, the depth estimation of traditional methods in these regions is often inaccurate, and the generated three-dimensional models lack details.
[0010] 2. Local optimal solution problem
[0011] Traditional methods usually perform depth estimation through feature matching within a local area, which is prone to falling into local optimal solutions, especially at the edges, occluded areas, and texture-poor areas of images, resulting in error accumulation and instability in the depth map.
[0012] 3. Insufficient handling of occlusion and illumination changes
[0013] In multi-view reconstruction, occlusion and illumination changes are inevitable phenomena. Traditional MVS methods often fail to effectively handle these problems. Especially under complex illumination conditions, the depth estimation results are easily affected, leading to a reduction in the quality of 3D reconstruction.
[0014] 4. Computational redundancy and efficiency issues
[0015] The calculations of traditional methods usually focus on local areas of images, resulting in a waste of computing resources. Especially in complex scenes, the computational cost is high and the processing speed is slow.
[0016] In response to the above problems, various improvements have been made in the prior art, such as:
[0017] 1. Optimization based on global consistency
[0018] Some methods propose to improve the depth estimation of low-texture areas through global optimization methods. For example, joint optimization is carried out using photometric consistency and geometric consistency constraints to improve the reconstruction accuracy of low-texture areas. However, these methods still rely on limited feature point matching and cannot fully solve the problems of occlusion and illumination changes in complex scenes.
[0019] 2. Utilization of non-local information
[0020] To avoid the trap of local optimal solutions, some methods introduce non-local information to enhance the robustness of matching. These methods improve the accuracy of depth estimation by increasing the sampling range of regions and using a wider range of pixel information. However, these methods often face computational efficiency problems, especially when dealing with large-scale data, the computational cost is high.
[0021] 3. Multi-scale optimization
[0022] Another type of method improves the accuracy of depth estimation by optimizing at different scales. These methods optimize the depth map from coarse to fine level by level, aiming to gradually recover the depth information of low-texture areas. Although this method improves the reconstruction quality to a certain extent, due to the dependence on high-resolution images, its computational cost is high and there are still certain limitations in detail recovery. Summary of the invention
[0023] The technical problem to be solved by the present invention is to provide a method for solving the depth estimation problem in 3D reconstruction, especially the balance problem of accuracy and efficiency in low-texture areas and complex scenes. Through a dual-pyramid structure with a cost pyramid and a depth pyramid as the core, and through the collaborative optimization of the two, the full utilization of multi-scale characteristics is achieved, taking into account global consistency and local detail retention, and significantly improving the quality and efficiency of 3D reconstruction. A 3D depth estimation method combining hierarchical prior and geometrically consistent points.
[0024] To solve the above technical problems, the technical solution provided by the present invention is: A 3D depth estimation method combining hierarchical prior and geometrically consistent points, characterized in that it includes the following contents:
[0025] Method 1: The specific implementation of the dual-pyramid structure;
[0026] Method 2: The specific implementation of the extended search strategy;
[0027] Method 3: The specific implementation of the hierarchical prior strategy;
[0028] Method 4: The specific implementation of geometrically consistent points for planar prior.
[0029] Further, the specific implementation of the dual-pyramid structure includes the following steps:
[0030] 1) Construction and optimization of the cost pyramid;
[0031] 2) Construction and optimization of the depth pyramid;
[0032] 3) Collaborative optimization of the dual pyramid.
[0033] Further, the construction and optimization of the cost pyramid includes the following steps:
[0034] 1) Multi-scale feature extraction: The image with the original resolution is downsampled level by level to generate a pyramid structure containing multi-scale images, defined as where s = 0 represents the coarsest scale and s = n represents the original resolution;
[0035] 2) Cost matrix calculation: At each scale, calculate the photometric consistency cost c photo and the geometric consistency cost C geom , and generate a comprehensive cost matrix through weighted fusion:
[0036] C s (p, h) = α · C photo (p, h) + β · C geom (p, h),
[0037] where α and β are weight parameters for photometric and geometric consistency;
[0038] 3), Gradual optimization: Optimize the cost matrix layer by layer from coarse to fine, providing reliable optimization guidance for the depth pyramid.
[0039] Furthermore, the construction and optimization of the depth pyramid include the following steps:
[0040] 1), Initial depth estimation: At the coarsest scale of the cost pyramid, select the cost optimal point using the comprehensive cost matrix to generate the initial depth map D 0 ;
[0041] 2), Gradual propagation and optimization: Upsample the depth map from the current scale s to the next scale s + 1 as the initial value D s+1 , Embed a joint bilateral propagator in each scale to repair the low-texture area and refine the edge depth information simultaneously;
[0042] 3), Final depth output: Complete the final detailed optimization at the highest resolution (s = n) to generate a high-precision depth map with global consistency and detail preservation.
[0043] Furthermore, the co-optimization of the double pyramid means that during the optimization process, the cost pyramid and the depth pyramid form a closed-loop optimization mechanism. The cost pyramid provides guidance for depth update, and the depth pyramid in turn corrects the cost matrix, gradually improving the accuracy and consistency of global depth estimation.
[0044] Furthermore, the specific implementation of the extended search strategy includes the following steps:
[0045] 1), Non-local expandable sampling mode: Set the initial sampling radius as R init , and dynamically adjust it to the maximum radius R during the matching cost evaluation process max ;
[0046] 2), Dynamic expansion strategy: Based on the distribution of the cost matrix C(p, h), dynamically generate a candidate hypothesis set:
[0047]
[0048] 3), Non-local optimization mechanism: Abandon the sampling points within the radius R of the central region local , and preferentially sample the non-local region far from the current pixel to ensure more robust depth estimation in the low-texture area.
[0049] Furthermore, the specific implementation of the hierarchical prior strategy includes the following steps:
[0050] 1), Prior generation in the low-resolution stage: Use the extended search and geometric consistency points to generate the plane prior model Pplane , determine the plane parameters by fitting the sparse point set G consistent ;
[0051] 2), Gradual propagation and optimization: Use the joint bilateral propagator to propagate the plane prior from low resolution to high resolution, refine the model and repair complex regions:
[0052] P prop (p)=∑ q∈N(p) w s (p,q)·w r (p,q)·P prior (q),
[0053] 3), Embedded depth estimation: In the depth optimization process at each scale, construct an optimization objective function by combining the plane prior:
[0054] C(p,h)=C photo (p,h)+λ·||h - P plane (p)||.
[0055] Furthermore, the specific implementation of the geometric consistency points for the plane prior includes the following steps:
[0056] 1), Generation of geometric consistency points: Screen out the geometric consistency point set G consistent by calculating the multi-view projection error, satisfying the following conditions:
[0057] Δe = ||p ref - Π i (P(h))||,
[0058] 2), Generation of the plane model: Perform triangulation calculation on the geometric consistency point set and generate the plane prior model P plane by least squares fitting;
[0059] 3), Dynamic optimization and detail repair: During the depth estimation process, dynamically embed the plane prior into the optimization objective, correct the depth assumptions in low-texture regions, and further repair the depth information in complex edge regions at the high-resolution stage.
[0060] After adopting the above method, the present invention has the following advantages:
[0061] 1. Collaborative optimization framework: The dual pyramid structure and the hierarchical prior strategy work together in multi-scale optimization, combined with the extended search strategy and geometric consistency points, constituting a closed-loop system for global depth estimation and local detail restoration;
[0062] 2. Significant improvement in depth estimation for low-texture regions: By using an extended search strategy and a geometric consistency point generation method, the matching performance of low-texture regions is optimized, avoiding the problem of traditional photometric consistency failure, and achieving depth estimation with higher robustness and accuracy;
[0063] 3. Enhanced adaptability to complex scenes: In occluded regions, complex lighting scenes, and large-scale uniform surfaces, by combining multi-scale prior propagation and non-local optimization, the reliability and global consistency of depth estimation results are ensured;
[0064] 4. Balance between global smoothness and local details: The hierarchical prior strategy constructs a global plane model at a coarse scale and gradually refines the detail regions at a fine scale, providing a high-quality 3D reconstruction solution from smooth surfaces to complex structures. Detailed implementation manner
[0065] The present invention will be further described in detail below.
[0066] I. Introduction of the dual pyramid structure
[0067] The present invention proposes an innovative dual pyramid structure for solving the depth estimation problem in 3D reconstruction, especially the balance between accuracy and efficiency in low-texture regions and complex scenes. The dual pyramid structure takes the cost pyramid and the depth pyramid as the core. Through the collaborative optimization of the two, the full utilization of multi-scale characteristics is achieved, taking into account global consistency and local detail retention, and significantly improving the quality and efficiency of 3D reconstruction.
[0068] A. Cost pyramid: The cost pyramid extracts photometric consistency and geometric consistency information from multi-scale image features to construct a cost matrix optimized layer by layer, providing reliable guidance for depth estimation. The specific implementation includes the following steps:
[0069] 1). Construction of multi-scale feature maps: The original image is downsampled layer by layer to generate a pyramid structure containing multi-scale images. The scale sequence is defined as where s = 0 represents the coarsest scale and s = n represents the original resolution.
[0070] At each scale, multi-level feature maps are extracted from the image for calculating the cost matrix.
[0071] 2). Calculation of the cost matrix:
[0072] The photometric consistency cost C photo : By calculating the photometric difference between the reference image and the source image, the matching cost of pixel points is measured:
[0073]
[0074] The geometric consistency cost Cgeom : Combine the front and back projection errors of multiple views to further constrain the geometric consistency of the cost:
[0075] C geom (p, h) = min(|d f - d b |, ∈),
[0076] where d f and d b are the depth values of the forward and backward projections respectively, and ∈ is the error tolerance threshold.
[0077] 3) Cost matrix optimization: Combine the photometric consistency and geometric consistency costs to generate a comprehensive cost matrix {C s}:
[0078] C s (p, h) = α · C photo (p, h) + β · C geom (p, h),
[0079] where α and β are weight coefficients used to balance the influence of photometric and geometric costs on the final depth estimation.
[0080] Optimize the cost matrix layer by layer, extract global features at low-resolution scales, and improve the matching accuracy of local details at high-resolution scales.
[0081] B. Depth pyramid:
[0082] Under the guidance of the cost pyramid, the depth pyramid optimizes the depth map from coarse to fine level by level, gradually solves the depth estimation problem in low-texture areas, and at the same time preserves edge details and complex structures. The specific implementation steps are as follows:
[0083] 1) Initial depth estimation
[0084] At the coarsest scale (s = 0) of the cost pyramid, use the comprehensive cost matrix {C 0} to extract the depth hypothesis with the optimal cost and generate the initial depth map {D 0}.
[0085] 2) Gradual propagation and optimization
[0086] Depth propagation: Upsample the depth map {D s} at the current scale to the next scale {D s+1} as the initial depth value.
[0087] Dynamic repair: Combine the cost difference map provided by the cost pyramid to repair the propagated depth map {D s+1}, especially improving the reliability of the depth in low-texture areas.
[0088] Fine-grained optimization: The optimal cost point of the cost pyramid is used to guide the fine-grained optimization of the depth map, ensuring the global consistency and local accuracy of the update results at the current scale.
[0089] Final depth output
[0090] At the finest scale (s=n), the features of the high-resolution image are used to optimize the details and enhance the edges of the depth map, and the final high-quality depth map {D n}.
[0091] C. Overall collaborative optimization of the double pyramid structure
[0092] A collaborative optimization relationship is formed between the cost pyramid and the depth pyramid.
[0093] 1) Cost-guided depth: The cost pyramid provides the optimal cost hypothesis for updating the depth pyramid, ensuring the reliability of depth estimation.
[0094] 2) Deep feedback cost: The optimization result of the deep pyramid is used to correct the error distribution of the cost pyramid and improve the accuracy of the cost matrix.
[0095] 3) Closed-loop optimization: The two form a feedback closed loop at multiple scales, transmitting and correcting information step by step to achieve a dynamic optimization process from coarse to fine.
[0096] D. Technical advantages of the double pyramid structure
[0097] 1) Balance global consistency with local details
[0098] The cost pyramid captures global information at a low-resolution scale, and the depth pyramid optimizes local details at a high-resolution scale. The two work together to achieve the unification of global depth estimation and local refinement.
[0099] 2) Enhanced robustness in low-texture areas
[0100] The cost pyramid combines photometric consistency and geometric consistency costs to establish reliable cost guidance in low-texture areas, significantly improving the accuracy and robustness of depth estimation.
[0101] 3) Adaptability to occlusion and lighting changes
[0102] The cost pyramid can effectively handle the occlusion problem between multiple views and enhance the matching ability in areas with changing illumination.
[0103] 4) Efficient computing and resource optimization
[0104] The dual pyramid structure decomposes complex problems through step-by-step optimization and reduces redundant calculations, making it particularly suitable for efficient implementation under the GPU architecture.
[0105] By introducing the dual pyramid structure, the present invention achieves high-precision, high-robustness depth estimation in low-texture regions during 3D reconstruction and efficient computation in complex scenes, and solves the technical bottlenecks in global consistency and local detail retention in the prior art.
[0106] E. Features and advantages of introducing the dual pyramid structure
[0107] This application proposes an innovative dual pyramid structure, including a cost pyramid and a depth pyramid, which cooperate with each other to form a multi-scale depth estimation framework from coarse to fine. This structure effectively solves problems such as difficult depth estimation in low-texture regions and insufficient restoration of details in complex scenes, and achieves a balance between global consistency and local detail retention.
[0108] 1). Features
[0109] 1. Cost pyramid
[0110] Based on the multi-scale feature maps, photometric consistency and geometric consistency costs are calculated layer by layer to generate a multi-scale cost matrix for guiding depth estimation.
[0111] Through dynamic cost optimization, the cost pyramid can capture global matching information at the low-resolution stage and provide precise detail guidance at the high-resolution stage.
[0112] 2. Depth pyramid
[0113] Depth maps are constructed from coarse to fine level by level, and the depth update at each scale is guided by the cost pyramid for optimization.
[0114] Through the joint bilateral propagator for progressive optimization of depth, depth correction in low-texture regions and detail restoration in edge regions are achieved.
[0115] 3. Cooperative optimization of the dual pyramid
[0116] The cost pyramid and the depth pyramid interact with each other to form a closed-loop optimization mechanism. The cost pyramid provides cost-optimal guidance information for the depth pyramid, and the optimization result of the depth pyramid in turn corrects the cost matrix of the cost pyramid, gradually improving the global consistency and local accuracy of depth estimation.
[0117] 2). Advantages and beneficial effects
[0118] 1. Unification of global consistency and local details
[0119] The dual pyramid structure achieves a perfect combination of global consistency and local optimization by capturing global information at the coarse scale and restoring detail information at the fine scale, and performs excellently especially in low-texture regions.
[0120] The cost pyramid provides guidance for global cost optimization, and the depth pyramid repairs details step by step at the high-resolution stage, significantly improving the depth estimation effect in complex scenes.
[0121] 2. Full utilization of multi-scale characteristics
[0122] At the coarse scale stage, the cost pyramid covers low-texture areas through large-scale cost optimization to ensure the depth accuracy of smooth surfaces; at the fine scale stage, the depth pyramid dynamically corrects depth hypotheses in combination with the cost matrix, enhancing the reconstruction accuracy of edge regions and complex structures.
[0123] 3. Enhancement of depth estimation in low-texture areas
[0124] By combining photometric consistency and geometric consistency costs, the cost pyramid can generate a reliable cost matrix in low-texture areas, providing high-quality optimization guidance for the depth pyramid and significantly improving the robustness and accuracy of depth estimation in low-texture areas.
[0125] 4. Adaptability to complex scenes
[0126] The dual pyramid collaborative optimization method has strong adaptability in occluded areas and complex lighting scenes, ensuring the reliability of depth estimation results globally.
[0127] 5. Efficient computing architecture
[0128] Through the step-by-step optimization framework, the complex depth estimation problem is decomposed into multi-level sub-problems from coarse to fine, reducing computational redundancy. At the same time, it adapts to the GPU parallel architecture, improving the overall operation efficiency of the algorithm.
[0129] By introducing the dual pyramid structure, this application has achieved a technical breakthrough in 3D reconstruction, especially showing significant advantages in depth estimation of low-texture areas and complex scenes, providing a strong technical guarantee for high-quality 3D reconstruction.
[0130] I. Specific implementation of the dual pyramid structure
[0131] 1. Construction and optimization of the cost pyramid
[0132] Multi-scale feature extraction:
[0133] The original-resolution image is downsampled step by step to generate a pyramid structure containing multi-scale images, defined as where s = 0 represents the coarsest scale and s = n represents the original resolution.
[0134] Cost matrix calculation:
[0135] At each scale, calculate the photometric consistency cost C photoand the geometric consistency cost C geom , and generate a comprehensive cost matrix through weighted fusion:
[0136] C s (p, h) = α · C photo (p, h) + β · C geom (p, h),
[0137] where α and β are the weight parameters for photometric and geometric consistency.
[0138] Hierarchical optimization:
[0139] Optimize the cost matrix layer by layer from coarse to fine, providing reliable optimization guidance for the depth pyramid.
[0140] 2. Construction and optimization of the depth pyramid
[0141] Initial depth estimation:
[0142] At the coarsest scale of the cost pyramid, select the optimal cost point using the comprehensive cost matrix to generate the initial depth map D 0 .
[0143] Hierarchical propagation and optimization:
[0144] Upsample the depth map from the current scale s to the next scale s + 1 as the initial value D s+1 .
[0145] Embed a joint bilateral propagator in each scale to repair low-texture regions and refine edge depth information simultaneously.
[0146] Final depth output:
[0147] Complete the final detailed optimization at the highest resolution (s = n) to generate a high-precision depth map with global consistency and detail retention.
[0148] 3. Cooperative optimization of the double pyramid
[0149] During the optimization process, the cost pyramid and the depth pyramid form a closed-loop optimization mechanism: the cost pyramid provides guidance for depth update, and the depth pyramid in turn corrects the cost matrix, gradually improving the accuracy and consistency of global depth estimation.
[0150] II. Extended search
[0151] The present invention proposes an innovative extended search method, which significantly improves the accuracy and robustness of depth estimation by utilizing non-local information to assist multi-view stereo matching (MVS), especially for the geometric recovery problem in low-texture regions. The extended search method effectively avoids the limitation of local optimal solutions and optimizes the overall effect of multi-view matching through non-local expandable sampling patterns, dynamic expansion strategies, and non-local optimization strategies.
[0152] A. Non-local expandable sampling pattern
[0153] 1). Adaptive sampling region
[0154] Traditional methods usually sample within a fixed local region, resulting in a lack of sufficient information support in low-texture regions and thus falling into local optimal solutions. The present invention proposes an expandable sampling pattern that can dynamically adjust the size of the sampling region according to the matching cost.
[0155] Definition of sampling region: The initial sampling region is defined as R init , and in each iteration, the radius R is dynamically adjusted according to the optimization degree of the matching cost. When the cost matrix indicates a lack of effective information in the current region, it is gradually expanded to a larger range R max .
[0156] 2). Combination of non-local expansion and local search
[0157] Within the local region, high-density sampling is performed to converge quickly;
[0158] Within the expanded region, non-local information is used to increase candidate hypothesis points, especially for low-texture regions and occluded regions, providing additional depth information support.
[0159] 3). Implementation steps
[0160] Initial sampling: Randomly sample a candidate point set within the initial search radius R init ;
[0161] Region expansion: Based on the photometric consistency and geometric consistency cost matrices, filter out the pixel regions with higher costs and gradually expand their sampling radii;
[0162] Dynamic adjustment: If the cost of the non-local sampling region is still higher than the preset threshold, further expand the search range until the matching cost meets the conditions or reaches the maximum search radius R max .
[0163] B. Expansion strategy
[0164] 1). Dynamic candidate hypothesis generation
[0165] The expansion strategy is guided by photometric consistency and geometric consistency, and dynamically adjusts the candidate hypothesis point set.
[0166] In each iteration, a new candidate hypothesis set is generated through the optimal points of the comprehensive cost matrix C(p, h), and random initialization and multi-view matching information are introduced to ensure that the candidate points cover the regions where the global optimal solution may appear.
[0167] 2) Expansion judgment based on the cost matrix
[0168] Using the matching cost distribution of each pixel in the cost matrix, the matching confidence of the candidate points is statistically calculated. If the confidence is insufficient, region expansion is initiated.
[0169] Judgment rule: Among them, H represents the current candidate hypothesis set, and τ extend is the expansion trigger threshold. Weighted voting is performed on the candidate points in the expansion region, and the optimal hypothesis point is selected to update the depth value.
[0170] 3) Voting mechanism
[0171] Weighted voting is performed on the candidate hypothesis points in the expansion region, and finally the hypothesis with the most votes is selected as the updated value of the depth hypothesis in the expansion region: h best = argmax h ∑ p∈R w p ·C(p, h),
[0172] C. Non-local optimization strategy
[0173] 1) Non-local region sampling
[0174] To further avoid local information redundancy, the present invention proposes a non-local strategy of abandoning the sampling points within the radius R local of the central pixel, and concentrating the sampling on more distant non-local regions:
[0175] R non-local = [R min , R max ,
[0176] Combining non-local sampling points with local sampling points makes the algorithm more adaptable in low-texture regions and complex scenes.
[0177] 2) Non-local candidate point guidance
[0178] Within the non-local region, the global optimal point of the matching cost matrix is used as the initial hypothesis, and high-confidence sampling update is performed on the expansion region.
[0179] The introduction of non-local points can effectively improve the depth estimation quality in low-texture regions and avoid local optimal traps.
[0180] 3), Dynamic non-local information fusion
[0181] Combine non-local information to perform dynamic weighted fusion on the candidate hypothesis points of the current pixel, and optimize the final depth hypothesis:
[0182] D(p) = ∑ h∈H w h ·h,
[0183] where w h is the weighting coefficient of the photometric consistency and geometric consistency costs.
[0184] D. Technical advantages and innovation points
[0185] 1), Efficient extended search mechanism
[0186] Dynamically adjust the size of the sampling area to ensure fast convergence in different scenarios and avoid redundant calculations.
[0187] 2), Full utilization of non-local information
[0188] By introducing non-local points, make up for the depth estimation defects in low-texture areas and occluded scenes, and significantly improve the effect of global depth optimization.
[0189] 3), Extended strategy combining photometric and geometric consistency
[0190] Use the multi-view matching cost matrix to dynamically guide the sampling of the extended area, effectively avoid the problem of local optimal solutions, and achieve global depth consistency.
[0191] 4), Adaptive optimization strategy
[0192] The method of the present invention can adaptively adjust the sampling range and strategy in low-texture areas, occluded areas, and complex illumination scenes to ensure high robustness and accuracy of depth estimation.
[0193] Through the introduction of extended search, the present invention has made a breakthrough in the depth estimation problem in low-texture areas and complex scenes, and significantly improved the overall quality and efficiency of 3D reconstruction.
[0194] E. Characteristics and advantages of using the extended search strategy
[0195] This application proposes an innovative extended search strategy. By combining non-local information with a dynamic region adjustment mechanism, it effectively solves the problem that traditional multi-view stereo matching (MVS) is prone to fall into local optimal solutions in low-texture areas and complex scenes. The extended search strategy significantly improves the robustness and accuracy of depth estimation through a non-local extensible sampling mode, a dynamic extension strategy, and a non-local optimization mechanism.
[0196] 1) Features
[0197] 1. Non-local extensible sampling mode
[0198] Provide an adaptive sampling area adjustment mechanism. The initial sampling area starts with a fixed radius R init and gradually expands to a larger radius R as the evaluation result of the matching cost max to ensure a wider search area coverage in complex scenarios.
[0199] By dynamically adjusting the sampling radius, the depth matching ability of the algorithm in low-texture areas and occlusion scenarios is enhanced.
[0200] 2. Dynamic expansion strategy
[0201] Based on the voting mechanism of the multi-view cost matrix, dynamically select the best hypothesis points to form a set of candidate hypothesis points.
[0202] When the local search fails to achieve the optimal cost, trigger area expansion. By analyzing the cost distribution in the current area, gradually introduce non-local information to avoid falling into local optimal solutions.
[0203] 3. Non-local optimization mechanism
[0204] Abandon the sampling points within the central pixel radius R local and concentrate on sampling non-local areas to improve the matching ability in low-texture areas.
[0205] Within the expanded area, use the global cost optimal point to guide the depth hypothesis update, effectively enhancing the robustness in low-texture areas.
[0206] 1) Advantages and beneficial effects
[0207] 1. Significantly improve the depth estimation quality in low-texture areas
[0208] The extended search strategy introduces more non-local information into the depth estimation process by expanding the sampling range, effectively making up for the deficiency of insufficient local information in low-texture areas.
[0209] The non-local optimization mechanism ensures high precision and high robustness of the depth estimation results in low-texture areas.
[0210] 2. Avoid falling into local optimal solutions
[0211] The dynamic expansion strategy adaptively adjusts the sampling area by analyzing the distribution of the cost matrix, gradually introducing more candidate hypothesis points to avoid the problem of incorrect matching dominated by local information.
[0212] The introduction of non-local sampling further enhances the global optimization ability of the algorithm, achieving higher depth estimation accuracy.
[0213] 3. Adaptability in complex scenarios
[0214] In scenarios with large occlusion areas and significant illumination changes, the extended search strategy can dynamically adjust the sampling range to ensure reliable depth estimation results in different regions.
[0215] The introduction of non-local information significantly improves the depth estimation effect in occlusion areas and effectively reduces the propagation of false matches.
[0216] 4. Efficient computation and resource optimization
[0217] The extended search strategy focuses on local high-density sampling in the initial stage, optimizes resource allocation through dynamic expansion and regional adjustment, and reduces unnecessary redundant computations.
[0218] Under the GPU parallel computing architecture, the dynamic adjustment mechanism of the extended search has strong adaptability and can significantly improve the running efficiency.
[0219] 5. Balancing global consistency and local details
[0220] The extended search strategy enhances global consistency through non-local optimization and dynamically adjusts the local sampling area to ensure the recovery of depth details in complex scenarios.
[0221] By using the extended search strategy, this application achieves a significant improvement in depth estimation in low-texture areas, occlusion areas, and complex illumination scenarios, providing a reliable guarantee for high-precision three-dimensional reconstruction of multi-view stereo matching.
[0222] II. Specific implementation of the extended search strategy
[0223] 1. Non-local expandable sampling mode
[0224] The initial sampling radius is set to R init , and it is dynamically adjusted to the maximum radius R max .
[0225] In each iteration, if the cost in the local area is higher than the set threshold, the extended search is triggered to gradually expand the sampling area to cover more candidate points.
[0226] 2. Dynamic expansion strategy
[0227] Based on the distribution of the cost matrix C(p, h), a candidate hypothesis set is dynamically generated:
[0228]
[0229] where T extendIt is the extended trigger threshold. Weighted voting is performed on the candidate points within the extended area to select the optimal hypothesis point to update the depth value.
[0230] 3. Non-local optimization mechanism
[0231] Abandon the sampling points within the radius R of the central area, and preferentially sample the non-local area far from the current pixel to ensure more robust depth estimation in the low-texture area. local
[0232] III. Hierarchical Prior Mining
[0233] The present invention proposes an innovative hierarchical prior mining method, which constructs and optimizes prior information step by step at different scales to assist in the recovery of 3D models, especially for the depth estimation problems in low-texture areas and complex surfaces. The hierarchical prior method adopts a coarse-to-fine optimization framework, generating smooth prior information using non-local reliable sparse correspondences in the early low-resolution stage, and repairing detailed information through a joint bilateral propagator in the high-resolution stage to achieve a balance between the accuracy of global depth estimation and the retention of details.
[0234] A. Hierarchical prior mining framework
[0235] The hierarchical prior mining framework mines different types of prior information at the coarse scale and the fine scale in a step-by-step optimized manner;
[0236] 1). Non-local prior mining in the low-resolution stage
[0237] At the coarsest scale (low resolution), based on the extended search strategy, a planar prior model is generated using reliable non-local sparse correspondences to expand the receptive field for estimating the depth of a large-scale smooth and uniform surface.
[0238] The mining of non-local prior is achieved by combining the photometric consistency and geometric consistency costs:
[0239] P prior (p) = w geo ·C geom (p) + w photo ·C photo (p),
[0240] where w gep and w photp are the weighting coefficients of geometric and photometric consistency.
[0241] 2). Detail depth recovery in the high-resolution stage
[0242] At the higher-resolution stage, based on the previously generated planar prior, the hypothesis is corrected and propagated to recover the detailed depth information.
[0243] When constructing the depth of details, a low-resolution planar model is embedded into the PatchMatch algorithm to optimize the depth estimation in complex edge regions.
[0244] 3) Prior information transfer between multiple scales
[0245] The planar prior information is gradually transferred between scales, and the prior is corrected and refined through a joint bilateral propagator to ensure a smooth transition from coarse to fine.
[0246] B. Multi-scale structure design
[0247] 1) Construction of multi-scale planar prior models
[0248] At each scale, a planar model is generated using extended search and geometrically consistent points and used as the prior information for the current scale.
[0249] The planar prior generated at the low-resolution stage has a better estimation accuracy for low-texture regions than the model directly constructed at high resolution.
[0250] 2) Step-by-step optimization strategy
[0251] Low-resolution optimization: Optimize the depth estimation of large-scale planar regions with the goal of global consistency.
[0252] Medium-resolution refinement: Combine the low-resolution prior model to perform local repair and optimization on medium scales.
[0253] High-resolution repair: Embed the prior model at the finest scale to recover details in complex structures and edge regions.
[0254] C. Iterative coarse-to-fine optimization process
[0255] The hierarchical prior achieves a dynamic balance between global depth estimation and detail repair through a coarse-to-fine iterative optimization framework:
[0256] 1) Low-resolution initialization
[0257] Apply the basic MVS method APD-MVS and extended search to generate initial hypotheses (depth and normal).
[0258] Downsample the initial hypotheses to the coarsest resolution and combine non-locally geometrically consistent points to generate a preliminary planar prior model.
[0259] 2) Planar prior propagation
[0260] Use a joint bilateral upsampler to propagate the planar prior from low resolution to high resolution, as shown in the following formula:
[0261] Pprop (p) = ∑ q∈N(p) w s (p, q)·w r (p, q)·P prior (q),
[0262] where w s is the spatial Gaussian weight, w r is the photometric range weight, and N(p) is the neighborhood of pixel p.
[0263] 3), Medium and high resolution optimization
[0264] At the medium resolution scale, new hypotheses are constructed using the propagated plane prior information and the current depth is optimized. The prior propagation and optimization process are repeated and gradually iterated to the original resolution.
[0265] 4), Detail enhancement and depth output
[0266] At the finest scale, the prior model is embedded in the APD-MVS algorithm to optimize the depth estimation in the edge and complex detail regions.
[0267] A globally consistent and detail-preserving high-quality depth map is output.
[0268] D, Technical advantages of the hierarchical prior
[0269] 1), Full exploration of non-local information
[0270] At the low resolution stage, a reliable non-local geometric consistency point spread receptive field is used to provide stronger support for low-texture regions.
[0271] 2), Dynamic balance of multi-scale smoothing and detail repair
[0272] The multi-scale structure allows the plane prior to be gradually optimized from global smoothing to local refinement, effectively balancing the depth estimation in low-texture regions and edge details.
[0273] 3), Efficient iterative optimization
[0274] The coarse-to-fine hierarchical optimization framework solves the large-scale smoothing problem in the initial stage, enhances details step by step in the later stage, and reduces redundant calculations and error propagation.
[0275] 4), High-precision reconstruction of edges and complex regions
[0276] The APD-MVS method embedded with the plane prior optimizes complex structures at the high resolution stage to ensure the accuracy of the 3D reconstruction results.
[0277] By introducing a hierarchical prior mining method, the present invention solves the accuracy problem of depth estimation in low-texture regions, while preserving the detailed structure, and has significant advantages in scenarios such as low-texture regions, complex scenes, and large-scale uniform surfaces.
[0278] E. Characteristics and advantages of the hierarchical prior strategy
[0279] This application proposes a hierarchical prior strategy. By gradually mining and optimizing non-local prior information at different scales, a depth estimation method is formed that gradually enhances from global smoothness to local details. This strategy takes hierarchical prior mining, multi-scale plane model construction, and dynamic optimization as the core, and effectively solves problems such as difficulties in depth estimation in low-texture regions and insufficient restoration of details in complex scenes.
[0280] 1). Characteristics
[0281] 1. Hierarchical prior mining framework
[0282] In the coarse-scale stage, by expanding the receptive field, reliable sparse geometric consistency points are mined from low-resolution non-local information to generate a globally smooth plane prior model.
[0283] In the fine-scale stage, combined with the assumptions corrected in the previous stage, the prior information is gradually refined to enhance the depth estimation of edges and complex structures.
[0284] 2. Construction and propagation of multi-scale plane models
[0285] The plane prior models generated at each scale are all fitted in combination with geometric consistency points to ensure the stability and robustness of the models in low-texture regions.
[0286] The joint bilateral propagator is used to transfer prior information between multi-scales, and the plane prior model is refined at the fine scale to avoid the accumulation of incorrect matches.
[0287] 3. Gradual optimization and dynamic repair
[0288] The plane prior model is dynamically embedded in depth estimation. By introducing plane constraints, the depth assumptions in low-texture regions are corrected.
[0289] In the gradual optimization framework, the low-resolution global prior guides the depth optimization in the high-resolution stage to ensure the global consistency and local accuracy of the estimation results.
[0290] 1). Advantages and beneficial effects
[0291] 1. Dynamic balance between global and local optimization
[0292] The hierarchical prior strategy captures global smooth features in the low-resolution stage, propagates them step by step to high resolutions, and finally achieves a dynamic balance between detail restoration and depth estimation.
[0293] Especially in low-texture regions, the step-by-step optimization of the hierarchical prior can significantly improve the quality of depth estimation.
[0294] 2. Enhancement of depth estimation in low-texture regions
[0295] The planar prior model generated in the coarse-scale stage can effectively constrain the depth estimation of large-scale uniform surfaces, avoiding the problem of traditional methods failing in low-texture regions.
[0296] Guided by geometrically consistent points, the planar prior model shows higher robustness in depth recovery for low-texture regions.
[0297] 3. Ability to preserve edges and details
[0298] In the fine-scale stage, the planar prior model is refined using a bilateral propagator to effectively repair the depth information in complex structures and edge regions, achieving the unity of global consistency and local detail preservation.
[0299] 4. Efficiency of multi-scale optimization
[0300] The hierarchical prior strategy decomposes the depth estimation problem into multi-level optimization tasks by gradually transmitting prior information, reducing computational redundancy, and significantly improving the adaptability and efficiency of the algorithm.
[0301] 5. Adaptability to complex scenes
[0302] Through expanding the receptive field and mining non-local information, the hierarchical prior strategy can achieve highly robust depth estimation results in occluded regions, low-light scenes, and large-scale uniform surfaces.
[0303] By adopting the hierarchical prior strategy, this application significantly improves the depth estimation ability in low-texture regions and complex scenes, while achieving a dynamic balance between global consistency and detail preservation, providing strong technical support for high-quality 3D reconstruction.
[0304] III. Specific implementation of the hierarchical prior strategy
[0305] 1. Prior generation in the low-resolution stage
[0306] Generate the planar prior model P using extended search and geometrically consistent points plane , and determine the planar parameters (depth d consistent and normal vector n p ) by fitting the sparse point set G p .
[0307] The plane prior model generated in the smooth region at the low-resolution stage is used for global depth estimation constraint.
[0308] 2. Hierarchical Propagation and Optimization
[0309] Use the joint bilateral propagator to propagate the plane prior from low resolution to high resolution, refine the model and repair complex regions:
[0310] P prop (p) = ∑ q∈N(p) w s (p, q)·w r (p, q)·P prior (q),
[0311] where w s and w r are spatial and photometric weights, and N(p) is the neighborhood range.
[0312] 3. Embedded Depth Estimation
[0313] During the depth optimization process at each scale, construct an optimization objective function by combining the plane prior:
[0314] C(p, h) = C photo (p, h) + λ·||h - P plane (p)||,
[0315] where λ is the plane constraint weight coefficient.
[0316] IV. Geometrically Consistent Points for Plane Prior
[0317] The present invention proposes an innovative method for generating a plane prior based on geometrically consistent points. By using the multi-view projection error to judge geometrically consistent points, it replaces the method that relies on triangulation of sparse matching points in traditional plane priors, thereby overcoming the limitations of depth estimation in low-texture regions and significantly improving the quality and robustness of depth estimation.
[0318] A. Limitations of Traditional Methods
[0319] 1). Generation of Sparse Matching Points Depends on Photometric Consistency
[0320] Traditional plane prior generation methods need to generate reliable sparse matching points from multiple views using photometric consistency. However, in low-texture regions, due to the small intensity differences between pixels, photometric consistency is difficult to provide sufficient matching constraints, resulting in the quality and quantity of sparse matching points being insufficient to meet the requirements for generating a plane model.
[0321] 2). Accuracy of Plane Model is Limited
[0322] The uneven distribution of sparse matching points will affect the stability of the plane model generated by triangulation. Especially in large-scale uniform surfaces and complex edge regions, the plane models generated by traditional methods are often not robust enough, resulting in a decrease in the accuracy of depth estimation.
[0323] B. Plane prior generation method based on geometrically consistent points
[0324] To solve the above problems, the present invention proposes a plane prior generation method based on geometrically consistent points of projection error, which uses multi-view geometric constraints instead of photometric consistency as the basis for generating sparse matching points, improving the accuracy and reliability of the plane model. The specific implementation steps are as follows:
[0325] 1). Calculation of multi-view projection error
[0326] For each pixel point p, assuming its three-dimensional coordinates under depth hypothesis h are P(h), then the projection position of this point in different views is Π i (P(h)), where Π i represents the projection matrix of the i-th view.
[0327] The projection error Δe is defined as the difference between the reference view pixel point p ref and the source view pixel point p src under the depth hypothesis:
[0328] Δe = ||p ref - Π i (P(h))||,
[0329] 2). Selection of geometrically consistent points
[0330] Set the projection error threshold τ geo . If the projection error of a certain pixel point is less than the threshold, it is determined as a geometrically consistent point:
[0331]
[0332] The point set G consistent selected through geometric consistency judgment is more evenly distributed and has a wider coverage range, especially suitable for low-texture regions.
[0333] 3). Construction of the plane model
[0334] For geometrically consistent points, use the triangulation method to generate three-dimensional coordinates and perform plane fitting on these points to generate a plane model. Plane fitting uses the least squares method to estimate the plane parameters (depth d p and normal vector n p ):
[0335] P plane : n p·(X - X 0 ) = 0,
[0336] where X 0 is the center point of the fitting point set, and n p is the plane normal vector.
[0337] 4), Dynamic optimization and detail repair
[0338] Embed the generated plane model into the depth estimation process to constrain and optimize the depth hypotheses in low-texture regions.
[0339] In the high-resolution stage, use the joint bilateral propagator to repair the details of the plane model and enhance the depth estimation in the edge regions.
[0340] C, Technical advantages and improvements
[0341] 1), Geometric consistency points replacing photometric consistency
[0342] The selection of geometric consistency points is based on the multi-view projection error, which can effectively avoid the problem of photometric consistency failure in low-texture regions and ensure the quality and uniform distribution of sparse points.
[0343] 2), Generation of a highly robust plane model
[0344] The plane model generated using geometric consistency points shows higher robustness in large-scale smooth regions and complex scenes, effectively reducing mismatches and fluctuations in depth estimation.
[0345] 3), Optimization ability in low-texture regions
[0346] Geometric consistency points expand the coverage of the plane model in low-texture regions, significantly improving the accuracy of depth estimation, especially in large-scale uniform surfaces and occluded areas.
[0347] 4), Preservation of edges and details
[0348] Embed the plane model generated by geometric consistency points into the multi-scale optimization process of depth estimation, and achieve high-precision recovery of edge regions and details through step-by-step refinement.
[0349] D, Implementation steps
[0350] 1), Selection of initial geometric consistency points
[0351] Calculate the projection error for each pixel point in each view and screen out the initial set of geometric consistency points G init .
[0352] 2), Generation of the plane prior model
[0353] Utilize the geometrically consistent point set G init , and generate a preliminary plane model P through plane fitting plane .
[0354] 3), Plane model embedding depth estimation
[0355] During the depth estimation process, embed the plane prior into the cost function to restrict the solution space of depth hypotheses:
[0356] C(p, h) = C photo (p, h)+λ · ||h - P plane (p)||,
[0357] where λ is the weight coefficient of the plane prior
[0358] 4), Gradual optimization and repair
[0359] Under the multi-scale depth estimation framework, gradually optimize the depth values in low-texture regions and refine and repair the edge regions at the fine-scale stage
[0360] By introducing the plane prior method based on geometrically consistent points, the present invention shows significant performance improvement in depth estimation for low-texture regions and complex scenes, and exhibits higher robustness and accuracy especially in large-scale uniform surfaces and occluded regions
[0361] E, Characteristics and advantages of geometrically consistent points for plane prior
[0362] This application proposes a method for generating a plane prior based on geometrically consistent points, dynamically screening geometrically consistent points through multi-view projection error calculation, replacing the traditional method of generating plane prior relying on sparse matching points, and significantly improving the robustness and depth estimation accuracy of the plane model in low-texture regions and complex scenes
[0363] 1), Characteristics
[0364] 1. Selection of geometrically consistent points
[0365] Calculate geometrically consistent points based on multi-view projection error, and screen pixel points with errors lower than a preset threshold as geometrically consistent points to ensure the quality of sparse points and the uniformity of distribution
[0366] The projection error calculation formula is:
[0367] Δe = ||p ref - Π i (P(h))||,
[0368] where p ref is the pixel point in the reference image, Π i(P(h)) is the multi-view projection point under the deep hypothesis.
[0369] 2. Generation of high-quality plane model
[0370] Triangulation calculation is performed using geometrically consistent points, and a plane model is generated by least squares fitting, significantly improving the robustness and stability of the plane model.
[0371] The parameters of the plane model include depth d d and normal vector n p , which provides strong depth constraints in low-texture regions based on this.
[0372] 3. Dynamic embedding of plane prior
[0373] Embed the generated plane prior model into the multi-view matching cost function to impose geometric constraints on the depth hypothesis in low-texture regions:
[0374] C(p,h) = C photo (p,h) + λ·||h - P plane (p)||,
[0375] where λ is the plane constraint weight, and P plane represents the depth hypothesis of the plane model.
[0376] 4. Gradual optimization and detail repair
[0377] In a multi-scale framework, gradually optimize the plane prior model through a joint bilateral propagator to achieve dynamic repair and smooth transition from low resolution to high resolution.
[0378] In the high-resolution stage, further refine the depth estimation of edge regions and complex structures to enhance the ability to retain details.
[0379] 2). Advantages and beneficial effects
[0380] 1. High-robustness geometrically consistent points replacing photometric consistency
[0381] Geometrically consistent points depend on the calculation results of multi-view projection errors and are not affected by the failure of photometric consistency in low-texture regions, effectively enhancing the quality and coverage of sparse points.
[0382] Compared with traditional photometrically consistent points, geometrically consistent points perform better in large-scale uniform surfaces and occlusion regions.
[0383] 2. Significant improvement in depth estimation in low-texture regions
[0384] The plane prior model generated using geometrically consistent points can provide stable depth hypothesis constraints for low-texture regions, significantly improving the depth estimation accuracy and robustness of low-texture regions.
[0385] 3. Generation and Dynamic Optimization of High-Quality Plane Models
[0386] The plane model generated based on geometrically consistent points has higher robustness and can adapt to smooth regions and complex edges in complex scenes.
[0387] The dynamic optimization of the plane model in the multi-scale framework further ensures the continuity and accuracy of the depth estimation results from low resolution to high resolution.
[0388] 4. Enhancing Adaptability in Complex Scenes
[0389] In occluded regions and scenes with large illumination variations, geometrically consistent points effectively compensate for the limitations of photometrically consistent points and provide reliable support for depth estimation in complex scenes.
[0390] 5. Precise Recovery of Edges and Details
[0391] Embedding the plane prior model into the depth optimization process at the high-resolution stage can effectively recover the depth estimation results of edge and detail regions while avoiding the propagation of incorrect matches.
[0392] By introducing a method for generating plane priors based on geometrically consistent points, this application significantly enhances the depth estimation ability in low-texture regions and complex scenes, providing strong technical support for 3D reconstruction in low-light scenes, large-scale uniform surfaces, and complex edges.
[0393] IV. Specific Implementation of Geometrically Consistent Points for Plane Priors
[0394] 1. Generation of Geometrically Consistent Points
[0395] Through multi-view projection error calculation, a set of geometrically consistent points G is selected, consistent , satisfying the following conditions:
[0396] Δe = ||p ref - Π i (P(h))|| < τ geo ,
[0397] 2. Generation of Plane Model
[0398] The set of geometrically consistent points is triangulated and a plane prior model P is generated by least squares fitting. plane .
[0399] 3. Dynamic Optimization and Detail Repair
[0400] During the depth estimation process, the planar prior is dynamically embedded into the optimization objective to correct the depth hypothesis in low-texture regions, and the depth information in complex edge regions is further repaired at the high-resolution stage.
[0401] Through the above specific implementation manners, the present invention starts from four core innovation points: the dual pyramid structure, the extended search strategy, the hierarchical prior strategy, and the geometric consistency points, and realizes high-precision 3D reconstruction in low-texture regions and complex scenes. Those in the same field can reproduce the technical solution and achieve the expected effect according to the above steps, parameters, and processes.
[0402] The above description of the present invention and its implementation manners is not restrictive, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the gist of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. A 3D depth estimation method combining hierarchical priors and geometric consistency points, characterized by: It includes the following specific implementation methods: Method 1: Specific implementation of the double pyramid structure; Method 2: Specific implementation of the extended search strategy; Method 3: Specific implementation of the hierarchical prior strategy; Method 4: Specific implementation of geometric consistency points for plane priors.
2. The three-dimensional depth estimation method combining hierarchical priors and geometric consistency points according to claim 1, characterized in that: The specific implementation of the double pyramid structure includes the following steps: 1) Construction and optimization of cost pyramid; 2) Construction and optimization of deep pyramid; 3) Double pyramid collaborative optimization.
3. The three-dimensional depth estimation method combining hierarchical priors and geometric consistency points according to claim 2, characterized in that: The construction and optimization of the cost pyramid includes the following steps: 1) Multi-scale feature extraction: The original resolution image is downsampled step by step to generate a pyramid structure containing multi-scale images, which is defined as Where s = 0 represents the coarsest scale, and s = n represents the original resolution; 2) Cost matrix calculation: At each scale, calculate the photometric consistency cost C photo and the geometric consistency cost C geom , and generate a comprehensive cost matrix through weighted fusion: C s (p,h)=α·C photo (p,h)+β·C geom (p,h), Where α and β are the weight parameters of photometric and geometric consistency; 3) Level-by-level optimization: Optimize the cost matrix layer by layer from coarse to fine, providing reliable optimization guidance for the deep pyramid.
4. The three-dimensional depth estimation method combining hierarchical priors and geometric consistency points according to claim 2, characterized in that: The construction and optimization of the depth pyramid includes the following steps: 1) Initial depth estimation: At the coarsest scale of the cost pyramid, the comprehensive cost matrix is used to select the optimal cost point to generate the initial depth map D0; 2) Step-by-step propagation and optimization: Sample the depth map from the current scale s to the next scale s+1 as the initial value D s+1 , a joint bilateral propagator is embedded in each scale to inpaint low-texture areas while refining edge depth information; 3) Final depth output: Complete the final detail optimization at the highest resolution (s=n) to generate a high-precision depth map that is globally consistent and retains details.
5. The three-dimensional depth estimation method combining hierarchical priors and geometric consistency points according to claim 1, characterized in that: The dual-pyramid collaborative optimization is that during the optimization process, the cost pyramid and the depth pyramid form a closed-loop optimization mechanism, the cost pyramid provides guidance for depth update, and the depth pyramid in turn corrects the cost matrix, gradually improving the accuracy and consistency of global depth estimation.
6. The three-dimensional depth estimation method combining hierarchical priors and geometric consistency points according to claim 1, characterized in that: The specific implementation of the extended search strategy includes the following steps: 1) Non-local scalable sampling mode: the initial sampling radius is set to R init , dynamically adjusted to the maximum radius R during the matching cost evaluation process max ; 2) Dynamic expansion strategy: Based on the distribution of the cost matrix C(p, h), dynamically generate a set of candidate hypotheses: 3) Non-local optimization mechanism: abandon the central area radius R local The sampling points within the pixel are prioritized to sample non-local areas far away from the current pixel, ensuring that the depth estimation in low-texture areas is more robust.
7. The three-dimensional depth estimation method combining hierarchical priors and geometric consistency points according to claim 1, characterized in that: The specific implementation of the hierarchical prior strategy includes the following steps: 1) Prior generation at low resolution: Generate a plane prior model P using extended search and geometric consistency points plane , by fitting the sparse point set G consistent Determine plane parameters; 2) Step-by-step propagation and optimization: Use a joint bilateral propagator to propagate the plane prior from low resolution to high resolution, refine the model and repair complex areas: P prop (p)=Σ q∈N(p) w s (p,q)·w r (p,q)·P prior (q); 3) Embedded depth estimation: In the depth optimization process of each scale, the optimization objective function is constructed by combining the plane prior: C(p,h)=C photo (p,h)+λ·||h-P plane (p)||。 8. The three-dimensional depth estimation method combining hierarchical priors and geometric consistency points according to claim 1, characterized in that: The specific implementation of the geometric consistency point for plane prior includes the following steps: 1) Generation of geometrically consistent points: Through multi-view projection error calculation, the geometrically consistent point set G is screened out consistent , the following conditions are met: 2) Plane model generation: Triangulate the geometrically consistent point set and generate a plane prior model P by least squares fitting. plane ; 3) Dynamic optimization and detail restoration: In the depth estimation process, the plane prior is dynamically embedded into the optimization target to correct the depth assumption of the low-texture area, and the depth information of the complex edge area is further restored in the high-resolution stage.
Citation Information
Patent Citations
Three-dimensional reconstruction method based on texture information guidance and multi-scale prior assistance
CN117292056A
Multi-view dense three-dimensional reconstruction method based on normal information assistance
CN119131255A
Four dimensional reconstruction and characterization system
US20120070068A1