3D Reconstruction Method and System for High-Resolution Images

Through downsampling processing and mean-shift and superpixel segmentation algorithms, a priori depth map is generated, and combined with the prior constraint terms in the Patchmatch algorithm, the problem of mismatch in the weak texture area is solved, which significantly improves the accuracy of depth estimation and the integrity of the three-dimensional model.

CN114842145BActive Publication Date: 2025-06-24GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210497213.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-09
Publication Date
2025-06-24
Estimated Expiration
2042-05-09

AI Technical Summary

Technical Problem

The traditional three-dimensional reconstruction method based on photometric consistency measurement has a problem of mismatch in weak texture areas, resulting in the absence of depth values ​​and the inability to effectively restore the three-dimensional spatial information of large-area weak texture areas.

Method used

The downsampling process is used to combine the mean-shift algorithm and the superpixel segmentation algorithm to generate a priori depth map, and a priori constraint term is introduced in the Patchmatch algorithm. The depth map is calculated by combining photometric consistency and priori depth.

Benefits of technology

The accuracy of depth estimation in weak texture areas is significantly improved, the fuzzy loss problem caused by prior depth is avoided, and the mismatch problem of photometric consistency in weak texture areas can be better overcome, and a complete three-dimensional point cloud model can be obtained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114842145B_ABST
    Figure CN114842145B_ABST
Patent Text Reader

Abstract

The present invention provides a three-dimensional reconstruction method for high-resolution images, including: performing downsampling processing on the high-resolution images; clustering the images after downsampling processing, successively projecting the categories that retain 10% of the current image into the three-dimensional space, performing plane fitting on each category and back-projecting it to the corresponding category to assign depth values to obtain a mean-shift prior depth map; performing superpixel segmentation on the unprocessed image regions, successively projecting the retained superpixel blocks into the three-dimensional space, and performing plane fitting on each superpixel block to obtain the fitting plane back-projected to the corresponding superpixel block in the current image to assign depth values to obtain a superpixel prior depth map; synthesizing the obtained mean-shift prior depth map and superpixel prior depth map into a final prior depth map; calculating the prior depth values in the final prior depth map to obtain a final depth map, and performing three-dimensional reconstruction on the high-resolution images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a three-dimensional reconstruction method and system for high-resolution images. Background Art

[0002] The goal of a three-dimensional reconstruction system is to perform initial depth estimation using a set of input images with a certain degree of overlap, calculate the depth value corresponding to each pixel, and capture any scene. The obtained depth maps are fused into a dense point cloud, which can represent the scene distribution in three-dimensional space. Therefore, for each given input image, it is desired to calculate the depth estimation of each pixel, which essentially establishes the correspondence from a two-dimensional image to a three-dimensional space object. Three-dimensional reconstruction can also be understood as a process of recovering the three-dimensional information of space from the two-dimensional information of images.

[0003] Traditional three-dimensional reconstruction systems based on image information usually adopt photometric consistency metrics, such as bilateral weighted NCC, to evaluate different depth hypotheses. The Patchmatch algorithm is an efficient method for generating depth maps. For each pixel in each view, multiple depth hypotheses are tested, and the hypothesis that maximizes photometric consistency between the input views is selected. The Patchmatch algorithm based on photometric consistency metrics usually performs well in regions with rich textures and can meet the three-dimensional reconstruction requirements of most scenes. However, this method faces challenges in weakly textured image regions such as walls and floors, which may cause blurring because photometric consistency is prone to false matching in scenes with large areas of weak textures. To handle the outliers generated by fuzzy matching, explicit smoothing constraints are usually imposed on the depth estimation (based on the assumption that adjacent pixels should have similar depths). Even so, the depth values in large areas of weakly textured regions are still generally missing, which means that the three-dimensional space information in large areas of weakly textured regions has not been recovered, far from meeting the three-dimensional reconstruction requirements in large areas of weakly textured regions. Summary of the Invention

[0004] The present invention provides a three-dimensional reconstruction method and system for high-resolution images, aiming to overcome the false matching problem of photometric consistency in weakly textured regions, greatly improve the accuracy of depth estimation in weakly textured regions, and avoid the problem of detail blurring and loss caused by prior depth.

[0005] To achieve the above object, the present invention provides a three-dimensional reconstruction method for high-resolution images, including:

[0006] Step 1, perform downsampling processing on the high-resolution image to downsample the high-resolution image to different scales;

[0007] Step 2: Use the mean - shift algorithm to cluster the downsampled image, retain the categories that account for 10% or more of the current image area in the clustering results, project them into the three - dimensional space one by one, perform plane fitting on each category through the ransac algorithm, and project the fitted plane back into the corresponding category in the current image to update the depth values of the pixels in this category to obtain the mean - shift prior depth map;

[0008] Step 3: Use the super - pixel segmentation algorithm to segment the image area not processed by the mean - shift algorithm after downsampling, project the super - pixel blocks that account for 10% or more of the current image area in the segmentation results into the three - dimensional space one by one, and perform plane fitting on each super - pixel block through the ransac algorithm. Project the fitted plane back into the corresponding super - pixel block in the current image to update the depth values of the pixels in the super - pixel block to obtain the super - pixel prior depth map;

[0009] Step 4: Synthesize the mean - shift prior depth map and the super - pixel prior depth map obtained in Step 2 and Step 3 into the final prior depth map;

[0010] Step 5: Combine with the patchmatch algorithm to calculate the prior depth values in the final prior depth map to obtain the final depth map and perform three - dimensional reconstruction on the high - resolution image.

[0011] Among them, the image after downsampling in Step 2 is 1 / 8 of the size of the original image.

[0012] Among them, the image after downsampling in Step 3 is 1 / 4 of the size of the original image, and the size of the super - pixel block is set to 50 pixels.

[0013] Among them, Step 4 includes: Sampling the mean - shift prior depth map and the super - pixel prior depth map obtained in Step 2 and Step 3 to the scale of the original image, calculating the prior depth values for the mean - shift prior depth map, and filling the pixels without prior depth values in the mean - shift prior depth map with the corresponding prior depth values in the super - pixel prior depth map to obtain the final prior depth map, and each pixel of the prior depth map is assigned a prior depth value.

[0014] Add prior depth information in the patchmatch algorithm to calculate the prior depth values for the weak - texture regions, and combine with the cost function to constrain the process of calculating the prior depth values. Calculate the depth map of the current image under the openMVS framework as:

[0015] cost = cost ph + cost pri

[0016] Among them, costph is the photometric consistency term, cost pri is the prior constraint term;

[0017] The definition of the photometric consistency term is:

[0018] cost ph = 1 - NCC

[0019] The definition of the prior constraint term is:

[0020] where, d prior is the prior depth value of the current pixel; d can is the candidate depth value; σ3 is a constant, usually set to 0.5.

[0021] The present invention also provides a three-dimensional reconstruction system for high-resolution images, including: an image processing module, a synthesis module, and a depth calculation module;

[0022] The image processing module downsamples the high-resolution image to different scales; uses the mean-shift algorithm to cluster the downsampled image, retains the categories that account for 10% or more of the current image area in the clustering result, projects them into the three-dimensional space one by one, performs plane fitting on each category through the ransac algorithm, and back-projects the fitted plane into the corresponding category in the current image to update the depth values of the pixels in the category to obtain the mean-shift prior depth map;

[0023] Uses the superpixel segmentation algorithm to downsample and segment the image area not processed by the mean-shift algorithm, projects the superpixel blocks that account for 10% or more of the current image area in the segmentation result into the three-dimensional space one by one, and performs plane fitting on each superpixel block through the ransac algorithm, and back-projects the fitted plane into the corresponding superpixel block in the current image to update the depth values of the pixels in the superpixel block to obtain the superpixel prior depth map;

[0024] The synthesis module is used to synthesize the mean-shift prior depth map and the superpixel prior depth map obtained through the processing of the image processing module into the final prior depth map;

[0025] The depth calculation module is used to calculate the prior depth values in the final prior depth map in combination with the patchmatch algorithm to obtain the final depth map and perform three-dimensional reconstruction on the high-resolution image.

[0026] The above solution of the present invention has the following beneficial effects:

[0027] 1. By combining the mean-shift clustering algorithm and the superpixel segmentation algorithm, it can better identify the planes in the image, make full use of the plane information in the image, thereby obtaining a more accurate prior depth value, and can obtain an accurate prior depth map with a relatively large coverage area to the greatest extent.

[0028] 2. The introduction of the prior depth can well assist in depth estimation for weakly textured regions, provide prior constraints for depth estimation of large areas of weakly textured regions, greatly improve the accuracy of depth estimation in weakly textured regions. At the same time, photometric consistency is used to protect details, avoiding the problem of detail blurring and loss caused by the prior depth. For the 3D reconstruction of high-resolution images, it can well overcome the problem of false matching of photometric consistency in weakly textured regions, thereby obtaining accurate depth values and obtaining a complete 3D point cloud model. Brief Description of the Drawings

[0029] Figure 1 is a flowchart of the present invention;

[0030] Figure 2 is a flowchart of an embodiment of the present invention;

[0031] Figure 3 is a framework diagram of the 3D reconstruction system of the present invention. Detailed Embodiments

[0032] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0034] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a locking connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0035] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0036] The present invention provides a three-dimensional reconstruction method and system based on high-resolution images in view of existing problems.

[0037] As Figure 1 , 2 shown, an embodiment of the present invention provides a three-dimensional reconstruction method based on high-resolution images, including:

[0038] Since the depth estimation of high-resolution images is much more difficult than that of low-resolution images because high-resolution images have more homogeneous pixels. Therefore, the depth maps obtained from high-resolution images often have a lot of noise, especially in large areas of weak texture regions.

[0039] The present invention uses the idea of an image pyramid to downsample high-resolution images to different scales; reducing the image resolution and then calculating the depth values can avoid the estimation error problem caused by high-resolution images to a certain extent, and thus more accurate prior depth values can be obtained.

[0040] Using the mean-shift algorithm to cluster the downsampled images, the downsampled images are 1 / 8 of the original image size. Retain the categories that account for 10% or more of the current image area in the clustering results, discard the remaining clustering results, use the patchmatch algorithm to calculate the depth map of the current image, and then project the retained categories into the three-dimensional space one by one as large regions in the three-dimensional space. Use the ransac algorithm to perform plane fitting on each category so that each category obtains a fitted plane, and backproject the obtained fitted plane to the current image to assign depth values to the corresponding categories, obtaining the mean-shift prior depth map.

[0041] The mean-shift prior depth map is for large planar regions in the image, and the prior depth values are calculated at the 1 / 8 original image scale. A smaller image scale can more effectively suppress the generation of incorrect depth estimates and can obtain better prior depth value results for large planar regions.

[0042] After downsampling the image regions not processed by the mean-shift algorithm using the superpixel segmentation algorithm and then performing segmentation, the downsampled image is 1 / 4 the size of the original image, ensuring that the pixels in each superpixel block are basically on the same plane. The size of the superpixel block is set to 50 pixels. Retain the superpixel blocks that account for 10% or more of the current image area in the segmentation result, and discard the remaining superpixel blocks. Use the patchmatch algorithm to calculate the depth map of the current image, and then project the retained superpixel blocks into the three-dimensional space one by one as small regions in the three-dimensional space. And perform plane fitting on each superpixel block through the ransac algorithm, and project the fitted plane back to the current image to assign a depth value to the corresponding superpixel block, obtaining a superpixel prior depth map; it can make up for the deficiencies of the mean-shift clustering algorithm, thereby obtaining a more complete prior depth map with a larger coverage area.

[0043] Synthesize the mean-shift prior depth map and the superpixel prior depth map obtained above into the final prior depth map; first sample the mean-shift prior depth map and the superpixel prior depth map to the scale of the original image. It can be noted that not every pixel in the mean-shift prior depth map is assigned a prior depth value, while each superpixel block in the superpixel prior depth map is assigned a prior depth value.

[0044] Since the mean-shift prior depth map calculates the prior depth value for large planes in the image and believes that the result of the mean-shift prior depth map is more accurate. Also, since mean-shift only processes the categories that account for 10% or more of the current image area, for the pixels without prior depth values in the mean-shift prior depth map, fill them with the corresponding prior depth values in the superpixel prior map, thereby obtaining the final prior depth map. At this time, each pixel in the prior depth map is assigned a prior depth value.

[0045] In this embodiment, the ransac algorithm is used to perform plane fitting on it, rather than directly performing ransac clustering on the depth values. The advantage of this is that we can also obtain accurate prior depth for inclined planes. The prior depth value is not affected by whether the plane is inclined.

[0046] Combine the patchmatch algorithm at the scale of the original image to calculate the prior depth values in the final prior depth map. Under the framework of openMVS, modify the cost function of the patchmatch algorithm based on photometric consistency, and add prior constraints on the basis of photometric consistency to calculate the depth map of the current image:.

[0047] cost = costph + cost pri

[0048] where cost ph is the photometric consistency term, and cost pri is the prior constraint term.

[0049] The photometric consistency term is defined as:

[0050] cost ph = 1 - NCC

[0051] The prior constraint term is defined as:

[0052]

[0053] where d prior is the prior depth value of the current pixel. d can is the candidate depth value, and σ3 is a constant.

[0054] Based on this cost function, the patchmatch algorithm calculates the final depth map under the openMVS framework, which integrates the prior depth values calculated previously. It can effectively calculate the depth values in weak texture regions, effectively solve the problem of missing depth values caused by photometric consistency mismatches, and greatly improve the integrity of the 3D model.

[0055] For example Figure 3 as shown, the embodiment of the present invention also provides a 3D reconstruction system for high-resolution images, including: an image processing module, a synthesis module, and a depth calculation module;

[0056] The image processing module downsamples the high-resolution image using the image pyramid idea to different scales; clusters the downsampled image using the mean-shift algorithm, retains the categories that account for 10% or more of the current image area in the clustering results, projects them into the three-dimensional space one by one, and performs plane fitting on each category using the ransac algorithm. The fitted plane is back-projected into the corresponding category in the current image to update the depth values of the pixels in this category to obtain the mean-shift prior depth map;

[0057] Downsamples and segments the image regions not processed by the mean-shift algorithm using the superpixel segmentation algorithm. Projects the superpixel blocks that account for 10% or more of the current image area in the segmentation results into the three-dimensional space one by one, and performs plane fitting on each superpixel block using the ransac algorithm. The fitted plane is back-projected into the corresponding superpixel block in the current image to update the depth values of the pixels in the superpixel block to obtain the superpixel prior depth map;

[0058] A synthesis module for synthesizing the final prior depth map from the mean-shift prior depth map and the superpixel prior depth map obtained after being processed by the image processing module;

[0059] A depth calculation module for calculating the prior depth value in the final prior depth map in combination with the patchmatch algorithm to obtain the final depth map and perform 3D reconstruction on the high-resolution image.

[0060] In this embodiment, by combining the mean-shift clustering algorithm and the superpixel segmentation algorithm, planes in the image can be better recognized, and the plane information in the image can be fully utilized, so as to obtain a more accurate prior depth value and a prior depth map with a large coverage area to the greatest extent; moreover, the introduction of the prior depth can well assist in depth estimation for weak texture regions, provide prior constraints for depth estimation of large-area weak texture regions, and greatly improve the accuracy of depth estimation for weak texture regions. At the same time, photometric consistency is used to protect details and avoid the problem of detail blurring and loss caused by the prior depth. For 3D reconstruction of high-resolution images, the problem of false matching of photometric consistency in weak texture regions can be well overcome, so as to obtain accurate depth values and a complete 3D point cloud model.

[0061] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A three-dimensional reconstruction method for high-resolution images, characterized in that, Including: Step 1: Downsample the high-resolution image to different scales. Step 2: Use the mean-shift algorithm to cluster the downsampled image, retain the categories that account for 10% or more of the current image area in the clustering results, project them into the three-dimensional space one by one, perform plane fitting on each category through the ransac algorithm, and back-project the fitted plane into the corresponding category in the current image to update the depth values of the pixels in this category to obtain the mean-shift prior depth map. Step 3: Use the superpixel segmentation algorithm to segment the image area not processed by the mean-shift algorithm after downsampling, project the superpixel blocks that account for 10% or more of the current image area in the segmentation results into the three-dimensional space one by one, and perform plane fitting on each superpixel block through the ransac algorithm, and back-project the fitted plane into the corresponding superpixel block in the current image to update the depth values of the pixels in the superpixel block to obtain the superpixel prior depth map. Step 4: Synthesize the mean-shift prior depth map and the superpixel prior depth map obtained in Step 2 and Step 3 into the final prior depth map. Step 5: Combine the patchmatch algorithm to calculate the prior depth values in the final prior depth map to obtain the final depth map and perform three-dimensional reconstruction on the high-resolution image.

2. The three-dimensional reconstruction method for high-resolution images according to claim 1, wherein The scale of the image after downsampling in Step 2 is 1 / 8 of the original image size.

3. The three-dimensional reconstruction method for high-resolution images according to claim 1, wherein The scale of the image after downsampling in Step 3 is 1 / 4 of the original image size, and the size of the superpixel block is set to 50 pixels.

4. The three-dimensional reconstruction method for high-resolution images according to claim 1, wherein Step 4 includes: Sampling the mean-shift prior depth map and the superpixel prior depth map obtained in Step 2 and Step 3 to the original image scale, calculating the prior depth values for the mean-shift prior depth map, and filling the pixels without prior depth values in the mean-shift prior depth map with the corresponding prior depth values in the superpixel prior depth map to obtain the final prior depth map, and each pixel of the prior depth map is assigned a prior depth value.

5. The three-dimensional reconstruction method for high-resolution images according to claim 4, wherein Based on adding prior depth information in the patchmatch algorithm to calculate the prior depth values for the weak texture regions, and combining the cost function to constrain the prior depth value calculation process, the depth map of the current image is calculated under the openMVS framework as: cost=cost ph +cost pri Among them, cost ph is the photometric consistency term, and cost pri is the prior constraint term; The definition of the photometric consistency term is: cost ph = 1 - NCC The definition of the prior constraint term is: Among them, d prior is the prior depth value of the current pixel; d can is the candidate depth value; σ3 is a constant, usually set to 0.

5.

6. A three-dimensional reconstruction system for high-resolution images, characterized in that, Including: An image processing module, a synthesis module, and a depth calculation module; The image processing module downsamples the high-resolution image using the image pyramid idea to different scales. Use the mean-shift algorithm to cluster the downsampled image, retain the categories that account for 10% or more of the current image area in the clustering results, project them into the three-dimensional space one by one, perform plane fitting on each category through the ransac algorithm, and back-project the fitted plane into the corresponding category in the current image to update the depth values of the pixels in this category to obtain the mean-shift prior depth map. After downsampling the image regions not processed by the mean-shift algorithm using the superpixel segmentation algorithm, perform segmentation. Sequentially project the superpixel blocks that account for 10% or more of the current image area in the segmentation result into three-dimensional space, and perform plane fitting on each superpixel block through the RANSAC algorithm. The obtained fitted plane is back-projected into the corresponding superpixel blocks of the current image to update the depth values of the pixels in the superpixel blocks to obtain a superpixel prior depth map; The synthesis module is used to synthesize the mean-shift prior depth map and the superpixel prior depth map obtained after being processed by the image processing module into a final prior depth map; The depth calculation module is used to calculate the prior depth values in the final prior depth map in combination with the patchmatch algorithm to obtain a final depth map and perform three-dimensional reconstruction on the high-resolution image.

Citation Information

Patent Citations

  • Cooperative significance detection method based on superpixel clustering

    CN107103326A

  • Method for compact 3D scene reconstruction based on semantic priori and progressive optimization of wide baseline

    CN109255833A