A three-dimensional reconstruction method, device, equipment and storage medium

By extracting features and grouping and matching multiple images of the target object, sparse point clouds are generated and dense matching is performed, which solves the problems of low efficiency and low accuracy in 3D reconstruction and realizes efficient and high-precision 3D model reconstruction in general scenes.

CN115546401BActive Publication Date: 2026-07-31WUHAN DASHI SMART TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN DASHI SMART TECH CO LTD
Filing Date
2022-09-26
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies for 3D reconstruction are inefficient and lack precision. Ordinary users cannot reconstruct a complete 3D model in a typical scenario by flipping a small object and taking a picture.

Method used

By extracting features from multiple images of the target object, combining image pairs, determining the target image grouping method, performing foreground feature matching and aerial triangulation, sparse point clouds are obtained, and a 3D model is generated through dense matching.

Benefits of technology

It improves the accuracy and efficiency of 3D reconstruction, avoids reconstruction failure by grouping images, clearly distinguishes the foreground and background, and enhances the accuracy and efficiency of reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115546401B_ABST
    Figure CN115546401B_ABST
Patent Text Reader

Abstract

This application provides a 3D reconstruction method, apparatus, device, and storage medium, relating to the field of image processing technology. The method involves extracting features from multiple images of a target object to obtain feature points for each image; generating image pairs; determining a target image grouping method from multiple image grouping methods based on the feature points of the image pairs; if two images in an image pair are located in different image groups corresponding to the target image grouping method, matching feature points of a preset foreground region in the image pair to obtain a foreground feature matching pair; performing aerial triangulation based on the foreground feature matching pair to obtain a sparse point cloud of the target object; performing dense matching on the sparse point cloud of the target object corresponding to the image grouping method to obtain a dense point cloud of the target object; and performing 3D modeling based on the dense point cloud to obtain a 3D model of the target object. This improves the accuracy and efficiency of 3D reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically, to a three-dimensional reconstruction method, apparatus, device, and storage medium. Background Technology

[0002] Photogrammetry is a technique that uses images of a subject to reconstruct its spatial location and three-dimensional shape. Photogrammetric-based 3D reconstruction technology is widely used in Virtual Reality (VR), Industrial Surveying, and Cultural Heritage Protection. Typically, to obtain a complete 3D model of a small object from an image, it needs to be placed in a scene with negligible background, such as a black-and-white turntable or green screen, then rotated and photographed. This setup requires specialized knowledge and equipment, making it inaccessible to the average user. For them, the most direct method is to place the small object in a common scene, such as a tabletop or flat outdoor ground, then rotate it, photograph it, and reconstruct the model. However, this operation violates the assumption of "scene staticity" in traditional 3D reconstruction techniques. Therefore, currently, ordinary users cannot use photogrammetry to reconstruct a complete 3D model of a small object by rotating it in a general scene and photographing it.

[0003] For reconstruction cases that do not meet the assumption of "scene remaining static," most existing solutions are based on videos containing moving objects. Techniques such as Multi-body SfM applied to dynamic three-dimensional reconstruction largely focus on issues such as moving object recognition, segmentation, and removal. However, video-based dynamic scene methods have a fixed shooting perspective, and video data acquisition is relatively difficult, reducing the efficiency and accuracy of 3D reconstruction. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of the prior art by providing a three-dimensional reconstruction method, apparatus, device, and storage medium to solve the problems of low efficiency and low accuracy in the prior art.

[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:

[0006] In a first aspect, embodiments of this application provide a three-dimensional reconstruction method, the method comprising:

[0007] Feature extraction is performed on multiple images of the target object to obtain feature points of the multiple images;

[0008] The multiple images are combined in pairs to obtain image pairs;

[0009] Based on the feature points of the image pairs, the target image grouping method is determined from multiple image grouping methods;

[0010] If the two images in the image pair are respectively in different image groups corresponding to the target image grouping method, then the feature points of the preset foreground region in the image pair are matched to obtain the foreground feature matching pair of the image pair;

[0011] Aerial triangulation is performed based on the foreground feature matching pairs to obtain the sparse point cloud of the target object corresponding to the image pair;

[0012] Based on the image grouping results corresponding to the target image grouping method, dense matching is performed on the sparse point cloud of the target object corresponding to the image pair to obtain the dense point cloud of the target object.

[0013] A 3D model of the target object is obtained by performing 3D modeling based on the dense point cloud.

[0014] Optionally, determining the target image grouping method from multiple image grouping methods based on the feature points of the image pair includes:

[0015] Based on the feature points of the image pair, feature matching is performed on the feature points of the two images in the image pair to obtain multiple sets of feature matching pairs of the image pair.

[0016] The target image grouping method is determined from the plurality of image grouping methods based on multiple feature matching pairs of the image pairs.

[0017] Optionally, determining the target image grouping method from the plurality of image grouping methods based on multiple feature matching pairs of the image pair includes:

[0018] The image pair is thinned out by multiple feature matching pairs to obtain the target feature matching pair of the image pair, such that the ratio of the number of feature points in the preset foreground region to the number of feature points in the preset background region in the target feature matching pair is at a preset ratio.

[0019] Based on the target feature matching pairs of the image pairs, the target image grouping method is determined from the plurality of image grouping methods.

[0020] Optionally, before performing thinning processing on multiple sets of feature matching pairs of the image pair to obtain the target feature matching pair of the image pair, the method further includes:

[0021] Based on multiple feature matching pairs of each image pair, determine the feature matching status of the preset foreground region and the preset background region;

[0022] Based on the feature matching results, image pairs that meet the preset feature matching conditions are determined as target image pairs from each image pair;

[0023] The step of thinning multiple feature matching pairs of the image pair to obtain the target feature matching pair of the image pair includes:

[0024] The target image pair is subjected to a thinning process on multiple feature matching pairs to obtain the target feature matching pair of the target image pair.

[0025] Optionally, determining the target image grouping method from the plurality of image grouping methods based on the target feature matching pairs of the image pairs includes:

[0026] Calculate the edge weights of each target image pair based on the target feature matching pairs of the target image pairs;

[0027] Based on the edge weights of each target image pair, calculate the grouping cost function value corresponding to the multiple image grouping methods;

[0028] Based on the grouping cost function values ​​corresponding to the multiple image grouping methods, the image grouping method with the smallest grouping cost function value is selected as the target image grouping method.

[0029] Optionally, the step of performing dense matching on the sparse point cloud of the target object corresponding to the image pair based on the image grouping result corresponding to the target image grouping method to obtain the dense point cloud of the target object includes:

[0030] Based on the image grouping results corresponding to the target image grouping method, select multiple images from the same group and images from different groups for each image to form stereo image pairs from the same group and stereo image pairs from different groups for each image.

[0031] Dense matching is performed on the sparse point clouds of the same group of stereo image pairs and the sparse point clouds of different groups of stereo image pairs.

[0032] The dense point cloud of the target object is obtained by combining the same set of stereo image pairs and the different sets of stereo image pairs after dense matching.

[0033] Optionally, the step of performing dense matching on the sparse point clouds of the same set of stereo image pairs and the sparse point clouds of different sets of stereo image pairs includes:

[0034] The depth values ​​of corresponding points in the two images of the sparse point cloud of the same stereo image pair are matched and verified to obtain the same stereo image pair after dense matching.

[0035] The depth values ​​of corresponding points in two images of the sparse point cloud of the different sets of stereo image pairs are matched and verified to obtain the different sets of stereo image pairs after dense matching.

[0036] Secondly, embodiments of this application provide a three-dimensional reconstruction apparatus, the apparatus comprising:

[0037] The extraction module is used to extract features from multiple images of a target object to obtain feature points from the multiple images.

[0038] The combination module is used to combine the multiple images in pairs to obtain image pairs;

[0039] The determination module is used to determine the target image grouping method from multiple image grouping methods based on the feature points of the image pair;

[0040] The matching module is used to match feature points of a preset foreground region in the image pair if the two images in the image pair are respectively in different image groups corresponding to the target image grouping method, so as to obtain a foreground feature matching pair of the image pair.

[0041] The measurement module is used to perform aerial triangulation based on the foreground feature matching pair to obtain the sparse point cloud of the target object corresponding to the image pair;

[0042] The dense matching module is used to perform dense matching on the sparse point cloud of the target object corresponding to the image pair according to the image grouping result corresponding to the target image grouping method, so as to obtain the dense point cloud of the target object.

[0043] The modeling module is used to perform three-dimensional modeling based on the dense point cloud to obtain a three-dimensional model of the target object.

[0044] Thirdly, embodiments of this application provide an electronic device, including: a processor and a storage medium, wherein the processor and the storage medium are connected via a bus for communication, and the storage medium stores program instructions executable by the processor, wherein the processor calls the program stored in the storage medium to perform the steps of the three-dimensional reconstruction method as described in any of the first aspects.

[0045] Fourthly, embodiments of this application provide a storage medium storing a computer program, which, when executed by a processor, performs the steps of the three-dimensional reconstruction method as described in any of the first aspects.

[0046] Compared with the prior art, this application has the following beneficial effects:

[0047] This application provides a 3D reconstruction method, apparatus, device, and storage medium. The method involves extracting features from multiple images of a target object to obtain feature points; combining the multiple images in pairs to obtain image pairs; determining a target image grouping method from multiple image grouping methods based on the feature points of the image pairs; if two images in an image pair belong to different image groups corresponding to the target image grouping method, matching feature points of a preset foreground region in the image pair to obtain a foreground feature matching pair; performing aerial triangulation based on the foreground feature matching pair to obtain a sparse point cloud of the target object corresponding to the image pair; performing dense matching on the sparse point cloud of the target object corresponding to the image pair based on the image grouping result corresponding to the target image grouping method to obtain a dense point cloud of the target object; and performing 3D modeling based on the dense point cloud to obtain a 3D model of the target object. Thus, by grouping images, the method avoids reconstruction failure caused by directly processing all images, improving the accuracy of 3D reconstruction; by judging the foreground and clarifying background noise, the method achieves better foreground and background differentiation; and by integrating foreground and background differentiation into the 3D reconstruction process, the method improves the efficiency of 3D reconstruction. Attached Figure Description

[0048] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 A flowchart illustrating a three-dimensional reconstruction method provided in this application;

[0050] Figure 2 A flowchart illustrating a method for determining the grouping method of target images provided in an embodiment of this application;

[0051] Figure 3 A flowchart illustrating another method for determining the grouping of target images provided in an embodiment of this application;

[0052] Figure 4 A flowchart illustrating a method for filtering image pairs provided in an embodiment of this application;

[0053] Figure 5 A flowchart illustrating another method for determining the grouping method of target images provided in an embodiment of this application;

[0054] Figure 6 A flowchart illustrating a method for determining a dense point cloud of a target object, provided in an embodiment of this application;

[0055] Figure 7 A flowchart illustrating a method for dense matching of sparse point clouds provided in an embodiment of this application;

[0056] Figure 8 A schematic diagram of a three-dimensional reconstruction device provided in an embodiment of this application;

[0057] Figure 9 This is a schematic diagram of an electronic device provided in an embodiment of this application.

[0058] Icons: 801-Extraction Module, 802-Combination Module, 803-Determining Module, 804-Matching Module, 805-Measurement Module, 806-Dense Matching Module, 807-Modeling Module, 901-Processor, 902-Storage Medium. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. The components of the embodiments of the present application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0060] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0061] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0062] Furthermore, the terms "first" and "second" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.

[0063] It should be noted that, where there is no conflict, the features in the embodiments of the present invention can be combined with each other.

[0064] To achieve accurate and efficient 3D reconstruction, this application provides a 3D reconstruction method, apparatus, device, and storage medium.

[0065] The following specific examples illustrate a three-dimensional reconstruction method provided in this application. Figure 1 This is a flowchart illustrating a three-dimensional reconstruction method provided in this application. The method is executed by an electronic device, which can be a device with computing processing capabilities, such as a desktop computer or laptop computer. Figure 1 As shown, the method includes:

[0066] S101. Perform feature extraction on multiple images of the target object to obtain feature points of multiple images.

[0067] To achieve 3D reconstruction of the target object, the small object needs to be intermittently flipped and photographed multiple times from both sides until multiple images are obtained that cover all shooting angles of the target object. Each image contains a foreground area and a background area; the foreground area is the image portion of the target object, and the background area is the image portion other than the target object.

[0068] After acquiring multiple images of the target object, these images need to be preprocessed to facilitate further processing. Feature extraction is performed on each of the multiple images to obtain feature points. These feature points are the characteristic pixels in the images. For example, feature extraction methods include, but are not limited to, Scale-Invariant Feature Transform (SIFT), which is applied to extract features from each of the multiple images of the target object.

[0069] S102. Combine multiple images in pairs to obtain image pairs.

[0070] To facilitate the complete display of the features of the target object, this embodiment uses image pairs for 3D reconstruction. An exhaustive combination method is used to combine multiple images in pairs to obtain multiple image pairs. That is, an image pair includes two images.

[0071] S103. Determine the target image grouping method from multiple image grouping methods based on the feature points of the image pair.

[0072] An image pair consists of two images, each containing multiple feature points. An image pair also contains multiple feature points.

[0073] Multiple images of the target object are divided into several groups, resulting in multiple grouping methods. Based on the feature points of the image pairs, the optimal grouping method is determined from these methods and used as the target image grouping method.

[0074] After determining the target image grouping method, multiple images are grouped based on the target image grouping method.

[0075] S104. If the two image pairs in the image pair are respectively in different image groups corresponding to the target image grouping method, then the feature points of the preset foreground region in the image pair are matched to obtain the foreground feature matching pair of the image pair.

[0076] After grouping multiple images based on the target image grouping method, if two image pairs within an image pair are located in different image groups, it indicates that the foreground of that image pair has been flipped. Furthermore, these two image pairs are respectively a foreground matching pair and a background matching pair.

[0077] Then, combining the prior knowledge of the foreground and background positions in the image (foreground is mostly located in the image center; background is mostly located in the image edge), an initial type determination is made for the image pairs. If the matching pair set of the central region of one of the two image pairs is located in the image center, then the matching pair set of the central region is determined to be a foreground matching pair, and the matching pair set of the central region of the other image pair is determined to be a background matching pair. If the matching pair set of the central region of one of the two image pairs is located in the image edge, then the matching pair set of the central region is determined to be a background matching pair, and the matching pair set of the central region of the other image pair is determined to be a foreground matching pair. If the matching pair set of the central region of both image pairs cannot be determined, or in other cases, then it is determined to be an unknown matching pair.

[0078] For example, using the center point of the image as the base point, a region that is similar in shape to the image and occupies a preset proportion of the image area is defined as the image center, and other regions are defined as the image edges.

[0079] After obtaining the initial type of the matching pair set in the central region of the image pair, the matching pair set of the foreground matching pair can be further calculated and determined according to the voting algorithm. The specific calculation method is shown in the following formula (1):

[0080]

[0081] Wherein, the initial score of the feature matching pair {m,n} The initial type of the matching pair is determined by the initial type of the pair: if {m,n} come from different image pairs, their initial type has been preliminarily determined by the above steps. Scores for foreground matching pairs, background matching pairs, and unknown matching pairs. The values ​​are 1, -1, and 0, respectively; for other cases, feature matching pairs {m, n} are all considered unknown matching pairs. S is the set of matching pairs in the central region of all image pairs. t The voting score for the feature matching pair {m,n}.

[0082] Ultimately, those with a high score (S) t All feature matching pairs in the set >δ) are identified as foreground feature matching pairs of the image pair. Here, δ can be 0.7.

[0083] S105. Perform aerial triangulation based on foreground feature matching to obtain sparse point clouds of the target objects corresponding to the image pairs.

[0084] The position coordinates of the feature pixels in the foreground feature matching pair are determined, and aerial triangulation is performed based on these coordinates. Specifically, based on the position coordinates of the control points and their positional relationships with all feature pixels, the image parameters of all feature pixels are solved, and a point cloud of all feature pixels in 3D space is constructed, resulting in a sparse point cloud of the target object corresponding to the image pair. For example, the solution method includes, but is not limited to, bundle adjustment.

[0085] S106. Based on the image grouping results corresponding to the target image grouping method, perform dense matching on the sparse point cloud of the target object corresponding to the image pair to obtain the dense point cloud of the target object.

[0086] After grouping multiple images based on the target image grouping method, and according to the image grouping results, multiple dense matching operations are performed on the sparse point clouds of the corresponding target objects in each image. Then, background noise removal processing is applied to the densely matched point clouds. The point cloud after background noise removal is the dense point cloud of the target object.

[0087] S107. Perform 3D modeling based on dense point cloud to obtain the 3D model of the target object.

[0088] The dense point cloud can already display the outline of the target object. Then, 3D modeling is performed based on the dense point cloud to obtain the 3D model of the target object, completing the 3D reconstruction. For example, 3D modeling methods include, but are not limited to, constructing a 3D spatial surface mesh and texture mapping.

[0089] In this way, by grouping images, the reconstruction failure caused by directly processing all images is avoided, thus improving the accuracy of 3D reconstruction; by judging the foreground and clarifying the background noise, the foreground and background are better distinguished; at the same time, the foreground and background distinction is integrated into the 3D reconstruction process, thus improving the efficiency of 3D reconstruction.

[0090] In summary, in this embodiment, feature points of multiple images of the target object are obtained by extracting features from each image; image pairs are obtained by combining multiple images in pairs; based on the feature points of the image pairs, the target image grouping method is determined from multiple image grouping methods; if two images in an image pair are in different image groups corresponding to the target image grouping method, feature points of a preset foreground region in the image pair are matched to obtain a foreground feature matching pair; aerial triangulation is performed based on the foreground feature matching pair to obtain a sparse point cloud of the target object corresponding to the image pair; based on the image grouping results corresponding to the target image grouping method, dense matching is performed on the sparse point cloud of the target object corresponding to the image pair to obtain a dense point cloud of the target object; and 3D modeling is performed based on the dense point cloud to obtain a 3D model of the target object. Thus, by grouping images, the reconstruction failure caused by directly processing all images is avoided, improving the accuracy of 3D reconstruction; foreground judgment and background noise removal result in better foreground and background differentiation; and the integration of foreground and background differentiation into the 3D reconstruction process improves the efficiency of 3D reconstruction.

[0091] In the above Figure 1 Based on the corresponding embodiments, this application also provides a method for determining the grouping method of target images. Figure 2 This is a flowchart illustrating a method for determining the grouping method of target images, provided in an embodiment of this application. Figure 2 As shown, determining the target image grouping method from multiple image grouping methods in S103 includes:

[0092] S201. Based on the feature points of the image pair, perform feature matching on the feature points of the two images in the image pair to obtain multiple sets of feature matching pairs of the image pair.

[0093] When performing feature matching on image pairs, the matching methods include, but are not limited to, K-nearest neighbor (KNN).

[0094] In addition, when performing feature matching on image pairs, the following steps are also taken: removing gross errors in feature point matching, performing grid-based motion statistics (GMS) to pair the neighboring grids of the central grid, changing the globally consistent pairing mode to a locally adaptive pairing mode, and removing erroneous feature point matching. In object-flipped image pairs, there are still valid feature matches between the objects and the background in the image pairs.

[0095] S202. Based on multiple feature matching pairs of image pairs, determine the target image grouping method from multiple image grouping methods.

[0096] Multiple images of the target object are divided into several groups, resulting in multiple grouping methods. Based on multiple feature matching pairs of the image pairs, the optimal image grouping method is determined as the target image grouping method.

[0097] After determining the target image grouping method, multiple images are grouped based on the target image grouping method.

[0098] In summary, in this embodiment, feature matching is performed on the feature points of the two images in the image pair to obtain multiple sets of feature matching pairs. Based on these multiple sets of feature matching pairs, the target image grouping method is determined from multiple image grouping methods. Therefore, by performing feature matching on the feature points of the two images in the image pair, more accurate image pairs can be obtained.

[0099] In the above Figure 2 Based on the corresponding embodiments, this application also provides another method for determining the target image grouping method. Figure 3 This is a flowchart illustrating another method for determining the grouping method of target images provided in an embodiment of this application. Figure 3 As shown, in S202, determining the target image grouping method from multiple image grouping methods based on multiple feature matching pairs of image pairs includes:

[0100] S301. Perform thinning processing on multiple feature matching pairs of the image pair to obtain the target feature matching pair of the image pair, so that the ratio of the number of feature points in the preset foreground region to the number of feature points in the preset background region in the target feature matching pair is at a preset ratio.

[0101] Multiple feature matching pairs undergoing feature matching and gross error removal may result in a significant difference in the number of feature points between the preset foreground region and the preset background region, reducing the accuracy of 3D reconstruction. Therefore, thinning is performed on multiple feature matching pairs of the image pair to obtain the target feature matching pair, ensuring that the ratio of feature points in the preset foreground region to the preset background region in the target feature matching pair is within a preset ratio (e.g., between 1:2 and 2:1). For example, thinning methods include, but are not limited to, non-maximum suppression (NMS).

[0102] The specific preset foreground and preset background regions can be determined based on the prior positions in the image (the foreground is mostly located in the center of the image; the background is mostly located at the edge of the image), similar to the determination method in the above embodiments, and will not be repeated here.

[0103] S302. Based on the target feature matching pairs of the image pairs, determine the target image grouping method from multiple image grouping methods.

[0104] Multiple images of the target object are divided into several groups, resulting in multiple grouping methods. Based on the target feature matching pairs of the image pairs, the optimal image grouping method is determined from the multiple image grouping methods as the target image grouping method.

[0105] After determining the target image grouping method, multiple images are grouped based on the target image grouping method.

[0106] In summary, in this embodiment, by thinning multiple feature matching pairs of an image pair, a target feature matching pair of the image pair is obtained, such that the ratio of feature points in the preset foreground region to feature points in the preset background region of the target feature matching pair is at a preset ratio. Based on the target feature matching pair of the image pair, a target image grouping method is determined from multiple image grouping methods. Thus, the feature matching pair is thinned to improve the accuracy of 3D reconstruction.

[0107] In the above Figure 3 Based on the corresponding embodiments, this application also provides a method for screening image pairs. Figure 4 This is a flowchart illustrating a method for filtering image pairs provided in an embodiment of this application. Figure 4 As shown, before performing thinning processing on multiple feature matching pairs of the image pair in S301 to obtain the target feature matching pair of the image pair, the method further includes:

[0108] S401. Based on multiple feature matching pairs of each image pair, determine the feature matching status of the preset foreground region and the preset background region.

[0109] In multiple feature matching pairs for each image pair, the two feature points in each feature matching pair may have various distributions, such as: one feature point in the foreground region and one feature point in the background region, both feature points in the foreground region, and both feature points in the background region.

[0110] The process of 3D reconstruction involves distinguishing the foreground and background regions. Therefore, based on multiple feature matching pairs for each image pair, the feature matching status of the preset foreground and background regions is determined. Specifically, this involves obtaining the number of feature matching pairs in the preset foreground and background regions, or their proportion within the multiple feature matching pairs of the image pair.

[0111] The specific division of the preset foreground area and preset background area is similar to the determination method in the above embodiments, and will not be repeated here.

[0112] S402. Based on the feature matching results, determine the image pairs that meet the preset feature matching conditions from each image pair as the target image pair.

[0113] To more accurately distinguish between foreground and background regions, image pairs with a particularly large number of feature matching pairs in both the preset foreground and preset background regions can be selected as target image pairs, facilitating foreground and background differentiation. Specifically, image pairs that meet preset feature matching conditions are selected as target image pairs from among all image pairs.

[0114] For example, the preset feature matching condition can be that the number of feature matching pairs is greater than or equal to a preset number, or the preset feature matching condition can be that the proportion of feature matching pairs in multiple feature matching pairs of an image pair is greater than or equal to a preset proportion.

[0115] For example, methods for measuring the richness of feature matching pairs in a preset foreground region and a preset background region include, but are not limited to, the number of matches, the distribution pattern of matches on the image, etc.

[0116] S301 performs thinning processing on multiple feature matching pairs of the image pair to obtain the target feature matching pairs of the image pair, including:

[0117] S403. Perform thinning processing on multiple feature matching pairs of the target image pair to obtain the target feature matching pair of the target image pair.

[0118] The filtered target images are easier to accurately reconstruct in 3D. Therefore, the feature matching pairs of the target image pairs are further thinned to obtain the target feature matching pairs of the target image pairs. The specific thinning method is similar to the above embodiment and will not be repeated here.

[0119] In summary, in this embodiment, the feature matching status of the preset foreground region and the preset background region is determined based on multiple feature matching pairs of each image pair; based on the feature matching status, image pairs that meet the preset feature matching conditions are selected as target image pairs; and the multiple feature matching pairs of the target image pairs are thinned to obtain the target feature matching pairs of the target image pairs. In conclusion, by filtering image pairs, the 3D reconstruction becomes more accurate.

[0120] In the above Figure 4 Based on the corresponding embodiments, this application also provides another method for determining the grouping method of target images. Figure 5 This is a flowchart illustrating another method for determining the grouping method of target images provided in an embodiment of this application. Figure 5 As shown, in S302, the target image grouping method is determined from multiple image grouping methods based on the target feature matching pairs of image pairs, including:

[0121] S501. Calculate the edge weights of each target image pair based on the target feature matching pairs of the target image pairs.

[0122] Before calculating the edge weights of the target image pairs, clustering is performed on them. Feature matching pairs within the image pairs are randomly sampled multiple times and quaternion calculations are performed, resulting in multiple quaternion sets (denoted as Q). Each feature matching pair's quaternion includes the 3D coordinates of the feature point and the angle value from which that feature point is rotated to another feature point in the matching pair. Cluster analysis is then performed on each quaternion set to obtain cluster centers. The clustering methods include, but are not limited to, the Clustering by Fast Search and Find of DensityPeak (CFSFDP) algorithm.

[0123] The calculation of edge weights is related to the clustering results of the quaternion set. For a target image pair {i,j}, its corresponding quaternion set is denoted as Q. ij Let w be the weight of the edge. ij The specific calculation formula is shown in formula (2) below:

[0124]

[0125] Among them, w ij For edge weights, Q represents ij The maximum distance between any two cluster centers in a plurality of cluster centers; σ i σ j Used to adjust the calculation of edge weights. σ i For all d of multiple image pairs containing image i max The Kth non-zero minimum value, σ j It is the Kth non-zero minimum value among all dmax values ​​of multiple image pairs containing image j, where K is a preset value and is not limited here.

[0126] S502. Calculate the grouping cost function value corresponding to multiple image grouping methods based on the edge weights of each target image pair.

[0127] The specific cost function is calculated as shown in formula (3) below:

[0128]

[0129] Where G1, G2, ..., G s Indicates the 1st, 2nd, ..., sth image groups; This represents all images except the m-th image group. Multiple image grouping methods result in s image groups G1, G2, ..., G... s Since they are different, the grouping cost function value corresponding to multiple image grouping methods can be calculated based on the multiple image grouping methods.

[0130] S503. Based on the grouping cost function values ​​corresponding to multiple image grouping methods, select the image grouping method with the smallest grouping cost function value as the target image grouping method.

[0131] The cost function is introduced to select the optimal target image grouping method. The smaller the grouping cost function value, the better the image grouping method. Therefore, based on the grouping cost function values ​​corresponding to multiple image grouping methods, the image grouping method with the smallest grouping cost function value is selected as the target image grouping method.

[0132] After determining the target image grouping method, multiple images are grouped based on the target image grouping method.

[0133] In summary, in this embodiment, the edge weights of each target image pair are calculated based on the target feature matching pairs; the grouping cost function values ​​corresponding to multiple image grouping methods are calculated based on the edge weights of each target image pair; and the image grouping method with the smallest grouping cost function value is selected as the target image grouping method based on the grouping cost function values ​​corresponding to multiple image grouping methods. Thus, by introducing edge weights and grouping cost function values, the optimal target image grouping method is obtained.

[0134] In addition, in the above Figure 5 Based on the corresponding embodiment, the set of matching pairs in the central region of the image pair in S104 above can be determined by the cluster center of the image pair, which is the neighborhood of the cluster center. For example, the cluster centers of the two image pairs are c p c q Cluster center c p c q The set of interior point matching pairs corresponding to the passage is M. p M q Therefore, it can be determined that M p M q The initial type.

[0135] In the above Figure 1 Based on the corresponding embodiments, this application also provides a method for determining dense point clouds of a target object. Figure 6 This is a flowchart illustrating a method for determining a dense point cloud of a target object, provided in an embodiment of this application. Figure 6 As shown, in S106, based on the image grouping results corresponding to the target image grouping method, dense matching is performed on the sparse point cloud of the target object corresponding to the image pair to obtain the dense point cloud of the target object, including:

[0136] S601. Based on the image grouping results corresponding to the target image grouping method, select multiple images from the same group and different groups for each image to form stereo image pairs from the same group and different groups for each image.

[0137] In this case, the image is taken as the left image, and multiple selected images from the same group and different groups are taken as the right images, that is, one left image corresponds to multiple right images.

[0138] To facilitate 3D reconstruction, the following conditions must be met when constructing stereo image pairs: a large number of shared feature points in the left and right images, similar scales between the left and right images, and a moderate baseline length formed by the left and right images. Specifically, preset conditions can be set: the number of shared feature points in the left and right images is greater than or equal to a preset number; the similarity between the left and right images is greater than or equal to a preset similarity; and the baseline length formed by the left and right images is within a preset length range.

[0139] S602. Perform dense matching on the sparse point clouds of the same set of stereo image pairs and the sparse point clouds of different sets of stereo image pairs respectively.

[0140] After obtaining stereo image pairs within the same group and stereo image pairs outside the same group for each image, multiple dense matching operations are performed on the sparse point clouds of the same group of stereo image pairs and the sparse point clouds of different groups of stereo image pairs. Background noise is then removed from the densely matched point clouds to complete the dense matching process.

[0141] S603. Obtain the dense point cloud of the target object based on the same group of stereo image pairs and different groups of stereo image pairs after dense matching.

[0142] By combining the point clouds corresponding to the same set of densely matched stereo image pairs with the electric clouds corresponding to different sets of stereo image pairs, a dense point cloud of the target object is obtained.

[0143] In summary, in this embodiment, by selecting multiple images from the same group and different groups for each image according to the image grouping results corresponding to the target image grouping method, stereo image pairs from the same group and different groups are formed for each image. Dense matching is then performed on the sparse point clouds of the stereo image pairs from the same group and the sparse point clouds of the stereo image pairs from different groups. Based on the densely matched stereo image pairs from the same group and the stereo image pairs from different groups, a dense point cloud of the target object is obtained. This results in richer point cloud data and improved accuracy of 3D reconstruction.

[0144] In the above Figure 6 Based on the corresponding embodiments, this application also provides a method for dense matching of sparse point clouds. Figure 7 This is a flowchart illustrating a method for dense matching of sparse point clouds provided in an embodiment of this application. Figure 7As shown, S602 performs dense matching on sparse point clouds in the same set of stereo image pairs and on sparse point clouds in different sets of stereo image pairs, including:

[0145] S701. Match and verify the depth values ​​of corresponding points in the two images of the sparse point cloud of the same stereo image pair to obtain the same stereo image pair after dense matching.

[0146] The depth values ​​of feature points in the image can be obtained from the sparse point cloud of the same stereo image pair.

[0147] If the depth values ​​of corresponding points in two images meet the left-right consistency matching check, they are considered a valid match; otherwise, they are considered an invalid match. The specific left-right consistency matching check formula is shown in formula (4) below:

[0148]

[0149] in, This represents the depth value of the corresponding point in the left image. S is the depth value of the corresponding point in the right image. depth This is the depth verification value. Specifically, the coordinates (pt) of the object's 3D point are recovered from the depth value of the corresponding point in the left image and the parameters of the left image. n Based on the right image parameters, pt n After projecting onto the right image, the equivalent depth in the formula is obtained.

[0150] Preset verification values ​​can be set. If the depth check value is greater than or equal to the preset check value, then the depth value of the corresponding point is determined to meet the left-right consistency matching check.

[0151] Furthermore, the number of valid matches formed by each pixel in the left image across all the right images is n. Based on n, it can be determined whether the depth value of that pixel in the left image is optimized or discarded. For example, a preset optimization ratio, such as 0.3, is set. When n < 0.3N (where N represents the total number of stereo pairs formed with the left image), the depth value is discarded.

[0152] Then, the depth values ​​of corresponding points in the two images of the sparse point cloud of the same stereo image pair are matched and verified to obtain the same stereo image pair after dense matching.

[0153] S702. Match and verify the depth values ​​of corresponding points in two images in the sparse point cloud of different stereo image pairs to obtain different stereo image pairs after dense matching.

[0154] The method for matching and verifying the depth values ​​of corresponding points in two images of sparse point clouds of different stereo image pairs is similar to the method for matching and verifying the depth values ​​of corresponding points in two images of sparse point clouds of the same stereo image pair, and will not be repeated here.

[0155] In summary, in this embodiment, by matching and verifying the depth values ​​of corresponding points in two images within the sparse point cloud of the same stereo image pair, densely matched stereo image pairs are obtained. Similarly, by matching and verifying the depth values ​​of corresponding points in two images within the sparse point cloud of different stereo image pairs, densely matched stereo image pairs are obtained. Thus, depth verification makes the 3D reconstruction more accurate.

[0156] The following describes the three-dimensional reconstruction apparatus, equipment, and storage medium provided in this application for execution. The specific implementation process and technical effects are described above and will not be repeated below.

[0157] Figure 8 This is a schematic diagram of a three-dimensional reconstruction device provided in an embodiment of this application. Figure 8 As shown, the device includes:

[0158] The extraction module 801 is used to extract features from multiple images of the target object to obtain feature points of the multiple images.

[0159] The combination module 802 is used to combine multiple images in pairs to obtain image pairs.

[0160] The determination module 803 is used to determine the target image grouping method from multiple image grouping methods based on the feature points of the image pair.

[0161] The matching module 804 is used to match feature points of a preset foreground region in the image pair if the two images in the image pair are in different image groups corresponding to the target image grouping method, so as to obtain a foreground feature matching pair of the image pair.

[0162] The measurement module 805 is used to perform aerial triangulation based on foreground feature matching pairs to obtain sparse point clouds of the target objects corresponding to the image pairs.

[0163] The dense matching module 806 is used to perform dense matching on the sparse point cloud of the target object corresponding to the image grouping result corresponding to the target image grouping method, so as to obtain the dense point cloud of the target object.

[0164] Modeling module 807 is used to perform 3D modeling based on dense point clouds to obtain a 3D model of the target object.

[0165] Furthermore, the determining module 803 is specifically used to perform feature matching on the feature points of the two images in the image pair based on the feature points of the image pair, to obtain multiple sets of feature matching pairs of the image pair; and to determine the target image grouping method from multiple image grouping methods based on the multiple sets of feature matching pairs of the image pair.

[0166] Furthermore, the determining module 803 is specifically used to perform thinning processing on multiple sets of feature matching pairs of the image pair to obtain the target feature matching pair of the image pair, so that the ratio of the number of feature points in the preset foreground region to the number of feature points in the preset background region in the target feature matching pair is at a preset ratio; and to determine the target image grouping method from multiple image grouping methods based on the target feature matching pair of the image pair.

[0167] Furthermore, the determining module 803 is specifically used to determine the feature matching situation in the preset foreground region and the preset background region based on the multiple feature matching pairs of each image pair; to determine the image pair that meets the preset feature matching conditions from each image pair as the target image pair based on the feature matching situation; and to perform thinning processing on the multiple feature matching pairs of the image pair to obtain the target feature matching pair of the image pair, including: performing thinning processing on the multiple feature matching pairs of the target image pair to obtain the target feature matching pair of the target image pair.

[0168] Furthermore, the determining module 803 is specifically used to calculate the edge weights of each target image pair based on the target feature matching pairs of the target image pairs; calculate the grouping cost function values ​​corresponding to multiple image grouping methods based on the edge weights of each target image pair; and select the image grouping method with the smallest grouping cost function value as the target image grouping method based on the grouping cost function values ​​corresponding to multiple image grouping methods.

[0169] Furthermore, the dense matching module 806 is specifically used to select multiple images from the same group and different groups for each image according to the image grouping results corresponding to the target image grouping method, forming stereo image pairs from the same group and stereo image pairs from different groups for each image; to perform dense matching on the sparse point clouds of the stereo image pairs from the same group and the sparse point clouds of the stereo image pairs from different groups respectively; and to obtain the dense point cloud of the target object based on the densely matched stereo image pairs from the same group and the stereo image pairs from different groups.

[0170] Furthermore, the dense matching module 806 is specifically used to perform matching and verification of the depth values ​​of corresponding points in two images in the sparse point cloud of the same group of stereo image pairs to obtain densely matched stereo image pairs in the same group; and to perform matching and verification of the depth values ​​of corresponding points in two images in the sparse point cloud of different groups of stereo image pairs to obtain densely matched stereo image pairs in different groups.

[0171] Figure 9This is a schematic diagram of an electronic device provided in an embodiment of this application. The electronic device may be a device with computing processing capabilities.

[0172] The electronic device includes a processor 901 and a storage medium 902. The processor 901 and the storage medium 902 are connected via a bus.

[0173] Storage medium 902 is used to store programs, and processor 901 calls the programs stored in storage medium 902 to execute the above method embodiments. The specific implementation and technical effects are similar, and will not be described in detail here.

[0174] Optionally, the present invention also provides a storage medium including a program, which, when executed by a processor, is used to perform the above-described method embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0175] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0176] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0177] The integrated units implemented as software functional units described above can be stored in a storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A three-dimensional reconstruction method, characterized by, The method includes: Feature extraction is performed on multiple images of the target object to obtain feature points of the multiple images; The multiple images are combined in pairs to obtain image pairs; Based on the feature points of the image pairs, the target image grouping method is determined from multiple image grouping methods; If the two images in the image pair are respectively in different image groups corresponding to the target image grouping method, then the feature points of the preset foreground region in the image pair are matched to obtain the foreground feature matching pair of the image pair. Aerial triangulation is performed based on the foreground feature matching pairs to obtain the sparse point cloud of the target object corresponding to the image pair; Based on the image grouping results corresponding to the target image grouping method, dense matching is performed on the sparse point cloud of the target object corresponding to the image pair to obtain the dense point cloud of the target object. Based on the dense point cloud, a 3D model of the target object is obtained; The step of performing dense matching on the sparse point cloud of the target object corresponding to the image pair based on the image grouping result corresponding to the target image grouping method to obtain the dense point cloud of the target object includes: Based on the image grouping results corresponding to the target image grouping method, select multiple images from the same group and images from different groups for each image to form stereo image pairs from the same group and stereo image pairs from different groups for each image. Dense matching is performed on the sparse point clouds of the same group of stereo image pairs and the sparse point clouds of different groups of stereo image pairs. The dense point cloud of the target object is obtained by combining the same set of stereo image pairs and the different sets of stereo image pairs after dense matching.

2. The method of claim 1, wherein, The step of determining the target image grouping method from multiple image grouping methods based on the feature points of the image pair includes: Based on the feature points of the image pair, feature matching is performed on the feature points of the two images in the image pair to obtain multiple sets of feature matching pairs of the image pair. The target image grouping method is determined from the plurality of image grouping methods based on multiple feature matching pairs of the image pairs.

3. The method of claim 2, wherein, The step of determining the target image grouping method from the multiple image grouping methods based on multiple feature matching pairs of the image pair includes: The image pair is thinned out by multiple feature matching pairs to obtain the target feature matching pair of the image pair, such that the ratio of the number of feature points in the preset foreground region to the number of feature points in the preset background region in the target feature matching pair is at a preset ratio. Based on the target feature matching pairs of the image pairs, the target image grouping method is determined from the plurality of image grouping methods.

4. The method of claim 3, wherein, Before performing thinning processing on multiple feature matching pairs of the image pair to obtain the target feature matching pair of the image pair, the method further includes: Based on multiple feature matching pairs of each image pair, determine the feature matching status of the preset foreground region and the preset background region; Based on the feature matching results, image pairs that meet the preset feature matching conditions are determined as target image pairs from each image pair; The step of thinning multiple feature matching pairs of the image pair to obtain the target feature matching pair of the image pair includes: The target image pair is subjected to a thinning process on multiple feature matching pairs to obtain the target feature matching pair of the target image pair.

5. The method of claim 4, wherein, The step of determining the target image grouping method from the plurality of image grouping methods based on the target feature matching pairs of the image pairs includes: Calculate the edge weights of each target image pair based on the target feature matching pairs of the target image pairs; Based on the edge weights of each target image pair, calculate the grouping cost function value corresponding to the multiple image grouping methods; Based on the grouping cost function values ​​corresponding to the multiple image grouping methods, the image grouping method with the smallest grouping cost function value is selected as the target image grouping method.

6. The method according to claim 1, characterized in that, The step of performing dense matching on the sparse point clouds of the same set of stereo image pairs and the sparse point clouds of different sets of stereo image pairs includes: The depth values ​​of corresponding points in the two images of the sparse point cloud of the same stereo image pair are matched and verified to obtain the same stereo image pair after dense matching. The depth values ​​of corresponding points in two images of the sparse point cloud of the different sets of stereo image pairs are matched and verified to obtain the different sets of stereo image pairs after dense matching.

7. A three-dimensional reconstruction device, characterized in that, The device includes: The extraction module is used to extract features from multiple images of a target object to obtain feature points from the multiple images. The combination module is used to combine the multiple images in pairs to obtain image pairs; The determination module is used to determine the target image grouping method from multiple image grouping methods based on the feature points of the image pair; The matching module is used to match feature points of a preset foreground region in the image pair if the two images in the image pair are respectively in different image groups corresponding to the target image grouping method, so as to obtain a foreground feature matching pair of the image pair. The measurement module is used to perform aerial triangulation based on the foreground feature matching pair to obtain the sparse point cloud of the target object corresponding to the image pair; The dense matching module is used to perform dense matching on the sparse point cloud of the target object corresponding to the image pair according to the image grouping result corresponding to the target image grouping method, so as to obtain the dense point cloud of the target object. The modeling module is used to perform three-dimensional modeling based on the dense point cloud to obtain a three-dimensional model of the target object. The dense matching module is specifically used to select multiple images from the same group and different groups for each image according to the image grouping results corresponding to the target image grouping method, forming stereo image pairs from the same group and stereo image pairs from different groups for each image; perform dense matching on the sparse point clouds of the stereo image pairs from the same group and the sparse point clouds of the stereo image pairs from different groups respectively; and obtain the dense point cloud of the target object based on the densely matched stereo image pairs from the same group and the stereo image pairs from different groups.

8. An electronic device, characterized in that, include: The processor and the storage medium are connected via a bus for communication. The storage medium stores program instructions executable by the processor. The processor calls the program stored in the storage medium to execute the steps of the three-dimensional reconstruction method as described in any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, performs the steps of the three-dimensional reconstruction method as described in any one of claims 1 to 6.