Deep learning-based broken object recombination method and system

By combining deep learning and RANSAC algorithms, using MaskNet++ network segmentation and registration point clouds, the problems of information loss and feature matching difficulties in three-dimensional fragment reconstruction are solved, and more accurate and robust reorganization of broken objects is achieved.

CN119941575APending Publication Date: 2025-05-06SHANGHAI INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411787643.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has lost information, incomplete or errors in the three-dimensional fragment reconstruction process, and it is difficult to match features in areas with low texture and contrast, resulting in artifacts and missing parts.

Method used

Using a deep learning-based method, combined with the RANSAC algorithm and MaskNet++ network, the precise reorganization of broken objects is achieved through the internal and external points segmentation and registration of the point cloud.

Benefits of technology

It improves the accuracy and robustness of three-dimensional fragment splicing, reduces artifacts and missing parts, and enhances the integrity and reliability of reconstruction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941575A_ABST
    Figure CN119941575A_ABST
Patent Text Reader

Abstract

The invention discloses a broken object recombination method and system based on deep learning. The method comprises the following steps: S1, obtaining a data set required by an experiment; step S2, estimating an optimal alignment position and a transformation matrix between the two groups of point clouds by using an RANSAC (Random Sample Consensus) algorithm; s3, performing visual display on the data subjected to the RANSAC algorithm by using a visual tool; s4, dividing the two groups of point clouds into inner points and outer points by using a MaskNet + + network; and S5, performing registration between the fragment pairs by using an RANSAC registration method, and restoring the object. By adopting the method, the optimal alignment position between two groups of point clouds and the transformation relation between matrixes can be estimated so as to realize accurate space alignment and carry out a series of hidden adjacency relation identification on input broken object fragments, and preparation is made for the next step of restoration. In addition, the method pays more attention to division of the fracture area and extraction of key points of the fracture area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular to a method and system for reconstructing broken objects based on deep learning. Background Art

[0002] With the rapid development of LiDAR scanning technology, people can more easily obtain three-dimensional models of real objects. These models have been widely used in various fields, such as computer-aided design and manufacturing, medical diagnosis, molecular biology research, games and film and television animation, cultural relics and archaeology, etc. In the process of obtaining three-dimensional models, due to various reasons (such as the limitations of scanning equipment, the complexity of objects, etc.), incomplete or fragmented three-dimensional models are often obtained. Therefore, how to realize the automatic splicing of these fragments has become an urgent problem to be solved.

[0003] Since traditional 3D reconstruction algorithms have occluded areas in the process of 3D fragment reconstruction, this will lead to serious information loss, making the reconstruction results incomplete or erroneous. In the process of splicing fragments, the topological structure of the object may change, such as breaking and reorganizing, which also puts higher requirements on the robustness and accuracy of the algorithm. At the same time, in 3D reconstruction based on multiple views, the accuracy of feature matching is crucial to the reconstruction effect. In areas with low texture and contrast, the difficulty of feature matching may lead to obvious artifacts and missing parts in the reconstruction results, such as incomplete fragment matching.

[0004] In recent years, deep learning, as a new research direction in machine learning algorithms, has achieved fruitful results in the fields of image processing, natural language processing, etc. Applying deep learning technology to 3D fragment splicing can extract more effective features, thereby improving the accuracy and robustness of splicing. It can also be combined with traditional algorithms to solve a series of problems. Summary of the invention

[0005] The purpose of the present invention is to provide a method and system for reconstructing broken objects based on deep learning to solve the problems raised in the above background technology.

[0006] To achieve the above-mentioned object of the invention, one aspect of the present invention provides a method for reconstructing broken objects based on deep learning, comprising the following steps:

[0007] Step S1, obtaining the data set required for the experiment;

[0008] Step S2, using the RANSAC algorithm (Random Sample Consensus) to estimate the optimal alignment position and transformation matrix between the two sets of point clouds;

[0009] Step S3, using a visualization tool to visualize the data after the RANSAC algorithm is performed;

[0010] Step S4, using the MaskNet++ network to divide the two groups of point clouds into internal points and external points;

[0011] Step S5: Use the RANSAC registration method to perform registration between fragment pairs and restore the object.

[0012] Further, step S1 includes the following steps:

[0013] Step S101, using a file opening tool to analyze the header and main data part of the point cloud file, and then using pytorch to read the file according to its file format;

[0014] Step S102, using PCL-related algorithm libraries and point cloud file visualization software to read the point cloud file and test the data set;

[0015] Step S103, obtain the fragmented data set to be used for the experiment, including the FragTag, Breaking Bad, and Fantastic Breaks data sets.

[0016] Further, step S103 includes the following steps:

[0017] Step S131, downloading the original data set or using laser radar to construct a three-dimensional object to generate original data;

[0018] Step S132, using CloudCompare software to use the seg command to delete useless points and set the color;

[0019] Step S133, using the CloudCompare tool to select regions or objects for segmentation, label them, set object names and label values, and then synthesize the labeled data into a point cloud dataset;

[0020] Step S134: export the point cloud data set in a file format.

[0021] Further, step S2 includes the following steps:

[0022] Step S201, randomly selecting m points from the point cloud data set as samples;

[0023] Step S202, estimating the pose transformation of the m points in step S201, and calculating a local optimal rigid body transformation T, rotation matrix R and translation vector t;

[0024] Step S203, perform model evaluation, calculate the alignment error of point pairs under transformation T, and determine which ones are inliers;

[0025] Step S204, updating the optimal solution, and considering the transformation T that generates more points as the current optimal transformation.

[0026] Furthermore, the MaskNet++ network consists of five spatial self-attention blocks and a group of multi-layer perceptrons, and the spatial self-attention blocks are used to share three channel cross-attention blocks for global feature extraction.

[0027] Further, step S4 includes the following steps:

[0028] Step S401, given two point clouds X = {x i ∈R 3 |i=1,…,N} and point cloud Y={y i ∈R 3 |i=1,…,M}, define the inner point X of two point clouds X and Y I , Y I and X O , Y O as follows:

[0029] X I =X∩Y,Y I =X∩YX O =XY,Y O =XY

[0030] Step S402, find two binary vectors C x ={c xi ∈{0,1}|j=1,…,N} and C y ={c yi ∈{0,1}|j=1,…,M}, so that the following conditions are satisfied:

[0031]

[0032] Where R∈SO(3) represents the rotation matrix, SO(3) represents all rotation transformations that preserve the length and angle of the vector in three-dimensional space, and t∈R 3 represents the translation between X and Y,

[0033] Operation Symbols The definition is as follows:

[0034]

[0035] Where X I is the interior point set of X, x i ∈X I , if C xi =1;Y I is the interior point set of Y, y i ∈YI ;

[0036] Step S403: Calculate the mask of point cloud X First, the shared global features of point cloud Y are repeated N times to have the same dimension as the point-by-point features φ(X) of point cloud X. These concatenated features are then input into the mask estimation subnetwork, where a sigmoid activation function is used in the last layer of h(·) to enforce C * ∈[0,1] and θ(·) that share global features to estimate the internal points. The mask algorithm of Y is similar and can be expressed as follows:

[0037]

[0038] Another aspect of the present invention provides a system for reconstructing broken objects based on deep learning, comprising a collection module, a data set module, a visualization module, a splitting module, and a registration module, wherein:

[0039] The collection module is used to obtain the data set required for the experiment;

[0040] The dataset module uses the RANSAC algorithm to estimate the optimal alignment position and transformation matrix between two sets of point clouds;

[0041] The visualization module uses visualization tools to visualize the data that has been processed by the RANSAC algorithm;

[0042] The splitting module uses the MaskNet++ network to divide the two groups of point clouds into internal points and external points;

[0043] The registration module uses the RANSAC registration method to register fragment pairs and restore objects.

[0044] Compared with the prior art, the present system and method have the following advantages:

[0045] 1. The present invention combines MaskNet++ and RANSAC algorithms to estimate the optimal alignment position between two sets of point clouds and the transformation relationship between matrices to achieve accurate spatial alignment, and perform a series of hidden adjacency relationship recognition on the input broken object fragments to prepare for the next step of restoration.

[0046] 2. The present invention overcomes the traditional reconstruction method that focuses on the geometric features of each broken object fragment. The present method focuses more on the division of the fracture area and the extraction of key points in the fracture area. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 Flowchart of a method for reconstructing broken objects based on deep learning.

[0048] Figure 2 Schematic diagram of the principle of a method for reconstructing broken objects based on deep learning.

[0049] Figure 3 This is the clustering result diagram of the RANSAC algorithm.

[0050] Figure 4 This is the result diagram of recognition and extraction of overlapping areas.

[0051] Figure 5 This is the neural network diagram of MaskNet++.

[0052] Figure 6 This is the final result of the experiment using this method. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] like Figure 1 , Figure 2 The flowchart and principle diagram of the method of the present invention are shown. The embodiment of the present invention provides a method for reconstructing broken objects based on deep learning, and the steps are as follows:

[0055] Step S1, obtain the data set required for the experiment. This includes the following steps:

[0056] Step S101, analyze the point cloud data format such as csv, txt, ply, pcd and other point cloud format files. You can use a file opening tool (such as a text editor) to analyze the header and main part of different point cloud files, and then use pytorch to read the file according to its file format.

[0057] Step S102: After installing CloudCompare and configuring the PCL library, you can modify the original data set or annotate the point cloud file yourself to create a data set. Use the PCL-related algorithm library and visualization software such as CloudCompare that can read point cloud files to test the data set.

[0058] Step S103, obtaining a data set to be used for the experiment, such as a fragmented data set such as FragTag, Breaking Bad, FantasticBreaks, etc., which includes the following steps:

[0059] Step S131, downloading original data sets (such as Thingi10K, Fantastic Breaks, etc.) or using laser radar to construct three-dimensional objects to generate original data.

[0060] Step S132, use the CloudCompare software to use the seg command to delete useless points, and set the color in Edit—Colors—Set Unique.

[0061] Step S133, use the CloudCompare tool to select regions or objects for segmentation. Assign a label to each cut-out part, giving it an object name and a value. This label value is the value after network training. Then combine these labeled data into a complete point cloud dataset. Use the "Merge MultipleClouds" option to merge.

[0062] Step S134, export these data sets in a suitable file format for later use.

[0063] Step 2: Use the RANSAC algorithm to estimate the optimal alignment position and transformation matrix R between the two sets of point clouds. RANSAC is a robust parameter estimation method that is widely used to process data sets with a high proportion of outliers. The RANSAC algorithm is used to estimate the optimal alignment or transformation matrix between two sets of point clouds. Its steps can be summarized as follows:

[0064] In step S201, two point cloud data P and Q to be reassembled and matched are used as input data, where P and Q are two sets of point cloud data in the data set or three-dimensional objects composed of other point clouds scanned by laser radar, and m point pairs are randomly selected as samples.

[0065] Step S202: perform pose transformation estimation. Calculate a local optimal rigid body transformation T, rotation matrix R and translation vector t based on the m point set pairs.

[0066] Step S203, evaluating the model: For all point pairs, calculate their alignment errors under transformation T and determine which ones are inliers.

[0067] Step S204, updating the optimal solution. If the current transformation T can generate more points, it is regarded as the current optimal transformation.

[0068] The process described above is expressed in the following formula:

[0069] For a point p in sets P and Q i and q i , for p i Get q by rotating the matrix R and translating the variable ti :

[0070] q i ≈Rp i +t

[0071] Calculate the centroid of two sets of points

[0072]

[0073] Then for each point set p i ,q i Decentralize and move the center of mass to the origin:

[0074]

[0075] The covariance matrix C is then constructed using the centered point set:

[0076]

[0077] Perform singular value decomposition on the covariance matrix C:

[0078] C=USV T

[0079] Among them, U and V are orthogonal matrices, S is a diagonal matrix, and then U and V are used to calculate the rotation matrix R:

[0080] R=VU T

[0081] Finally, the rotation matrix and translation vector are combined into the second change matrix T:

[0082]

[0083] Repeat the above steps until the stopping condition is met (if the number of iterations reaches the set threshold, a good enough transformation is found). The best transformation T output at the end will be the transformation found during the entire iteration process that can make the most point pairs become inliers.

[0084] Step 3: Use visualization tools to visualize the data after the RANSAC algorithm, such as Figure 3 , as shown in 4.

[0085] Step 4: Use the MaskNet++ network to classify the two groups of point clouds into inliers and outliers based on the recognition results of the RANSAC algorithm. Figure 5The figure shows a schematic diagram of the MaskNet++ network. It consists of five spatial self-attention (SSA) blocks, three channel cross-attention (CCA) blocks for shared global feature extraction, and a set of multilayer perceptrons (MLP). Inliers are points in the overlapping area of ​​two point clouds, while outliers are points in the non-overlapping area.

[0086] Step S4 includes the following steps:

[0087] Step S401: Given two point cloud data X = {x i ∈R 3 |i=1,…,N} and point cloud Y={y i ∈R 3 |i=1,…,M} as input data, define the inner point X of two point clouds X and Y I , Y I and X O , Y O as follows:

[0088] X I =X∩Y,Y I =X∩Y

[0089] X O =XY,Y O =XY

[0090] Step S402, find two binary vectors C x ={c xi ∈{0,1}|j=1,…,N} and C y ={c yi ∈{0,1}|j=1,…,M}, so that the following conditions are satisfied:

[0091]

[0092] Where R∈SO(3) represents the rotation matrix, SO(3) represents all rotation transformations that preserve the length and angle of the vector in three-dimensional space, and t∈R 3 Indicates the amount of translation between X and Y.

[0093] Operation Symbols The definition is as follows:

[0094]

[0095] Where X I is the interior point set of X, x i ∈X I , if Cxi =1;Y I is the interior point set of Y, y i ∈Y I .

[0096] Step S403: Calculate the mask of point cloud X First, the shared global features of point cloud Y are repeated N times to have the same dimension as the point-by-point features φ(X) of point cloud X. These concatenated features are then input into the mask estimation subnetwork, where a sigmoid activation function is used in the last layer of h(·) to enforce C * ∈[0,1] and θ(·) that share global features to estimate the inliers, and the mask algorithm of Y is the same.

[0097]

[0098]

[0099] Step 5: Use the RANSAC registration method to register the fragment pairs. In this experiment, the relevant parameters are set to scale S = 5, and the number of key point pairs used for registration is n f =30, and the threshold of the control matrix validity ε=5. In the experiment of quantitative comparison, considering that the fragment reconstruction framework includes the registration of multiple fragments, the average value of the chamfer distance and the average value of the normal consistency after the registration of all fragments are used as the measurement indicators of the final matching effect. In addition, the experiment also observed the effect of the final reconstruction, as well as the success rate of the overall reconstruction, and recorded the time required for the reconstruction process. This method allows for accurate evaluation of the performance of the fragment reconstruction framework, while providing an intuitive understanding of the reconstruction effect and efficiency of different object models. Then, experiments were conducted on feature extraction in different regions, and the best effect of feature extraction in overlapping regions was analyzed from a quantitative perspective, which was used for overall reconstruction of objects containing complex fragmentation relationships. Comparative experiments were also conducted to demonstrate the superiority of feature matching based on overlapping regions.

[0100] like Figure 6 The final result of the experiment using this method can be seen in the figure. It can be seen that the effect of restoring multiple broken objects using overlapping region features is still the most outstanding. The average normal consistency between the broken object fragments is improved by 56.49% and 17.95% respectively when using overlapping region features compared with global region features and fracture region features, and the average chamfer distance is reduced by 20.55% and 15.40% respectively. The registration success rate is also increased from 16.67% (2 out of 12 were successfully restored) and 50.34% (7 out of 12 were successfully restored) to 89.77% (11 out of 12 were successfully restored).

[0101] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for reconstructing broken objects based on deep learning, characterized in that: The following steps are involved: Step S1, obtaining the data set required for the experiment; Step S2, using the RANSAC algorithm to estimate the optimal alignment position and transformation matrix between the two sets of point clouds; Step S3, using a visualization tool to visualize the data after the RANSAC algorithm is performed; Step S4, using the MaskNet++ network to divide the two groups of point clouds into internal points and external points; Step S5: Use the RANSAC registration method to perform registration between fragment pairs and restore the object.

2. The method for reconstructing broken objects based on deep learning according to claim 1, characterized in that: Step S1 includes the following steps: Step S101, using a file opening tool to analyze the header and main data part of the point cloud file, and then using pytorch to read the file according to its file format; Step S102, using PCL-related algorithm libraries and point cloud file visualization software to read the point cloud file and test the data set; Step S103, obtain the fragmentation data set to be tested, including FragTag, Breaking Bad,Fantastic Breaks dataset.

3. The method for reconstructing broken objects based on deep learning according to claim 2, characterized in that: Step S103 includes the following steps: Step S131, downloading the original data set or using laser radar to construct a three-dimensional object to generate original data; Step S132, using CloudCompare software to use the seg command to delete useless points and set the color; Step S133, using the CloudCompare tool to select regions or objects for segmentation, label them, set object names and label values, and then synthesize the labeled data into a point cloud dataset; Step S134: export the point cloud data set in a file format.

4. The method for reconstructing broken objects based on deep learning according to claim 1, characterized in that: Step S2 includes the following steps: Step S201, randomly selecting m points from the point cloud data set as samples; Step S202, estimating the pose transformation of the m points in step S201, and calculating a local optimal rigid body transformation T, rotation matrix R and translation vector t; Step S203, perform model evaluation, calculate the alignment error of point pairs under transformation T, and determine which ones are inliers; Step S204, updating the optimal solution, and considering the transformation T that generates more points as the current optimal transformation.

5. The method for reconstructing broken objects based on deep learning according to claim 1, characterized in that: The MaskNet++ network consists of five spatial self-attention blocks and a set of multi-layer perceptrons. The spatial self-attention blocks are used to share three channel cross-attention blocks for global feature extraction.

6. The method for reconstructing broken objects based on deep learning according to claim 1, characterized in that: Step S4 includes the following steps: Step S401, given two point clouds X = {x i ∈R 3 |i=1,…,N} and point cloud Y={y i ∈R 3 |i=1,…,M}, define the inner point X of two point clouds X and Y I , Y I and X O , Y O as follows: X I =X∩Y,Y I =X∩Y X O =XY,Y O =XY; Step S402, find two binary vectors and So that the following conditions are met: Where R∈SO(3) represents the rotation matrix, SO(3) represents all rotation transformations that preserve the length and angle of the vector in three-dimensional space, and t∈R 3 represents the translation between X and Y, Operation Symbols The definition is as follows: Where X I is the interior point set of X, x i ∈X I ,if Y I is the interior point set of Y, y i ∈Y I ; Step S403: Calculate the mask of point cloud X First, the shared global features of point cloud Y are repeated N times to have the same dimension as the point-by-point features φ(X) of point cloud X. These concatenated features are then input into the mask estimation subnetwork, where a sigmoid activation function is used in the last layer of h(·) to enforce C * ∈[0,1] and θ(·) that share global features to estimate the internal points. The mask algorithm of Y is similar and can be expressed as follows:

7. A system for reconstructing broken objects based on deep learning, characterized in that: It includes collection module, data set module, visualization module, segmentation module and registration module, among which: The collection module is used to obtain the data set required for the experiment; The dataset module uses the RANSAC algorithm to estimate the optimal alignment position and transformation matrix between two sets of point clouds; The visualization module uses visualization tools to visualize the data that has been processed by the RANSAC algorithm; The splitting module uses the MaskNet++ network to divide the two groups of point clouds into internal points and external points; The registration module uses the RANSAC registration method to register fragment pairs and restore objects.

Citation Information

Patent Citations

  • Electrical switchboard.

    US760077A