Diversified shape completion method based on component analysis and decomposition
Through the diversity shape completion method of component analysis and decomposition, the problem of difficulty in recovering complete three-dimensional information from defective cloud data in the prior art is solved, and high-quality three-dimensional object diversity completion is achieved on the basis of retaining the input shape.
Patent Information
- Application Number
- CN202510110268.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The prior art is difficult to effectively restore complete three-dimensional information from incomplete point cloud data, especially on the basis of retaining the input shape, to achieve diversity completion of three-dimensional objects.
Through component analysis and decomposition, a complete particle size reduction is generated to achieve a clear division of existing and missing areas. The template structural parameters are predicted using deep learning networks to derive the complete structural representation of the object, and restore the complete details of the three-dimensional object through three-view completion and diversity generation.
It realizes the high-quality completion of the diverse shapes of three-dimensional objects on the basis of retaining the input shape, improves the integrity and interpretability of point cloud completion, and significantly improves the ability to recover details in complex categories.
Smart Images

Figure CN120031835A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of point cloud generation, and in particular to a diversity shape completion method based on component analysis and decomposition. Background Art
[0002] This section merely provides background information related to the present disclosure and is not necessarily prior art.
[0003] Point cloud data is three-dimensional spatial data obtained through sensors such as laser scanning and depth cameras, which can accurately describe the geometric shape of objects or environments. Compared with traditional two-dimensional image data, point clouds not only provide the location of the surface of objects, but also have rich spatial information. Therefore, they are widely used in autonomous driving, robot perception, architectural modeling, AR / VR, digital twins and other fields. However, in real-world application scenarios, incomplete point clouds are inevitably obtained, resulting in the inability to effectively utilize point cloud data. Therefore, how to restore complete three-dimensional information from these incomplete point cloud data is currently a major challenge in computer vision and computer graphics.
[0004] Existing shape completion methods use incomplete signed distance fields or placeholder voxels as input to predict the complete signed distance field of an object. However, in real scenarios, it is impractical to obtain a partial signed distance field from a partial point cloud. At the same time, it is also extremely important to achieve diverse predictions of missing areas. Existing methods often have difficulty retaining the input shape, resulting in poor robustness when applied to real scanned objects. Summary of the invention
[0005] Purpose of the invention: In view of the shortcomings of the prior art, the present invention discloses a diversified shape completion method based on component analysis and decomposition, which aims to outline the details inside the structure through the three-view images after component decomposition, complete the missing details in the form of two-dimensional image generation, and restore the three-dimensional object from the complete three-view images. The proposed technology reduces the granularity of the generated completion through structural analysis and decomposition of components, while achieving a clear division of existing and missing areas, and can achieve diversified completion of three-dimensional objects on the basis of retaining the input shape.
[0006] Technical solution: A diverse shape completion method based on component analysis and decomposition, comprising the following steps:
[0007] Step 1: Represent the object as an abstract template structure. The object is incomplete input data represented by a three-dimensional point cloud. The template structure parameters matched with the incomplete point cloud are predicted by a deep learning network, and the complete structure representation of the object is derived based on the template structure parameters.
[0008] Step 2: Based on the complete structural representation of the object obtained in step 1, the object is decomposed into components, the projection details of the internal points of each component on three orthogonal planes are obtained, and the projection details are input into the network to generate the solid detail information of each component.
[0009] Step 3: Complete the three-view drawing, combine the detailed information of the symmetrical parts, and deduce and complete the complete component details of the object.
[0010] Step 4: By comparing the incomplete and complete three-view images obtained in step 3, the missing area is distinguished and identified, and it is determined whether diversity generation is required based on the missing area.
[0011] Step 5: Mask out the area that needs diversity generation in the complete detail image, perform image-to-image diffusion generation, and obtain diversified complete detail three-view images.
[0012] Step 6: Reconstruct the 3D region based on the three views, obtain the corresponding signed distance field, and obtain a watertight completion result through the marching cube algorithm.
[0013] Furthermore, in step 1, the abstract template structure of object matching is a general structure based on category design, which includes M interrelated primitive boxes, each of which corresponds to a part in the object, and each box encloses a part of the object. The structural relationship, position and size of the boxes reflect the structural relationship, relative size and rough shape between the parts of the object.
[0014] Each primitive box is represented by 8 parameters, including three coordinate parameters of the starting point, three coordinate parameters of the ending point, and two dimension parameters representing the length and width of the primitive box. The relationship between boxes is defined by the connection constraints between key points.
[0015] There are 26 key points in total, corresponding to the center points of the 6 faces, 8 vertices and the midpoints of the 12 edges of the unit cube. In addition to the connection relationship between components, the definition of the template structure also includes the semantic information and symmetry of the components, which helps to more accurately describe the overall structure and properties of the object.
[0016] Furthermore, after matching the corresponding abstract template structure, through the network N s To learn the parameters of each box that best fits the input residual point cloud. Get the complete object structure that matches the input point cloud. s Taking the three-channel coordinates of the input point cloud as input features, the network outputs n values after the corresponding encoding and decoding process, which correspond to the parameters of its structural template. The general structural template is adjusted through the output parameters to make the template structure more fit the input point cloud, thereby obtaining the complete structure of the input point cloud.
[0017] The deep network N s The data augmentation strategy is used in the training scheme of , which is to randomly rotate the complete point cloud in the dataset and remove some points according to the z-axis segmentation to obtain the input residual point cloud. The selection of the removed points can be achieved by setting a threshold and cutting off some points whose z-axis is greater than the threshold after rotation. This enhances the robustness of the network to various residual point cloud distributions.
[0018] Furthermore, step 2 divides the input point cloud according to the box position of each component in the structure, and obtains the projection of each component point cloud on three orthogonal planes, thereby extracting the detailed information inside the input area structure. A network is used to process these projections to supplement and improve the closed contour lines and solid three-view images. The specific process can be divided into the following two sub-steps:
[0019] Step 2-1: According to the position and proportion of each box obtained in step 1, the input point cloud is segmented into components, and the projection of each component point cloud on three orthogonal planes is obtained, and the projection is used as the original multi-hole three-view image, which describes the structural details inside the input area;
[0020] Step 2-2: Use the component projection image obtained in step 2-1 as input to train the network N f , through the network N f The solid three-view details of each component are inferred from the porous three-view detail prediction. This makes the detail description more refined and accurate. The three-view obtained from the point cloud projection of the original input component in step 2-1 is porous and cannot outline the details of the input structure with a solid area surrounded by an accurate outline, making it difficult for subsequent steps to clearly define the input area and generate a watertight mesh result. The network N described in step 2-2 needs to be introduced. f To predict the details of each part's solidity.
[0021] Furthermore, the prediction process described in step 2-2 uses a data augmentation strategy, which allows the method to fill in any porous details. This strategy divides the image by a random straight line, and then randomly selects the area above or below the straight line as a mask. The mask is applied to both the porous detail map obtained by projecting the complete point cloud and the solid three-view image that comes with the dataset, which are represented as and , as the network N f The data pairs required for training, during the training process, the network N f Learning by recover .
[0022] Furthermore, step 3 aims to complete the completion of component details. Although the solid three-view obtained in step 2 has provided a preliminary part outline, there are still details in the missing area. Through this step, the method can supplement and improve these missing details and generate complete part information. The specific process can be divided into the following two sub-steps:
[0023] Step 3-1: For each box, obtain the projected three-view image of the point cloud within its symmetric component as an additional input to the network.
[0024] Step 3-2: Train the network N c ,The network takes the original porous three view details of the part as input, predicts and outputs the complete solid point cloud details, completing the missing areas.
[0025] The symmetry information between the components in step 3-1 is included in the general template structure definition corresponding to the category. After the symmetrical detail diagram is obtained, it will be concatenated with the original detail diagram as the network input.
[0026] The step 3-2 requires the complete point cloud to be input into the network N described in step 2 f The solid detail map obtained in the training network N c Tags. Network N c Learning to recover full details from porous details in residual clouds.
[0027] The network N c A data augmentation strategy is used during the training process; the data augmentation strategy is the same as the data augmentation strategy in step 1. This increases the robustness of the network.
[0028] Furthermore, step 4 identifies the exact area by comparing the missing details of the image, providing a basis for preserving the input shape. The specific process can be divided into the following sub-steps:
[0029] Step 4-1: Comparatively complete solid detail D c and incomplete solid details D s , identify the overlapping area as the original input area A o .
[0030] Step 4-2: For each part P, complete the three views Symmetrical mirroring is performed according to three orthogonal planes to obtain the solid three-view image of the mirrored part Check its complete three views Whether it is a solid three-view image with its mirror image Overlap, calculate the overlap rate of the three views separately If one of the overlap ratios r of the three views does not exceed the threshold T 1 , then the component is not a symmetrical component and its missing area needs to be diversified.
[0031] The mirroring refers to symmetrically mirroring the three views according to three orthogonal planes. For example, if you need to obtain the mirrored three views about the xy plane, you only need to mirror the xy view vertically and mirror the yz and xz views horizontally.
[0032] Step 4-3: For each symmetrical part P, obtain the symmetry relationship of the parts from the complete structure of the object. If part P has a symmetrical component, merge the solid three-view image of its mirrored part and the solid three-view image of its symmetrical component as the area to be inspected, and then define the overlapping area of the solid three-view image of part P and the area to be inspected as the associated input area A. r ; Component details d c and d s The overlapping area is taken as the original input area A o . Calculation area A o ∪A r The proportion of complete part details exceeds the threshold T 2 , it means that the missing region can be directly filled by symmetric information without the need for diversity generation; otherwise, it means that the missing region needs diversity generation.
[0033] Furthermore, the threshold T in step 4-2 1 The value range is 0.7-0.9. The threshold T in step 4-3 2 The value range is 0.5-0.7.
[0034] Furthermore, step 5 is based on the diffusion model N of the graph. g Realize diversity completion, which is the diffusion model N of graph-generated graphs g Image information is injected, and the noisy source image is used as initialization. The source image is a complete detail image masked by a random mask; during the denoising process, the diffusion model learns to recover the complete detail image without the mask from the noise. The randomness and data complexity of this process enable the model to capture more details during the generation process, thereby enhancing the diversity and accuracy of the output.
[0035] Step 6: Recover the 3D face from the complete three-view image. A point in the 3D space can be uniquely determined based on the rays of the three orthogonal plane projections. The distance between the projection point on the three views and the boundary is calculated, and the minimum value is selected as the spatial distance. If the point is uniquely determined by the points in the figure enclosed by the contour in the three projection views, it means that it is inside the closed 3D figure and the value is negative; if the point is uniquely determined by the points on the contour in the three projection views, it means that it is exactly on the surface of the closed 3D figure and the value is 0; otherwise, it is outside the closed 3D figure and the value is positive. Based on this, the corresponding signed distance field representation is obtained, and then the watertight face is recovered according to the marching cube algorithm.
[0036] The signed distance field is a function that represents distance. For any point in three-dimensional space, the signed distance field represents the shortest distance from the point to the surface of the object. If the point is outside the object, the value is positive; if it is inside the object, the value is negative; if it is on the surface of the object, the value is zero. Through the signed field, the shape and position of the object in space can be determined, providing a basis for subsequent surface extraction and mesh reconstruction. For each three-dimensional query point, the distance between the projection of the query point on the three views and the boundary is selected, and the minimum value is the signed distance field of the query point.
[0037] Further, step 6 includes the following steps,
[0038] Step 6-1: Obtain the signed distance field and reconstruct the 3D region based on the three-view image. For each 3D query point, the distance between the projection of the query point on the three-view image and the boundary is selected, and the minimum value is taken as the signed distance field of the query point. The shape and position of the object in space can be determined through the signed distance field.
[0039] Step 6-2: Extract isosurfaces from the signed distance field obtained in step 6-1, and restore watertight three-dimensional surfaces using a marching cube algorithm based on the isosurfaces. The marching cube algorithm is a classic three-dimensional surface reconstruction algorithm used to restore a closed and complete three-dimensional surface without holes.
[0040] Beneficial effects:
[0041] The present invention provides a method for completing a diverse shape based on component analysis and decomposition, which has the following beneficial effects:
[0042] This method makes full use of structural information for point cloud completion. By analyzing the structure of three-dimensional objects, this method can complete the incomplete point cloud into a complete mesh. In this process, the existing structural information in the input point cloud is fully utilized to ensure that the completed mesh is geometrically coherent. This completion strategy based on structural information also helps to improve the integrity and interpretability of the completed point cloud data, and can still effectively restore important details when there are large missing areas.
[0043] The proposed method of component decomposition completion and generation focuses only on the diverse prediction of a single component, thereby effectively preserving the input shape. In this way, the system can not only cope with the morphological changes of different components, but also reduce the risk of over-filling missing areas or shape deviation from the prototype, ensuring that the generated parts are highly consistent with the original input.
[0044] This method has been demonstrated to be superior through a large number of experiments, achieving SOTA performance on both synthetic and real scan datasets. Whether in terms of point cloud completion accuracy, the diversity of generated effects, or the ability to recover details of complex categories, this method has shown superior performance.
[0045] The present invention draws inspiration from the structural information of the object and introduces an innovative diversity shape completion method using local structural analysis. After the structural analysis, each component can be separated, the relationship between the components can also be determined, and the missing parts can be more accurately completed and reconstructed based on the existing related parts (such as symmetrical corresponding parts). In addition, focusing on the diversity prediction of the missing area of a single component not only makes the completion simpler and more accurate, but also brings richer changes to the final completion result. This method uses projected three-view images to describe the details of the components, and realizes the diversity completion of three-dimensional objects through the completion and detail generation of two-dimensional three-view images, and can obtain reconstruction results on the basis of effectively retaining the input shape. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0047] Figure 1 The basic flow chart of the method of the present invention is shown in FIG.
[0048] Figure 2 The present invention is summarized in the following.
[0049] Figure 3 Schematic diagram of the component structure analysis process.
[0050] Figure 4 Schematic diagram of the process of component decomposition and detail extraction.
[0051] Figure 5 Schematic diagram of the component completion process.
[0052] Figure 6 Schematic diagram of the missing region detection process.
[0053] Figure 7 Schematic diagram of the process generated for diversity details.
[0054] Figure 8 Schematic diagram of the process of recovering triangular facets from detail images.
[0055] Fig. 9 Schematic diagram of the comparison of point cloud completion results on synthetic datasets.
[0056] Fig.10A schematic diagram showing the comparison of point cloud completion results on a real scanned dataset.
[0057] Fig.11 Schematic diagram of the data augmentation strategy. DETAILED DESCRIPTION
[0058] like Figure 1 and Figure 2 As shown, a diverse shape completion method based on component analysis and decomposition includes the following steps:
[0059] Step 1: Represent the object as an abstract template structure. The object is incomplete input data represented by a three-dimensional point cloud. The template structure parameters matched with the incomplete point cloud are predicted by a deep learning network, and the complete structure representation of the object is derived based on the template structure parameters.
[0060] Step 2: Based on the complete structural representation of the object obtained in step 1, the object is decomposed into components, the projection details of the internal points of each component on three orthogonal planes are obtained, and the projection details are input into the network to generate the solid detail information of each component.
[0061] Step 3: Complete the three-view drawing, combine the detailed information of the symmetrical parts, and deduce and complete the complete component details of the object.
[0062] Step 4: By comparing the incomplete and complete three-view images obtained in step 3, the missing area is distinguished and identified, and it is determined whether diversity generation is required based on the missing area.
[0063] Step 5: Mask out the area that needs diversity generation in the complete detail image, perform image-to-image diffusion generation, and obtain diversified complete detail three-view images.
[0064] Step 6: Reconstruct the 3D region based on the three views, obtain the corresponding signed distance field, and obtain a watertight completion result through the marching cube algorithm.
[0065] Figure 2 Medium 1 -M 4 Results generated for 4 diversities.
[0066] like Figure 3 As shown in Figure 1, in step 1, the abstract template structure for object matching is a general structure designed based on categories, which contains M interrelated primitive boxes, each of which encloses a part of the object. The structural relationship, position and size of these boxes reflect the structural relationship, relative size and rough shape between the parts of the object.
[0067] Each primitive box is represented by 8 parameters, including three coordinate parameters of the starting point, three coordinate parameters of the ending point, and two dimension parameters representing the length and width of the primitive box. The relationship between boxes is defined by the connection constraints between key points.
[0068] There are 26 key points in total, corresponding to the center points of the 6 faces, 8 vertices and the midpoints of the 12 edges of the unit cube. In addition to the connection relationship between components, the definition of the template structure also includes the semantic information and symmetry of the components, which helps to more accurately describe the overall structure and properties of the object.
[0069] After matching the corresponding abstract template structure, through the network N s To learn with input residual cloud P p The best fitting box parameters p. Get the complete object structure that matches the input point cloud. Deep network N s Taking the three-channel coordinates of the input point cloud as the input features, the network outputs n values after the corresponding encoding and decoding process, which correspond to the parameters of its structural template. The general structural template is adjusted through the output parameters to make the template structure more suitable for the input point cloud, thereby obtaining the complete structure S of the input point cloud. c .
[0070] The deep network N s The data enhancement strategy used in its training scheme is as follows Fig.11 As shown, the strategy is to randomly rotate the complete point cloud in the data set according to the rotation matrix R, remove some points according to the z-axis segmentation, and then rotate the point cloud according to the rotation matrix R. -1 Flip it back to get the residual cloud of the input. 1 and P 2 is the set z-axis threshold plane, i.e. z=a, a 1 Area is the removal area, a 2 For reserved area.
[0071] This enhances the network's robustness to various defective cloud distributions.
[0072] N s The network is optimized by the mean square error of the predicted structural parameters and label parameters. s The pseudo code of the network is as follows:
[0073]
[0074] like Figure 4As shown in Figure 2, step 2 divides the input point cloud according to the box position of each component in the structure, and obtains the projection of each component point cloud on three orthogonal planes, thereby extracting the detailed information inside the input area structure. A network is used to process these projections to supplement and improve the closed contour lines and solid three-view images. The specific process can be divided into the following two sub-steps:
[0075] Step 2-1: According to the position and scale of each box obtained in step 1, the input point cloud is segmented into components, and the projection of each component point cloud on three orthogonal planes is obtained, and the projection is used as the original porous three-view detail D p , the porous three-view describes the structural details inside the input area.
[0076] Step 2-2: Use the component projection image obtained in step 2-1 as input to train the network N f , through the network N f Detail D of the porous three-view p Predictive reasoning obtains the solid three-view details D of each component s . Make the detailed description more refined and accurate. The three-view image obtained by step 2-1 based on the point cloud projection of the original input component is porous and cannot outline the details of the input structure with a solid area surrounded by an accurate outline, making it difficult for subsequent steps to clearly define the input area and generate a watertight mesh result. It is necessary to introduce the network N described in step 2-2 f To predict the solid details of each part D s .
[0077] The prediction process described in step 2-2 uses a data augmentation strategy, which allows the method to fill in any porous details. This strategy divides the image by a random straight line, and then randomly selects the area above or below the straight line as a mask. The mask is applied to both the porous detail map obtained by projecting the complete point cloud and the solid three-view image that comes with the dataset, represented as and , as the network N f The data pairs required for training, during the training process, N f Learning by recover . Training N f The pseudo code of the network is as follows:
[0078]
[0079] like Figure 5 As shown in Figure 3, step 3 aims to complete the completion of component details. Although the solid three-view obtained in step 2 has provided a preliminary part outline, there are still details in the missing area. Through this step, the method can supplement and improve these missing details and generate complete details D c, that is, complete part information. The specific process can be divided into the following two sub-steps:
[0080] Step 3-1: For each box, obtain the projected three-view image of the point cloud within its symmetric component as an additional input to the network.
[0081] Step 3-2: Train the network N c , the network uses the original porous three-view details D p As input, it predicts and outputs complete solid point cloud details to complete the missing areas.
[0082] The symmetry information between the components in step 3-1 is included in the general template structure definition corresponding to the category. After the symmetrical detail diagram is obtained, it will be concatenated with the original detail diagram as the network input. Figure 5 Taking a chair as an example, after obtaining the multi-hole three-view images of the backrest, seat and four legs according to the structural information, the three-view images of the symmetrical components are obtained and merged as the network N according to the symmetry information of the structure. c Input.
[0083] The step 3-2 requires the complete point cloud to be input into the network N described in step 2 f The solid three-view detail D obtained in s , and use it as the training network N c Tags. Network N c Learning to recover complete details from porous details of residual clouds c .
[0084] The network N c A data augmentation strategy is used during training; the data augmentation strategy is the same as the data augmentation strategy in step 1. This increases the robustness of the network. c The process uses the pre-trained N s and N f The training pseudo code of the network is as follows:
[0085]
[0086] like Figure 6 As shown, step 4 clearly identifies the missing area by comparing the detail map, providing a basis for preserving the input shape. The specific process can be divided into the following sub-steps:
[0087] Step 4-1: Comparatively complete solid detail D c and incomplete solid details D s , identify the overlapping area as the original input area A o .
[0088] Step 4-2: For each part P, complete the three views Symmetrical mirroring is performed on three orthogonal planes to obtain the solid three-view image of the mirrored part. Check its complete three views Whether it is a solid three-view drawing with its mirror image part Overlap, calculate the overlap rate of the three views separately If one of the overlap ratios r of the three views does not exceed the threshold T 1 , then the component is not a symmetrical component and its missing area needs to be diversified.
[0089] The mirroring refers to symmetrically mirroring the three views according to three orthogonal planes. For example, if you need to obtain the mirrored three views about the xy plane, you only need to mirror the xy view vertically and mirror the yz and xz views horizontally.
[0090] Step 4-3: For each symmetrical part P, obtain the symmetry relationship of the parts from the complete structure of the object. If part P has a symmetrical component, merge the solid three-view image of its mirrored part and the solid three-view image of its symmetrical component as the area to be inspected, and then define the overlapping area of the solid three-view image of part P and the area to be inspected as the associated input area A. r ; Component details d c and d s The overlapping area is taken as the original input area A o . Calculation area A o ∪A r The proportion of complete part details exceeds the threshold T 2 , it means that the missing region can be directly filled by symmetric information without the need for diversity generation; otherwise, it means that the missing region needs diversity generation.
[0091] The threshold T in step 4-2 1 The value of is 0.8, and the threshold T in step 4-3 2 The value of is 0.6.
[0092] like Figure 7 As shown, step 5 is based on the diffusion model N of the graph. g Realize diversity completion, which is the diffusion model N of graph-generated graphs g Image information is injected, and the noisy source image is used as initialization. The source image is a complete detail image masked by a random mask; during the denoising process, the diffusion model learns to recover the complete detail image without the mask from the noise. The randomness and data complexity of this process enable the model to capture more details during the generation process, thereby enhancing the diversity and accuracy of the output. Training N g The pseudo code of the network is as follows:
[0093]
[0094] like Figure 8 As shown, step 6 recovers the three-dimensional facet from the complete three-view image. A point in three-dimensional space can be uniquely determined based on the rays of the three orthogonal plane projections. The distance between its projection point on the three views and the boundary is calculated, and the minimum value is selected as the spatial distance. If a point is uniquely determined by a point in the figure enclosed by the contour in all three projection views, it means that it is inside the closed three-dimensional figure and the value is negative; if a point is uniquely determined by a point on the contour in all three projection views, it means that it is exactly on the surface of the closed three-dimensional figure and the value is 0; otherwise, it is outside the closed three-dimensional figure and the value is positive. Based on this, the corresponding signed distance field representation is obtained, and then the watertight facet is recovered according to the marching cube algorithm.
[0095] The pseudo code for getting the signed distance field is as follows:
[0096]
[0097] Fig. 9 The figure shows a comparison between the proposed method (Ours) and the cGAN, SDFusion, and ShapeFormer methods on the ShapNet dataset. The figure shows the input, label GT, and three diverse completion results of each method.
[0098] Fig.10 The comparison between the proposed method (Ours) and the existing techniques (cGAN, SDFusion and ShapeFormer) on real scan datasets Scannet, LiDAR-Net and scans from consumer-grade devices is shown. The first three rows show the diversity completion results of the swivel chair extracted from ScanNet. The bottom three rows show the completion results of the chair extracted from ScanNet, LiDAR-Net and scans using consumer-grade devices.
[0099] The present invention solves the problem that the current methods cannot obtain high-quality output results under the premise of retaining the input shape in the field of three-dimensional completion. After structural analysis and decomposition of the components, the method performs diversified generation and completion of the three-view images of the incomplete input components. By completing the three-view images of each component of the incomplete three-dimensional object, the method realizes the conversion from the complete three-view image to the signed distance field and then to the watertight triangular facet, thus realizing diversified completion of the three-dimensional object.
[0100] The present invention provides a method and idea for a diversified shape completion method based on component analysis and decomposition. There are many methods and approaches to implement the technical solution. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies.
Claims
1. A diverse shape completion method based on component analysis and decomposition, characterized in that: The following steps are involved: Step 1: Represent the object as an abstract template structure. The object is incomplete input data represented by a 3D point cloud. s Predicting template structural parameters that match the residual defect cloud, and deriving a complete structural representation of the object based on the template structural parameters; Step 2: Based on the complete structural representation of the object obtained in step 1, the object is decomposed into components, the projection details of the internal points of each component on three orthogonal planes are obtained, and the projection details are input into the network to generate the solid detail information of each component; Step 3: Complete the three-view image, combine the detailed information of the symmetrical parts, and deduce and complete the complete component details of the object; Step 4: By comparing the incomplete and complete three-view images obtained in step 3, the missing area is distinguished and identified, and whether diversity generation is required is determined according to the missing area; Step 5: Mask out the area that needs diversity generation in the complete detail image, perform image-to-image diffusion generation, and obtain diversified complete detail three-view images; Step 6: Reconstruct the 3D region based on the three views, obtain the corresponding signed distance field, and obtain a watertight completion result through the marching cube algorithm.
2. The method for diverse shape completion based on component analysis and decomposition according to claim 1, characterized in that: Step 1: The abstract template structure is a general structure designed for a three-dimensional object. The general structure describes the basic component structure relationship and the variation range of component parameters of the object. The structure is composed of M interrelated primitive boxes, and the primitive boxes correspond to the components in the object respectively. The original box is represented by 8 parameters, including three coordinate parameters of the starting point, three coordinate parameters of the ending point, and two dimension parameters representing the length and width of the original box; The association relationship between the boxes is controlled by the connection constraint relationship between the key points. There are 26 key points in total, corresponding to the center points of the 6 faces, 8 vertices and the midpoints of its 12 edges on the unit cube; the association relationship includes the connection relationship between the components and the semantic and symmetry information of the components.
3. The method for diverse shape completion based on component analysis and decomposition according to claim 2, characterized in that: Step 1: The deep network N s Taking the three-channel coordinates of the input point cloud as input features, the network outputs n values after the corresponding encoding and decoding process, which correspond to the parameters of its structural template. The general structural template is adjusted through the output parameters to make the template structure more suitable for the input point cloud, thereby obtaining the complete structure of the input point cloud; The deep network N s The data augmentation strategy is used in the training scheme of , which is to randomly rotate the complete point cloud in the dataset and remove some points according to the z-axis segmentation to obtain the input residual point cloud.
4. The method for diverse shape completion based on component analysis and decomposition according to claim 3, characterized in that: Step 2 includes the following steps: Step 2-1: According to the position and proportion of each box obtained in step 1, the input point cloud is segmented into components, and the projection of each component point cloud on three orthogonal planes is obtained, and the projection is used as the original multi-hole three-view image, which describes the structural details inside the input area; Step 2-2: Use the component projection image obtained in step 2-1 as input to train the network N f , through the network N f The solid details D of each part are obtained by inferring the porous three-view detail predictions. s .
5. The method for diverse shape completion based on component analysis and decomposition according to claim 4, characterized in that: Step 2-2 The prediction process uses a data enhancement strategy, which is to divide the image by a random straight line, randomly select the area above or below the straight line as a mask, and apply the mask to the porous detail map obtained by the projection of the complete point cloud and the solid three-view image of the white belt of the dataset, respectively. and As a network f The data pairs required for training, during the training process, the network N f Learning by recover 6. The method for diverse shape completion based on component analysis and decomposition according to claim 5, characterized in that: Step 3 includes the following steps: Step 3-1: For each box, obtain the projected three-view of the point cloud in its symmetric component as the network N c The symmetry information between the components in step 3-1 is included in the definition of the general template structure. After the symmetric detail map is obtained, it is concatenated with the original detail map as the network input; Step 3-2: Train the network N c , the network N c Taking the original porous three-view details of the part as input, predict and output the complete solid details D c , fill in the missing areas; The step 3-2 requires the complete point cloud to be input into the network N described in step 2 f The solid detail map obtained in the training network N c Tags, Network c Learn to recover full details from porous details in residual clouds; The network N c A data augmentation strategy is used during training; the data augmentation strategy is the same as the data augmentation strategy in step 1.
7. The method for diverse shape completion based on component analysis and decomposition according to claim 6, characterized in that: Step 4 includes the following steps: Step 4-1: Comparatively complete solid detail D c and incomplete solid details D s , identify the overlapping area as the original input area A o ; Step 4-2: For each part P, complete the three views Symmetrical mirroring is performed according to three orthogonal planes to obtain the solid three-view image of the mirrored part Check its complete three views Whether it is a solid three-view drawing with its mirror image part Overlap, calculate the overlap rate of the three views separately If one of the overlap rates r of the three views does not exceed the threshold T1, the component is not a symmetrical component and its missing area needs to be diversified. Step 4-3: For each symmetrical part P, obtain the symmetry relationship of the parts from the complete structure of the object. If part P has a symmetrical component, merge the solid three-view image of its mirrored part and the solid three-view image of its symmetrical component as the area to be inspected, and define the overlapping area of the solid three-view image of part P and the area to be inspected as the associated input area A. r ; Component details d c and d s The overlapping area is taken as the original input area A o , calculate area A o ∪A r If the proportion of complete part details exceeds the threshold T2, the missing area is directly completed by the symmetry information; otherwise, the missing area needs to be generated in a diverse manner.
8. The method for diverse shape completion based on component analysis and decomposition according to claim 7, characterized in that: The value range of the threshold value T1 in step 4-2 is 0.7-0.9, and the value range of the threshold value T2 in step 4-3 is 0.5-0.
7.
9. The method for diverse shape completion based on component analysis and decomposition according to claim 8, characterized in that: Step 5: Diffusion model based on graph N g Realize diversity completion, which is the diffusion model N of graph-generated graphs g Image information is injected and the noisy source image is used as initialization. The source image is a complete detail image masked by a random mask. In the denoising process, the diffusion model learns to recover the complete detail image without mask from the noise. The network achieves diversity generation through the randomness of the denoising process and the complexity of high-dimensional data.
10. The method for diverse shape completion based on component analysis and decomposition according to claim 9, characterized in that: Step 6 The following steps are included: Step 6-1: Obtain the signed distance field, reconstruct the 3D area based on the three-view image, and for each 3D query point, select the minimum value of the distance between the projection of the query point on the three-view image and the boundary as the signed distance field of the query point. The shape and position of the object in space can be determined through the signed distance field; Step 6-2: Extract isosurfaces using the signed distance field obtained in step 6-1, and restore watertight three-dimensional surfaces using a marching cubes algorithm based on the isosurfaces.
Citation Information
Patent Citations
Three-dimensional point cloud completion method based on deep learning and voxels
CN112927359A
Three-dimensional model completion and restoration method and system based on point cloud vectorization skeleton matching
CN115953563A
Three-dimensional model data processing method and device, equipment, medium and product
CN119206006A
Dataset generation method for self-supervised learning scene point cloud completion based on panoramas
US20230094308A1