3D point cloud defect detection method based on diffusion model

Through the methods of hierarchical graph coding and diffusion denoising, the shortcomings of existing 3D point cloud defect detection in complex geometric structure reconstruction and multi-scale feature fusion are solved, and efficient and accurate point cloud defect detection is achieved.

CN120411064APending Publication Date: 2025-08-01ZHEJIANG UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510596023.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing 3D point cloud defect detection method based on generative models lacks the ability to reconstruct complex geometric structures, insufficient capture of multi-scale anomaly features, and the post-processing steps cannot be optimized end-to-end, resulting in high error detection rates and low computational efficiency.

Method used

The method of hierarchical graph coding and diffusion denoising is adopted to extract multi-level geometric features through hierarchical graph convolution encoder, and high-fidelity point cloud reconstruction is used to use a diffusion model driven by feature conditions, and combined with multi-scale anomaly feature fusion strategy to achieve end-to-end optimization.

Benefits of technology

It improves the accuracy of point cloud defect detection, can capture microscopic details and macroscopic structural defects at the same time, reduces the error detection rate and improves computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411064A_ABST
    Figure CN120411064A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D point cloud defect detection method based on a diffusion model. The method comprises the following steps: 1) preprocessing an original point cloud; 2) extracting local to global geometric features of the preprocessed point cloud step by step; 3) inputting the extracted features into a diffusion model denoising module, and reconstructing point clouds by the model depending on geometric features; 4) extracting multi-level features of the reconstructed point cloud and the original point cloud; and 5) calculating the difference of each level, and fusing to obtain an exception score and an exception graph. According to the method, the microdefect and deformation area can be accurately positioned in a complex industrial scene, and meanwhile, the calculation burden introduced by a traditional post-processing algorithm is avoided; and the point cloud anomaly detection accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of industrial defect detection, and relates to a 3D point cloud defect detection method based on a diffusion model, mainly used to distinguish whether point cloud data contains defects. Background Art

[0002] As a digital expression form of the geometric information of the object surface, 3D point cloud has important application value in the fields of industrial quality inspection, autonomous driving, reverse engineering, etc. In the industrial manufacturing scenario, high-precision 3D sensors can quickly obtain the surface point cloud data of parts, and local defects in the point cloud are often the key factors affecting product quality. Traditional defect detection methods usually rely on geometric features designed manually or rule-based segmentation algorithms, but such methods have poor generalization ability for complex defects and are difficult to cope with problems such as point cloud data sparsity, noise interference, and non-uniform sampling.

[0003] In recent years, point cloud analysis methods based on deep learning have gradually become the mainstream. Among them, generative models are used to learn the data distribution of normal point clouds, and abnormal regions are located through reconstruction errors. However, such methods have significant limitations: on the one hand, traditional generative models have insufficient ability to reconstruct point cloud details, especially when dealing with complex geometric structures, problems such as over-smoothing or topological distortion are likely to occur, resulting in insignificant differences in reconstruction errors between defect regions and normal regions; on the other hand, existing methods usually adopt a single-scale feature comparison strategy, making it difficult to capture the abnormal feature differences of multi-level structures in point clouds. In addition, the irregularity and disorder of point clouds pose challenges to the feature extraction process. Conventional convolutional neural networks are difficult to effectively model the spatial relationship of point clouds, while graph neural networks can handle non-Euclidean data, but their shallow feature fusion mechanism may ignore the complementarity of different levels of semantic information.

[0004] Existing research attempts to optimize the detection effect through a two-stage method: first, a generative model reconstructs the point cloud, and then a rule-based geometric analysis algorithm is used to locate abnormalities. However, the defect location module of such methods is independently optimized with the generative model, and feature alignment cannot be achieved through end-to-end training, and additional post-processing steps significantly increase the computational complexity. For example, the method based on autoencoders needs to rely on threshold segmentation to distinguish abnormalities, but its reconstruction error distribution is vulnerable to noise interference, resulting in an increase in the false detection rate; the path optimization method based on graph search can capture structural continuity but cannot adapt to the unstructured characteristics of point clouds. Therefore, how to achieve multi-scale feature fusion and end-to-end optimization in the defect detection process while ensuring the generation accuracy has become the key challenge to improve the robustness and efficiency of 3D point cloud defect detection. Summary of the Invention

[0005] To overcome the problems of the existing 3D point cloud defect detection methods based on generative models, such as insufficient ability to reconstruct complex geometric structures, inadequate capture of multi-scale abnormal features, and high false detection rate and low computational efficiency caused by the inability to optimize the post-processing steps end-to-end, the present invention proposes a 3D point cloud defect detection method based on hierarchical graph encoding and diffusion denoising. This method extracts multi-level geometric features through a hierarchical graph convolutional encoder, and uses a feature-conditioned diffusion model to achieve high-fidelity point cloud reconstruction. Combining with a multi-scale abnormal feature fusion strategy, it can accurately locate microscopic defects and deformation regions in complex industrial scenarios, while avoiding the computational burden introduced by traditional post-processing algorithms.

[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0007] A 3D point cloud defect detection method based on a diffusion model, comprising the following steps:

[0008] 1) Preprocess the original point cloud data;

[0009] 2) Input the preprocessed point cloud into a hierarchical point cloud graph encoder module to extract local-to-global geometric features layer by layer;

[0010] 3) Input the features extracted by the encoder module into a diffusion denoising module to generate a reconstructed point cloud through noise prediction constrained by feature conditions;

[0011] 4) Perform point cloud registration on the reconstructed point cloud and the original point cloud, input them into an anomaly detection module, and perform multi-scale feature extraction;

[0012] 5) The anomaly detection module inputs the reconstructed point cloud and the original point cloud into a hierarchical point cloud graph encoder, fuses and compares the features of different layers, and uses the Euclidean distance between the features after fusing the original point cloud and the reconstructed point cloud to identify and locate the abnormal region.

[0013] Further, in the step 1), the preprocessing process is: performing farthest point sampling on the point cloud data.

[0014] Still further, in the step 2), the hierarchical feature extraction process is: constructing a dynamic local neighborhood for each point based on the K-nearest neighbor algorithm to encode the local geometric structure; the first level takes the coordinate difference as the input, generates node features through multi-layer perceptron mapping and max pooling; the second level fuses the node feature differences to further extract high-order geometric patterns; perform global max pooling on the node features, and compress them into a compact latent encoding through a fully connected layer.

[0015] Furthermore, in the step 3), the processing process of the diffusion denoising module is as follows: Concatenate the time step encoding with the shape features, dynamically adjust the noise prediction weights through a gating mechanism, and inject a bias term to guide the optimization of the local point density; Generate channel attention weights using the encoder features to suppress the noise prediction directions that deviate from the normal geometric patterns.

[0016] Furthermore, in the step 4), the process of registering the reconstructed point cloud with the original point cloud includes the following sub-steps:

[0017] 41) Align the reconstructed point cloud with the original point cloud spatially through the iterative closest point algorithm, and unify the feature scales of the point clouds through global spatial normalization;

[0018] 42) Input the aligned reconstructed point cloud and the original point cloud into a hierarchical point cloud map feature encoder for multi-scale feature extraction.

[0019] Furthermore, in the step 5), the processing process of the anomaly detection module includes the following sub-steps:

[0020] 51) Use the trained network to extract point-level features, region-level features, and graph-level features of the point cloud;

[0021] 52) Use the improved chamfer distance as the metric method for point-level features; use the attention weight cosine similarity as the metric method for region-level features; use the KL divergence as the metric method for graph-level features, and generate an anomaly score after weighted fusion;

[0022] 53) Adopt an adaptive threshold segmentation algorithm to generate a binary defect mask according to the anomaly score distribution.

[0023] Preferably, in the steps 2) and 3), the optimal parameters of the network model are obtained during training, and the network training method is: Backpropagation optimization is performed on the loss function value between the predicted noise and the real noise, and the mean square error loss sum is used to calculate the error between the predicted denoised data and the real data.

[0024] A 3D point cloud defect detection method based on a diffusion model proposed by the present invention is mainly characterized by the accurate defect detection of point cloud data. Due to the irregularity and disorder of point cloud data, it is difficult for ordinary convolutional neural networks to extract the features of point cloud data, and traditional point cloud reconstruction methods face the problem of being difficult to handle the partial occlusion sensitivity of point clouds. Therefore, the present invention proposes to use the features extracted by a graph neural network in combination with a diffusion model to achieve defect detection of point clouds. Since the diffusion model can perform undifferentiated masking on point cloud objects, it overcomes the limitation that the masked autoencoder cannot fully reconstruct point clouds when reconstructing point clouds. The feature learning obstacle brought by the disorder of point clouds is overcome through a dynamic graph construction mechanism, and the local geometric relationship in the non-Euclidean space is encoded into a stable topological representation, and then the features obtained by the graph neural network are used to guide the diffusion model to reconstruct the random and disordered point cloud into a normal sample. Furthermore, the present invention proposes a hierarchical point cloud graph encoder module, which constructs a dynamic local neighborhood graph for each point in the point cloud data through K-nearest neighbor search, providing a structural prior for subsequent graph convolution operations. Local geometric features are gradually extracted through two-level edge convolution layers. Specifically, first, the coordinate difference between the node coordinates and their neighbors is concatenated as the input feature, and after being mapped by a multi-layer perceptron, max pooling is used to aggregate the neighbor information to output node-level features; then, based on the output of the first layer, the node feature difference is concatenated to form the input, and high-order geometric features are further extracted through a multi-layer perceptron, and max pooling is used again for aggregation to obtain high-dimensional node features. This process gradually expands the receptive field by stacking graph convolution layers, capturing multi-level geometric patterns from local details to regional structures.

[0025] In the process of graph neural feature decoding, first, the time step encoding is fused with the shape feature to establish an explicit association between the noise level and the shape semantics. Then, the conditional gating and bias injection mechanism are used to dynamically adjust the sensitivity of the network to the shape feature. Finally, the shape semantics are injected into the density optimization process to realize denoising according to the graph neural feature-assisted diffusion model.

[0026] In the anomaly detection module, the graph neural features of the reconstructed point cloud and the original point cloud obtained by the previous hierarchical point cloud graph encoder are used to compare the anomaly scores between different hierarchical features. Specifically, the density-weighted chamfer distance is used as the measurement method for point-level features; the attention-weighted cosine similarity is used as the measurement method for region-level features; the KL divergence is used as the measurement method for graph-level features. Finally, these anomaly scores are weighted and averaged to obtain the anomaly score of the final point cloud.

[0027] The beneficial effects of the present invention are as follows: It overcomes the feature learning deviation brought by the disorder of point clouds, can capture both microscopic detail defects and macroscopic structure defects, and aims to improve the accuracy of point cloud anomaly detection. Description of the Drawings

[0028] Figure 1It is a flow chart of a 3D point cloud defect detection method based on a diffusion model.

[0029] Figure 2 It is a flow chart of a hierarchical point cloud graph encoder module.

[0030] Figure 3 It is a flow chart of denoising of a diffusion model. Specific implementation manners

[0031] The present invention will be further described below with reference to the flow chart.

[0032] Referring to Figures 1 to 3 , a 3D point cloud defect detection method based on a diffusion model includes the following steps:

[0033] 1) Preprocess the original point cloud data. The processing process is as follows: adjust the number of points in the point cloud data to a fixed size, such as 2048 points. In order to maintain the geometric features of the point cloud data, use the farthest point sampling method to sample the point cloud data, and then normalize the point cloud data;

[0034] 2) Input the preprocessed point cloud data into a hierarchical point cloud graph encoder. As Figure 2 shown, the hierarchical feature extraction process includes the following sub-steps:

[0035] 21) For each point, construct a dynamic local neighborhood graph through K-nearest neighbor search. The value of K is selected as 16. In the sparse point region, the value of K is adaptively reduced to 8, and in the dense point region, the value of K is increased to 32. Output the neighborhood relationship matrix [B, N, 16] (where B is the data batch size and N is the number of points in the point cloud) and the relative coordinate difference [B, N, 16, 3]. Finally, splice the central point coordinates and the neighborhood coordinate difference to generate an input tensor of [B, N, 16, 6];

[0036] 22) Use a two-layer multi-layer perceptron (6→64→64 dimensions), extract non-linear geometric patterns through a non-linear activation function to obtain features [B, N, 16, 64], and further perform max pooling to take the maximum value of the features along the neighborhood dimension (K = 16), and output point-level features with a shape of [B, N, 64];

[0037] Select the value of K as 64, use the K-nearest neighbor algorithm to construct a neighborhood graph of local features, output the high-order neighborhood relationship [B, N, 64] and the feature difference [B, N, 64, 64], and further splice the output features of the first layer and the neighborhood feature difference to generate an input of [B, N, 64, 128];

[0038] Obtain features [B, N, 64, 256] through a three-layer MLP (128→256→256→256 dimensions), and further perform attention pooling to weighted aggregate and output structural features with a shape of [B, N, 256];

[0039] 23) Perform max pooling on the structural features along the point dimension to obtain a global descriptor [B, 256]. Further, pass it through a two-layer fully connected network (256→512→128 dimensions), where after the first layer, non-linear activation and regularization are performed to generate graph-level features of [B, 128], fusing local-to-global geometric information;

[0040] 3) Input the latent encoding obtained by the hierarchical point cloud graph encoder into a diffusion model, as Figure 3 shown, and the specific process includes the following sub-steps:

[0041] 31) Concatenate the time step encoding with the shape features extracted by the encoder to ensure that each generation step is constrained by the features;

[0042] 32) Each layer of the denoising network receives the noisy point cloud, and uses a gating mechanism to generate weight coefficients to amplify the feature responses related to the target shape, and uses bias injection to add condition-related bias terms to the intermediate features of the network, implicitly injecting shape prior knowledge;

[0043] 33) The feature-conditioned driving network adjusts the noise prediction direction to guide the optimization of local point density and inputs the predicted noise;

[0044] 4) Register the reconstructed point cloud with the original point cloud. Specifically, the point cloud is spatially aligned through the iterative closest point algorithm, and the feature scales of the point cloud are unified through global spatial normalization. The process of registering the reconstructed point cloud with the original point cloud includes the following sub-steps:

[0045] 41) Spatially align the reconstructed point cloud with the original point cloud through the iterative closest point algorithm, and unify the feature scales of the point cloud through global spatial normalization;

[0046] 42) Input the aligned reconstructed point cloud and the original point cloud into the hierarchical point cloud graph feature encoder for multi-scale feature extraction;

[0047] 5) Input the registered reconstructed point cloud and the original point cloud into the hierarchical point cloud graph encoder to obtain the multi-scale features of the original point cloud and the reconstructed point cloud, and calculate the anomaly score between the features of the two point clouds. Specifically, when calculating the anomaly score, density-weighted chamfer distance is used as the metric for point-level features; attention-weighted cosine similarity is used as the metric for region-level features; KL divergence is used as the metric for graph-level features. Finally, these anomaly scores are weighted and averaged to obtain the anomaly score of the final point cloud. The anomaly detection process includes the following sub-steps:

[0048] 51) Use the trained network to extract point-level features, region-level features, and graph-level features of the point cloud;

[0049] 52) Use the improved chamfer distance as the metric method for point-level features; use the attention-weighted cosine similarity as the metric method for region-level features; use the KL divergence as the metric method for graph-level features, and generate an anomaly score after weighted fusion;

[0050] 53) Adopt an adaptive threshold segmentation algorithm to generate a binary defect mask according to the anomaly score distribution.

[0051] In this embodiment, the optimal network parameters are obtained during model training, and the network training method is: backpropagation optimization is performed on the loss function value between the predicted noise and the true noise. Specifically, the mean square error loss is used to calculate the error between the predicted denoised data and the true data.

[0052] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept and is only for illustrative purposes. The protection scope of the present invention should not be regarded as limited to the specific forms stated in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that can be conceived by those of ordinary skill in the art based on the inventive concept of the present invention.

Claims

1. A 3D point cloud defect detection method based on a diffusion model, characterized in that, The method includes the following steps: 1) Preprocess the original point cloud data; 2) Input the preprocessed point cloud data into the hierarchical point cloud graph encoder module to extract local-to-global geometric features layer by layer; 3) Input the features extracted by the encoder module into the diffusion denoising module to generate a reconstructed point cloud through noise prediction with feature conditional constraints; 4) Perform point cloud registration on the reconstructed point cloud and the original point cloud, input it into the anomaly detection module, and perform multi-scale feature extraction; 5) The anomaly detection module inputs the reconstructed point cloud and the original point cloud into the hierarchical point cloud graph encoder, fuses and compares the features of different layers, and uses the Euclidean distance between the features after fusing the original point cloud and the reconstructed point cloud to identify and locate the abnormal area.

2. The 3D point cloud defect detection method based on a diffusion model according to claim 1, wherein, In step 2), the hierarchical feature extraction process is as follows: Based on the K-nearest neighbor algorithm, construct a dynamic local neighborhood for each point to encode the local geometric structure; the first level takes the coordinate difference as the input, and generates node features through multi-layer perceptron mapping and max pooling; The second level fuses the node feature differences to further extract high-order geometric patterns; perform global max pooling on the node features and compress them into a compact latent encoding through a fully connected layer.

3. The 3D point cloud defect detection method based on the diffusion model according to claim 2, wherein, In step 2), the hierarchical feature extraction process includes the following sub-steps: 21) For each point, construct a dynamic local neighborhood graph through K-nearest neighbor search to provide a structural prior for subsequent graph convolution operations; 22) Gradually extract local geometric features through two-level edge convolution layers. First, concatenate the coordinate difference between the node coordinates and those of its neighbors as the input feature, after multi-layer perceptron mapping, use max pooling to aggregate neighbor information and output node-level features; then, use the K-nearest neighbor algorithm to construct a neighborhood graph of node features to obtain neighborhood relationships and feature differences, concatenate the output of the first layer and the node feature differences to form the input, further extract high-order geometric features through a multi-layer perceptron, and use attention pooling for weighted aggregation to obtain high-dimensional node features. This process gradually expands the receptive field by stacking graph convolution layers to capture multi-level geometric patterns from local details to regional structures. 23) To fuse the geometric information of the entire point cloud, perform global max pooling on the node features output by the second layer: Divide the nodes along the batch dimension, take the channel maximum value of the high-dimensional features of all points within each batch to obtain a global descriptor; then enhance the non-linear expression ability through a fully connected layer to generate the final latent encoding.

4. A 3D point cloud defect detection method based on a diffusion model according to any one of claims 1 to 3, characterized in that, In step 3), the processing process of the diffusion denoising module is as follows: Concatenate the time step encoding and the shape feature, dynamically adjust the noise prediction weight through a gating mechanism, and inject a bias term to guide the optimization of local point density; generate channel attention weights using the encoder features to suppress the noise prediction direction that deviates from the normal geometric pattern.

5. The 3D point cloud defect detection method based on a diffusion model according to claim 4, characterized in that, In step 3), the processing process of the diffusion denoising module includes the following sub-steps: 31) Concatenate the time step encoding and the shape feature extracted by the encoder to ensure that each step of generation is feature-constrained; 32) Each layer of the denoising network receives the noisy point cloud, and uses the gating mechanism to generate weight coefficients to amplify the feature responses related to the target shape, and uses bias injection to add condition-related bias terms to the intermediate features of the network, implicitly injecting shape prior knowledge; 33) The feature condition drives the network to adjust the noise prediction direction, guides the optimization of local point density, and generates the denoised point cloud.

6. A 3D point cloud defect detection method based on a diffusion model according to any one of claims 1 to 3, characterized in that, In the step 4), the process of registering the reconstructed point cloud with the original point cloud includes the following sub-steps: 41) Align the reconstructed point cloud with the original point cloud in space through the iterative closest point algorithm, and unify the feature scale of the point cloud through global spatial normalization; 42) Input the aligned reconstructed point cloud and the original point cloud into the hierarchical point cloud map feature encoder for multi-scale feature extraction.

7. A 3D point cloud defect detection method based on a diffusion model according to any one of claims 1 to 3, characterized in that, In the step 5), the processing process of the anomaly detection module includes the following sub-steps: 51) Use the trained network to extract the point-level features, region-level features, and graph-level features of the point cloud; 52) Use the improved chamfer distance as the measurement method for the point-level features; use the attention weight cosine similarity as the measurement method for the region-level features; use the KL divergence as the measurement method for the graph-level features, and generate an anomaly score after weighted fusion; 53) Adopt an adaptive threshold segmentation algorithm to generate a binary defect mask according to the anomaly score distribution.

8. A 3D point cloud defect detection method based on a diffusion model according to any one of claims 1 to 3, characterized in that, In the 2) and 3), the optimal parameters of the network model are obtained during training, and the network training method is: backpropagation optimization is performed on the loss function value between the predicted noise and the real noise, and the mean square error loss sum is used to calculate the error between the predicted denoised data and the real data.

Citation Information

Cited By

  • District construction quality defect identification and positioning method based on multivariate data fusion

    CN120781304A

  • Method for identifying and positioning quality defects of transformer area construction based on multi-element data fusion

    CN120781304B

  • Packing box anomaly detection method and device based on semantic constraint and geometric prior

    CN122089714A

  • Method and device for detecting abnormal packaging box based on semantic constraint and geometric prior

    CN122089714B