Unsupervised surface anomaly detection method and system based on diffusion probability model

By adopting an unsupervised method based on the diffusion probability model in surface anomaly detection, anomaly pattern generation module, reconstruction subnetwork and discriminant subnetwork are built, which solves the problems of insufficient data and noise sensitivity in traditional methods, and realizes efficient anomaly detection and robust model.

CN119963492APending Publication Date: 2025-05-09SHANGHAI SPACE PRECISION MACHINERY RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510003393.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Traditional surface anomaly detection methods are difficult to obtain sufficient data for modeling in industrial environments, lack of modeling capabilities, are sensitive to noise and complexity, and rely on the collection of abnormal samples.

Method used

Unsupervised surface anomaly detection method based on diffusion probability model is adopted, and anomaly detection of image data is carried out by constructing anomaly mode generation module, reconstructing sub-network and discriminant sub-network, and diffusion probability model and noise embedding technology are used.

Benefits of technology

It realizes surface anomaly detection under the unsupervised paradigm, reduces dependence on abnormal samples, improves the robustness and reconstruction capabilities of the model, and enhances the ability to identify severely deformed samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963492A_ABST
    Figure CN119963492A_ABST
Patent Text Reader

Abstract

The invention provides an unsupervised surface anomaly detection method and system based on a diffusion probability model. The unsupervised surface anomaly detection method comprises the following steps that S1, an industrial camera collects non-anomaly image data; s2, making an anomaly detection data set; s3, constructing a surface anomaly detection network based on a diffusion probability model; s4, training the network; and S5, detecting the predicted image, and outputting a detection result. According to the method, the limitation of a supervised method on exception definition can be overcome, and accurate positioning of surface exceptions is realized by fully utilizing non-exception samples which are easy to obtain in an industrial scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an unsupervised surface anomaly detection method and system based on a diffusion probability model. Background Art

[0002] With the continuous improvement of the degree of industrial intelligence, the importance of anomaly detection technology in industrial production has become increasingly prominent, especially in the fields of product quality inspection, safety monitoring, traffic detection, medical image processing, etc. Surface anomaly detection research mainly focuses on discovering and locating abnormal patterns in image data. Traditional methods mainly rely on artificial feature design and use statistical or rule-based models. However, there are some limitations in using it in industrial environments. It is difficult to obtain enough data for modeling, the model capacity is insufficient, and it is sensitive to noise and complexity. Some engineering practices use target detection-based methods to complete the location and classification of image anomalies, but they also face the problem of collecting abnormal samples, which requires a lot of resources, and the lack of training data is still a challenge. Summary of the invention

[0003] In view of the defects in the prior art, the object of the present invention is to provide an unsupervised surface anomaly detection method and system based on a diffusion probability model.

[0004] The unsupervised surface anomaly detection method based on the diffusion probability model provided by the present invention includes:

[0005] Step S1: collecting image data;

[0006] Step S2: Based on the collected image data, an anomaly detection data set is generated;

[0007] Step S3: constructing an unsupervised surface anomaly detection network based on a diffusion probability model;

[0008] Step S4: training an unsupervised surface anomaly detection network based on a diffusion probability model using anomaly detection dataset;

[0009] Step S5: Perform anomaly detection on the image to be inspected through the trained unsupervised surface anomaly detection network based on the diffusion probability model to obtain the detection result of the anomaly in the image, that is, the mask of the anomaly location.

[0010] Preferably, the step S1 comprises:

[0011] Step S101: using an industrial camera to collect image data of a scene to be inspected;

[0012] Step S102: classifying the image into abnormal patterns according to the preset abnormal pattern features in the image;

[0013] The step S2 comprises:

[0014] Step S201: According to the abnormal pattern classification, the abnormal detection data set is divided into a training set and a test set, wherein the training set contains all samples without abnormalities, and the test set sets different subfolders to respectively contain samples with different abnormal patterns;

[0015] Step S202: Annotate samples with abnormal patterns in the test set, and place the annotation files in subfolders of each abnormal pattern under the annotation folder according to the test set folder settings.

[0016] Preferably, step S3 comprises:

[0017] Step S301: construct an abnormal pattern generation module, and use the abnormal pattern generated by the deviation distribution to identify the abnormality; the noise image is generated by using the Perlin noise generator, and binarized by the threshold of uniform random sampling to generate an abnormal mask; the abnormal texture source image is sampled from an abnormal source image data set that is independent of the input image distribution, and then random enhancement sampling, tone separation, sharpening, equalization, brightness change, color change and contrast enhancement are performed on the abnormal source image data set; the enhanced texture image A is masked with the abnormal mask and compared with the training image I a Mixing, that is, generating abnormal patterns that deviate from the distribution; training image I a Defined as:

[0018]

[0019] Among them, I is the sample without abnormality in the data set, The mask M a The flip mask of , A is the enhanced texture image, and β is the opacity parameter;

[0020] Step S302: construct a reconstruction sub-network, where the reconstruction sub-network consists of a diffusion probability model, an encoder, and a decoder;

[0021] Among them, using the diffusion probability model, the training of the reconstruction subnetwork is transformed from the image pixel space to the low-dimensional potential representation space with the same perception; for the image in the image space, the encoder is used to learn the pixel-level reconstruction of the normal sample, the image is encoded into a potential space vector, and the diffusion process is performed, and then the diffusion model based on UNet is used to learn the semantic feature generation, and the potential space vector with noise added after the diffusion process is denoised. The training objective loss function L LDM It is expressed as:

[0022]

[0023] Among them, z and ∈ are noises that conform to the normal distribution, N(0,1) is the normal distribution, t is the number of time steps of random sampling, z t is the noise at time step t, ∈ θ The model UNet performs denoising operations, and E is the expectation under various conditions;

[0024] The sampling process is carried out in the latent space, and the reconstructed sample vector is sampled from the learned distribution and decoded by the decoder into a reconstructed sample image in the image space;

[0025] Step S303: construct a discriminant subnetwork, the input of which is the intermediate state image of the image space output by the reconstruction subnetwork, the reconstructed image and the abnormal sample synthesized by the normal image through the abnormal pattern generation module. The output of the network is the abnormal position mask, which is used to calculate the loss with the abnormal pattern mask used in the abnormal generation module to guide the update of the network.

[0026] Preferably, in step S3, an unsupervised surface anomaly detection network based on a diffusion model is constructed, using noise embedding and channel interpolation;

[0027] In the unsupervised scenario of anomaly detection, noise embedding is used to guide sample generation;

[0028] Given a normal sample x, a synthetic abnormal sample x is obtained through the abnormal pattern generation module a , the synthetic abnormal sample is encoded into a latent vector by the encoder and regarded as the initial state c 0 , and then after T times of iterative diffusion process, the variance of the T times of added noise is β 1 …β T , then:

[0029]

[0030]

[0031] Where t is the number of random sampling time steps, c 0 ,c 1 ,c t-1 ,c t are the potential vectors of 0, 1, t-1, and t time steps respectively, and q is the conditional probability, that is, given c 0 Under these conditions, c t The probability of

[0032] Select a random time step t and get the diffusion-processed noise vector c noisy , put c noisy As a control condition, the diffusion probability model learns the target loss function L LDM becomes:

[0033]

[0034] Using interpolation channels, the discriminant subnetwork has recognition diversity during the reconstruction process and can distinguish real anomalies;

[0035] The latent vector c and reconstructed sample z of the abnormal input image after being encoded by the encoder r Interpolate to get the intermediate state, the expression is: z inter =λc+(1-λ)z r , where λ is a weight parameter between 0 and 1.

[0036] Preferably, the step S4 comprises:

[0037] Step S401: the optimizer uses momentum-based stochastic gradient descent Adam, the learning rate decay strategy uses linear decay, the encoder and decoder in the reconstruction sub-network use autoencoder models, and the network regularization method uses divergence regularization;

[0038] Step S402: Select normal samples from the training set of the prepared anomaly detection data set, pass through the abnormal pattern generation module, synthesize corresponding abnormal samples, send the normal samples and abnormal samples to the encoder of the reconstruction subnetwork to be encoded as potential vectors, and after diffusion processing, obtain the noisy normal sample potential vector and the abnormal sample potential vector, send them to the denoising UNet network for denoising training, and after denoising for a randomly selected number of time steps, output the potential vector of the reconstructed sample, perform channel interpolation on the potential vector of the input abnormal sample and the potential vector of the reconstructed sample to obtain the potential vector of the intermediate state, send them to the decoder to obtain the reconstructed image and the intermediate state image of the image space, and send them together with the abnormal samples as input to the discriminant subnetwork, output the abnormal mask, calculate the loss with the abnormal pattern mask used in the abnormal generation module, calculate the gradient, and then guide the discriminant subnetwork and the reconstruction subnetwork parameter update to complete the network training.

[0039] The unsupervised surface anomaly detection system based on the diffusion probability model provided by the present invention comprises:

[0040] Module M1: collect image data;

[0041] Module M2: Produce anomaly detection dataset based on collected image data;

[0042] Module M3: Constructing an unsupervised surface anomaly detection network based on diffusion probability model;

[0043] Module M4: Train an unsupervised surface anomaly detection network based on a diffusion probability model using anomaly detection datasets;

[0044] Module M5: Perform anomaly detection on the image to be inspected through the trained unsupervised surface anomaly detection network based on the diffusion probability model to obtain the detection result of the anomaly in the image, that is, the mask of the anomaly location.

[0045] Preferably, the module M1 comprises:

[0046] Module M101: Use industrial cameras to collect image data of the scene to be inspected;

[0047] Module M102: classifying the image into abnormal patterns according to the preset abnormal pattern features in the image;

[0048] The module M2 comprises:

[0049] Module M201: According to the abnormal pattern classification, the anomaly detection data set is divided into a training set and a test set, where the training set contains all samples without abnormalities, and the test set sets different subfolders to place samples with different abnormal patterns;

[0050] Module M202: Label the samples with abnormal patterns in the test set, and place the labeling files in the subfolders of each abnormal pattern under the labeling folder according to the test set folder settings.

[0051] Preferably, the module M3 comprises:

[0052] Module M301: Construct an abnormal pattern generation module, and use the abnormal pattern generated by the deviation distribution to identify the abnormality; the noise image is generated by using the Perlin noise generator, and is binarized by the threshold of uniform random sampling to generate an abnormal mask; the abnormal texture source image is sampled from an abnormal source image dataset that is independent of the input image distribution, and then random enhancement sampling, tone separation, sharpening, equalization, brightness change, color change and contrast enhancement are performed on the abnormal source image dataset; the enhanced texture image A is masked with the abnormal mask and compared with the training image I a Mixing, that is, generating abnormal patterns that deviate from the distribution; training image I a Defined as:

[0053]

[0054] Among them, I is the sample without abnormality in the data set, The mask M a The flip mask of , A is the enhanced texture image, and β is the opacity parameter;

[0055] Module M302: construct a reconstruction sub-network, which consists of a diffusion probability model, an encoder, and a decoder;

[0056] Among them, using the diffusion probability model, the training of the reconstruction subnetwork is transformed from the image pixel space to the low-dimensional potential representation space with the same perception; for the image in the image space, the encoder is used to learn the pixel-level reconstruction of the normal sample, the image is encoded into a potential space vector, and the diffusion process is performed, and then the diffusion model based on UNet is used to learn the semantic feature generation, and the potential space vector with noise added after the diffusion process is denoised. The training objective loss function L LDM It is expressed as:

[0057]

[0058] Among them, z and ∈ are noises that conform to the normal distribution, N(0,1) is the normal distribution, t is the number of time steps of random sampling, z t is the noise at time step t, ∈ θ is the model UNet that performs the denoising operation, and E is the expectation calculated under various conditions;

[0059] The sampling process is carried out in the latent space, and the reconstructed sample vector is sampled from the learned distribution and decoded by the decoder into a reconstructed sample image in the image space;

[0060] Module M303: Construct a discriminative subnetwork. The input of the discriminative subnetwork is the intermediate state image of the image space output by the reconstruction subnetwork, the reconstructed image and the abnormal sample synthesized by the normal image through the abnormal pattern generation module. The output of the network is the abnormal position mask, which is used to calculate the loss with the abnormal pattern mask used in the abnormal generation module to guide the update of the network.

[0061] Preferably, the module M3 constructs an unsupervised surface anomaly detection network based on a diffusion model, using noise embedding and channel interpolation;

[0062] In the unsupervised scenario of anomaly detection, noise embedding is used to guide sample generation;

[0063] Given a normal sample x, a synthetic abnormal sample x is obtained through the abnormal pattern generation module a , the synthetic abnormal sample is encoded into a latent vector by the encoder and regarded as the initial state c 0 , and then after T times of iterative diffusion process, the variance of the T times of added noise is β 1 …β T , then:

[0064]

[0065]

[0066] Where t is the number of random sampling time steps, c 0 ,c1 ,c t-1 ,c t are the potential vectors of 0, 1, t-1, and t time steps respectively, and q is the conditional probability, that is, given c 0 Under these conditions, c t probability;

[0067] Select a random time step t and get the diffusion-processed noise vector c noisy , put c noisy As a control condition, the diffusion probability model learns the target loss function L LDM becomes:

[0068]

[0069] Using interpolation channels, the discriminant subnetwork has recognition diversity during the reconstruction process and can distinguish real anomalies;

[0070] The latent vector c and reconstructed sample z of the abnormal input image after being encoded by the encoder r Interpolate to get the intermediate state, the expression is: z inter =λc+(1-λ)z r , where λ is a weight parameter between 0 and 1.

[0071] Preferably, the module M4 comprises:

[0072] Module M401: The optimizer uses momentum-based stochastic gradient descent Adam, the learning rate decay strategy uses linear decay, the encoder and decoder in the reconstruction sub-network use autoencoder models, and the network regularization method uses divergence regularization;

[0073] Module M402: Select normal samples from the training set of the prepared anomaly detection dataset, pass through the abnormal pattern generation module, synthesize the corresponding abnormal samples, send the normal samples and the abnormal samples to the encoder of the reconstruction sub-network to be encoded as latent vectors, and after diffusion processing, obtain the noisy normal sample latent vector and the abnormal sample latent vector, which are sent to the denoising UNet network for denoising training. After denoising for a randomly selected number of time steps, output the latent vector of the reconstructed sample, perform channel interpolation on the latent vector of the input abnormal sample and the latent vector of the reconstructed sample to obtain the latent vector of the intermediate state, send it to the decoder to obtain the reconstructed image and the intermediate state image in the image space, and send them together with the abnormal samples as input to the discriminant sub-network, output the abnormal mask, calculate the loss with the abnormal pattern mask used in the abnormal generation module, calculate the gradient, and then guide the update of the parameters of the discriminant sub-network and the reconstruction sub-network to complete the network training.

[0074] Compared with the prior art, the present invention has the following beneficial effects:

[0075] (1) The present invention adopts an unsupervised paradigm, and there are no abnormal samples in the training data usage scenario, which solves the limitations of the traditional supervised paradigm in defining abnormalities and the problem that abnormal samples are difficult to obtain and insufficient in number;

[0076] (2) The present invention adopts a reconstruction sub-network based on a diffusion probability model to perform reconstruction operations in the latent space, reduce network resource consumption, and improve network reconstruction speed. At the same time, it uses noise condition embedding and channel interpolation methods to improve the reconstruction ability of the reconstruction sub-network and enhance the robustness of the model for severely deformed samples.

[0077] (3) The present invention uses the intermediate image reconstructed by channel interpolation as input to the discriminant subnetwork together with the abnormal image and the reconstructed image, so that the discriminant subnetwork can reduce the impact caused by non-abnormal differences and further improve the network's ability to locate abnormalities. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0079] Figure 1 It is a structural schematic diagram of the surface anomaly detection network based on the diffusion probability model of the present invention;

[0080] Figure 2a and Figure 2b is a schematic diagram of a random anomaly mask generated by an anomaly generation module of the present invention;

[0081] Figure 3 It is a structural diagram of a standard UNet network structure used by the reconstruction sub-network and the discrimination sub-network of the present invention;

[0082] Figure 4 This is a flow chart of the unsupervised surface anomaly detection method based on the diffusion probability model of the present invention. DETAILED DESCRIPTION

[0083] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0084] Example 1

[0085] Reference Figure 4 The present invention provides an unsupervised surface anomaly detection method based on a diffusion probability model, comprising the following steps:

[0086] Step S1, collecting image data;

[0087] S101, using an industrial camera to collect image data of the scene to be inspected;

[0088] S102, classifying the image according to the abnormal pattern features in the image; using the non-abnormal samples as training data and the abnormal samples as test sets to evaluate the model accuracy;

[0089] Step S2: creating anomaly detection data set;

[0090] S201. According to the abnormal pattern classification, the data set is divided into a training set and a test set, wherein the training set contains all samples without abnormalities, and the test set sets different subfolders to contain samples with different abnormal patterns, such as surface oil stains, surface scratches, surface damage, etc.;

[0091] S202: To evaluate the model, annotate samples with abnormal patterns in the test set, and place the annotated files in the subfolders of each abnormal pattern under the annotated folder according to the test set folder settings;

[0092] S203. The complete dataset format can refer to the Mvtec AD dataset; the training set of the dataset is all samples without abnormalities, and the test set is samples containing abnormal patterns in different subfolders;

[0093] Step S3, constructing an unsupervised surface anomaly detection network DiffusionADNet based on a diffusion probability model, which consists of an abnormal pattern generation module, a reconstruction subnetwork, and a discrimination subnetwork;

[0094] S301, constructing an abnormal pattern generation module, wherein the abnormal pattern does not need to use a real abnormal pattern, but can generate an abnormal pattern that deviates from the distribution, so that the abnormal pattern deviates from the normal distribution to identify the abnormality; the noise image is generated using a Perlin noise generator, and binarized by a uniformly randomly sampled threshold to generate an abnormal mask; sampling an abnormal texture source image from an abnormal source image data set that is independent of the input image distribution, and then performing random enhancement sampling, tone separation, sharpening, equalization, brightness change, color change, and contrast enhancement on the abnormal source image data set; masking the enhanced texture image A with an abnormal mask and mixing it with a training image I to generate an abnormal pattern that deviates from the distribution; training image I a It can be defined as:

[0095]

[0096] Among them, I is the sample without abnormality in the data set, M aThe flip mask of the mask, A is the enhanced texture image, β is the opacity parameter, which is reasonably set according to the characteristics of the abnormal pattern, and the default value is 0.2;

[0097] A set of random anomaly masks M a , as shown in Figure 2.

[0098] S302, constructing a reconstruction sub-network, where the reconstruction sub-network is based on a diffusion probability model and consists of a diffusion model, an encoder, and a decoder;

[0099] The reconstruction subnetwork uses the diffusion probability model technology. The detailed model structure can refer to the existing model latent diffusion. The training of the reconstruction subnetwork is transformed from the high-dimensional and expensive image pixel space to the low-dimensional latent representation space with the same perception. In order to adapt the diffusion probability model to the anomaly detection task, for the image in the image space, the encoder is used to learn the pixel-level reconstruction of normal samples, encode the image into a latent space vector, and perform diffusion processing. Then, the diffusion model based on UNet is used to learn semantic feature generation, and the latent space vector with noise added after diffusion processing is denoised. The training goal can be expressed as:

[0100]

[0101] Among them, z t and ∈ are noises that conform to the normal distribution, t is the number of random sampling steps, ranging from (1,1000), ∈ θ The model UNet that performs denoising operations;

[0102] The sampling process is also carried out in the latent space. The reconstructed sample vector is sampled from the learned distribution and decoded by the decoder into a reconstructed sample image in the image space.

[0103] S303, constructing a discriminant sub-network, where the discriminant sub-network is constructed based on the prior art UNet;

[0104] Among them, the specific structure of the discriminant subnetwork is basically consistent with the standard UNet, and other improved versions of UNet can also be used for specific implementation; the input of the network is the intermediate state image of the image space output by the reconstruction subnetwork, the reconstructed image and the abnormal sample synthesized by the normal image through the abnormal pattern generation module, and the output of the network is the abnormal position mask, which is used to calculate the loss with the abnormal mask used by the abnormal pattern generation module to guide the update of the network;

[0105] The specific structure of the unsupervised surface anomaly detection network DiffusionADNet based on the diffusion probability model is as follows Figure 1 As shown, a standard UNet network structure used by the reconstruction sub-network and the discrimination sub-network is as follows Figure 3 shown.

[0106] Step S4, training an unsupervised surface anomaly detection network DiffusionADNet based on a diffusion probability model;

[0107] S401, the optimizer uses momentum-based stochastic gradient descent Adam, the learning rate decay strategy uses linear decay linear, the encoder and decoder in the reconstruction sub-network use the existing autoencoder model, and the network regularization method uses divergence regularization;

[0108] S402, select normal samples from the prepared anomaly detection data set training set, pass through the abnormal pattern generation module, synthesize corresponding abnormal samples, send the normal samples and abnormal samples to the encoder of the reconstruction sub-network to be encoded as potential vectors, and after diffusion processing, use the noise embedding technology to generate a noisy vector, and finally obtain the noisy normal sample potential vector and the abnormal sample potential vector, send them to the denoising UNet network for denoising training, denoise after randomly selected time steps, output the potential vector of the reconstructed sample, perform channel interpolation operation, and perform channel interpolation on the potential vector of the input abnormal sample and the potential vector of the reconstructed sample to obtain the potential vector of the intermediate state:

[0109] z inter =λc+(1-λ)z r

[0110] Among them, c is the potential vector of the input abnormal sample pair, z r To reconstruct the potential vector of the sample, the interpolation control parameter λ is set to 0.5, z inter That is the potential vector of the intermediate state obtained;

[0111] Then, the latent vector of the reconstructed sample and the latent vector of the intermediate state are sent to the decoder to obtain the reconstructed image and the intermediate state image of the image space, and sent to the discriminant sub-network as input together with the abnormal sample, and the abnormal mask is output. The loss is calculated with the abnormal pattern mask used in the abnormal generation module, and the gradient is calculated. Then, the parameters of the discriminant sub-network and the reconstruction sub-network are guided to update to complete the network training.

[0112] Step S5: predicting the image to be inspected.

[0113] S501 , input the image to be detected into the trained surface anomaly detection network DiffusionADNet based on the diffusion model to obtain the detection result of the anomaly in the image, that is, the mask of the anomaly location.

[0114] Example 2

[0115] The present invention also provides an unsupervised surface anomaly detection system based on a diffusion probability model. The unsupervised surface anomaly detection system based on a diffusion probability model can be implemented by executing the process steps of the unsupervised surface anomaly detection method based on a diffusion probability model, that is, those skilled in the art can understand the unsupervised surface anomaly detection method based on a diffusion probability model as a preferred implementation of the unsupervised surface anomaly detection system based on a diffusion probability model.

[0116] The unsupervised surface anomaly detection system based on the diffusion probability model provided by the present invention includes: module M1: collecting image data; module M2: making anomaly detection data set based on the collected image data; module M3: building an unsupervised surface anomaly detection network based on the diffusion probability model; module M4: training the unsupervised surface anomaly detection network based on the diffusion probability model through the anomaly detection data set; module M5: performing anomaly detection on the image to be inspected through the trained unsupervised surface anomaly detection network based on the diffusion probability model, and obtaining the detection result of the anomaly in the image, that is, the mask for locating the anomaly.

[0117] The module M1 includes: module M101: using an industrial camera to collect image data of a scene to be inspected; module M102: classifying the image into abnormal patterns according to preset abnormal pattern features in the image;

[0118] The module M2 includes: module M201: according to the abnormal pattern classification, the abnormal detection data set is divided into a training set and a test set, wherein the training set places all samples without abnormalities, and the test set sets different subfolders to place samples with different abnormal patterns; module M202: annotating the samples with abnormal patterns in the test set, and placing the annotated files in the subfolders of each abnormal pattern under the annotated folder according to the test set folder settings.

[0119] The module M3 comprises: module M301: constructing an abnormal pattern generation module, using the abnormal pattern generated to deviate from the distribution to identify the abnormality; generating a noise image using a Perlin noise generator, and binarizing it through a uniformly randomly sampled threshold to generate an abnormal mask; sampling an abnormal texture source image from an abnormal source image data set that is independent of the input image distribution, and then performing random enhancement sampling, tone separation, sharpening, equalization, brightness change, color change and contrast enhancement on the abnormal source image data set; masking the enhanced texture image A with the abnormal mask, and comparing it with the training image I a Mixing, that is, generating abnormal patterns that deviate from the distribution; training image I a Defined as:

[0120]

[0121] Among them, I is the sample without abnormality in the data set, The mask M a The flip mask of , A is the enhanced texture image, and β is the opacity parameter;

[0122] Module M302: construct a reconstruction sub-network, which consists of a diffusion probability model, an encoder, and a decoder;

[0123] Among them, using the diffusion probability model, the training of the reconstruction subnetwork is transformed from the image pixel space to the low-dimensional potential representation space with the same perception; for the image in the image space, the encoder is used to learn the pixel-level reconstruction of the normal sample, the image is encoded into a potential space vector, and the diffusion process is performed, and then the diffusion model based on UNet is used to learn the semantic feature generation, and the potential space vector with noise added after the diffusion process is denoised. The training objective loss function L LDM It is expressed as:

[0124]

[0125] Among them, z and ∈ are noises that conform to the normal distribution, N(0,1) is the normal distribution, t is the number of time steps of random sampling, z t is the noise at time step t, ∈ θ The model UNet performs denoising operations, and E is the expectation under various conditions;

[0126] The sampling process is carried out in the latent space, and the reconstructed sample vector is sampled from the learned distribution and decoded by the decoder into a reconstructed sample image in the image space;

[0127] Module M303: Construct a discriminative subnetwork. The input of the discriminative subnetwork is the intermediate state image of the image space output by the reconstruction subnetwork, the reconstructed image and the abnormal sample synthesized by the normal image through the abnormal pattern generation module. The output of the network is the abnormal position mask, which is used to calculate the loss with the abnormal pattern mask used in the abnormal generation module to guide the update of the network.

[0128] In the module M3, an unsupervised surface anomaly detection network based on a diffusion model is constructed, using noise embedding and channel interpolation; in an unsupervised scenario of anomaly detection, noise embedding is used to guide sample generation; given a normal sample x, a synthetic anomaly sample x is obtained through the anomaly pattern generation module a , the synthetic abnormal sample is encoded into a latent vector by the encoder and regarded as the initial state c 0 , and then after T times of iterative diffusion process, the variance of the T times of added noise is β 1 …β T , then:

[0129]

[0130]

[0131] Where t is the number of random sampling time steps, c 0 ,c 1 ,c t-1 ,c t are the potential vectors of 0, 1, t-1, and t time steps respectively, and q is the conditional probability, that is, given c 0 Under these conditions, c t The probability of

[0132] Select a random time step t and get the diffusion-processed noise vector c noisy , put c noisy As a control condition, the diffusion probability model learns the target loss function L LDM becomes:

[0133]

[0134] Use the interpolation channel to make the discriminant subnetwork have recognition diversity in the reconstruction process and distinguish the real anomalies; the latent vector c after the abnormal input image is encoded by the encoder and the reconstructed sample z r Interpolate to get the intermediate state, the expression is: z inter =λc+(1-λ)z r , where λ is a weight parameter between 0 and 1.

[0135] The module M4 includes: module M401: the optimizer uses momentum-based stochastic gradient descent Adam, the learning rate decay strategy uses linear decay, the encoder and decoder in the reconstruction subnetwork use autoencoder models, and the network regularization method uses divergence regularization; module M402: normal samples are selected from the training set of the prepared anomaly detection data set, and the corresponding abnormal samples are synthesized through the abnormal pattern generation module. The normal samples and the abnormal samples are sent to the encoder of the reconstruction subnetwork to be encoded as potential vectors, and after diffusion processing, the noisy normal sample potential vectors and the abnormal sample potential vectors are obtained, which are sent to the denoising UNet network for denoising training, and after denoising for a randomly selected number of time steps, the potential vector of the reconstructed sample is output, the potential vector of the input abnormal sample and the potential vector of the reconstructed sample are channel interpolated to obtain the potential vector of the intermediate state, which is sent to the decoder to obtain the reconstructed image and the intermediate state image of the image space, which are sent to the discriminant subnetwork as input together with the abnormal sample, and the abnormal mask is output. The loss is calculated with the abnormal pattern mask used in the abnormal generation module, and the gradient is calculated, and then the parameters of the discriminant subnetwork and the reconstruction subnetwork are guided to update, and the network training is completed.

[0136] Those skilled in the art know that, in addition to implementing the system, device and its various modules provided by the present invention in a purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Therefore, the system, device and its various modules provided by the present invention can be considered as a hardware component, and the modules included therein for implementing various programs can also be considered as structures within the hardware component; the modules for implementing various functions can also be considered as both software programs for implementing the method and structures within the hardware component.

[0137] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. An unsupervised surface anomaly detection method based on a diffusion probability model, characterized in that: include: Step S1: collecting image data; Step S2: Based on the collected image data, an anomaly detection data set is generated; Step S3: constructing an unsupervised surface anomaly detection network based on a diffusion probability model; Step S4: training an unsupervised surface anomaly detection network based on a diffusion probability model using anomaly detection dataset; Step S5: Perform anomaly detection on the image to be inspected through the trained unsupervised surface anomaly detection network based on the diffusion probability model to obtain the detection result of the anomaly in the image, that is, the mask of the anomaly location.

2. The unsupervised surface anomaly detection method based on the diffusion probability model according to claim 1 is characterized in that: The step S1 comprises: Step S101: using an industrial camera to collect image data of a scene to be inspected; Step S102: classifying the image into abnormal patterns according to the preset abnormal pattern features in the image; The step S2 comprises: Step S201: According to the abnormal pattern classification, the abnormal detection data set is divided into a training set and a test set, wherein the training set contains all samples without abnormalities, and the test set sets different subfolders to respectively contain samples with different abnormal patterns; Step S202: Annotate samples with abnormal patterns in the test set, and place the annotation files in subfolders of each abnormal pattern under the annotation folder according to the test set folder settings.

3. The unsupervised surface anomaly detection method based on the diffusion probability model according to claim 1, characterized in that: The step S3 comprises: Step S301: construct an abnormal pattern generation module, and use the abnormal pattern generated by the deviation distribution to identify the abnormality; the noise image is generated by using the Perlin noise generator, and binarized by the threshold of uniform random sampling to generate an abnormal mask; the abnormal texture source image is sampled from an abnormal source image data set that is independent of the input image distribution, and then random enhancement sampling, tone separation, sharpening, equalization, brightness change, color change and contrast enhancement are performed on the abnormal source image data set; the enhanced texture image A is masked with the abnormal mask and compared with the training image I a Mixing, that is, generating abnormal patterns that deviate from the distribution; training image I a Defined as: Among them, I is the sample without abnormality in the data set, The mask M a The flip mask of , A is the enhanced texture image, and β is the opacity parameter; Step S302: construct a reconstruction sub-network, where the reconstruction sub-network consists of a diffusion probability model, an encoder, and a decoder; Among them, using the diffusion probability model, the training of the reconstruction subnetwork is transformed from the image pixel space to the low-dimensional potential representation space with the same perception; for the image in the image space, the encoder is used to learn the pixel-level reconstruction of the normal sample, the image is encoded into a potential space vector, and the diffusion process is performed, and then the diffusion model based on UNet is used to learn the semantic feature generation, and the potential space vector with noise added after the diffusion process is denoised. The training objective loss function L LDM It is expressed as: Among them, z and ∈ are noises that conform to the normal distribution, N(0,1) is the normal distribution, t is the number of time steps of random sampling, z t is the noise at time step t, ∈ θ is the model UNet that performs the denoising operation, and E is the expectation calculated under various conditions; The sampling process is carried out in the latent space, and the reconstructed sample vector is sampled from the learned distribution and decoded by the decoder into a reconstructed sample image in the image space; Step S303: construct a discriminant subnetwork, the input of which is the intermediate state image of the image space output by the reconstruction subnetwork, the reconstructed image and the abnormal sample synthesized by the normal image through the abnormal pattern generation module. The output of the network is the abnormal position mask, which is used to calculate the loss with the abnormal pattern mask used in the abnormal generation module to guide the update of the network.

4. The unsupervised surface anomaly detection method based on the diffusion probability model according to claim 3 is characterized in that: In the step S3, an unsupervised surface anomaly detection network based on a diffusion model is constructed, using noise embedding and channel interpolation; In the unsupervised scenario of anomaly detection, noise embedding is used to guide sample generation; Given a normal sample x, a synthetic abnormal sample x is obtained through the abnormal pattern generation module a , the synthetic abnormal sample is encoded into a latent vector by the encoder, which is regarded as the initial state c0, and then undergoes T-times iterative diffusion process, where the variance of the T-times added noise is β1…β T , then: Among them, t is the number of time steps of random sampling, c0, c1, c t-1 ,c t are the potential vectors of 0, 1, t-1, and t time steps respectively, and q is the conditional probability, that is, given c0, the probability of generating c t probability; Select a random time step t and get the diffusion-processed noise vector c noisy , put c noisy As a control condition, the diffusion probability model learns the target loss function L LDM becomes: Using interpolation channels, the discriminant subnetwork has recognition diversity during the reconstruction process and can distinguish real anomalies; The latent vector c and reconstructed sample z of the abnormal input image after being encoded by the encoder r Interpolate to get the intermediate state, the expression is: z inter =λc+(1-λ)z r , where λ is a weight parameter between 0 and 1.

5. The unsupervised surface anomaly detection method based on diffusion probability model according to claim 1, characterized in that: The step S4 comprises: Step S401: the optimizer uses momentum-based stochastic gradient descent Adam, the learning rate decay strategy uses linear decay, the encoder and decoder in the reconstruction sub-network use autoencoder models, and the network regularization method uses divergence regularization; Step S402: Select normal samples from the training set of the prepared anomaly detection data set, pass through the abnormal pattern generation module, synthesize corresponding abnormal samples, send the normal samples and abnormal samples to the encoder of the reconstruction subnetwork to be encoded as potential vectors, and after diffusion processing, obtain the noisy normal sample potential vector and the abnormal sample potential vector, send them to the denoising UNet network for denoising training, and after denoising for a randomly selected number of time steps, output the potential vector of the reconstructed sample, perform channel interpolation on the potential vector of the input abnormal sample and the potential vector of the reconstructed sample to obtain the potential vector of the intermediate state, send them to the decoder to obtain the reconstructed image and the intermediate state image of the image space, and send them together with the abnormal samples as input to the discriminant subnetwork, output the abnormal mask, calculate the loss with the abnormal pattern mask used in the abnormal generation module, calculate the gradient, and then guide the discriminant subnetwork and the reconstruction subnetwork parameter update to complete the network training.

6. An unsupervised surface anomaly detection system based on a diffusion probability model, characterized in that: include: Module M1: collect image data; Module M2: Produce anomaly detection dataset based on collected image data; Module M3: Constructing an unsupervised surface anomaly detection network based on diffusion probability model; Module M4: Train an unsupervised surface anomaly detection network based on a diffusion probability model using anomaly detection datasets; Module M5: Perform anomaly detection on the image to be inspected through the trained unsupervised surface anomaly detection network based on the diffusion probability model to obtain the detection result of the anomaly in the image, that is, the mask of the anomaly location.

7. The unsupervised surface anomaly detection system based on diffusion probability model according to claim 6, characterized in that: The module M1 comprises: Module M101: Use industrial cameras to collect image data of the scene to be inspected; Module M102: classifying the image into abnormal patterns according to the preset abnormal pattern features in the image; The module M2 comprises: Module M201: According to the abnormal pattern classification, the anomaly detection data set is divided into a training set and a test set, where the training set contains all samples without abnormalities, and the test set sets different subfolders to place samples with different abnormal patterns; Module M202: Label the samples with abnormal patterns in the test set, and place the labeling files in the subfolders of each abnormal pattern under the labeling folder according to the test set folder settings.

8. The unsupervised surface anomaly detection system based on diffusion probability model according to claim 7, characterized in that: The module M3 comprises: Module M301: Construct an abnormal pattern generation module, and use the abnormal pattern generated by the deviation distribution to identify the abnormality; the noise image is generated by using the Perlin noise generator, and is binarized by the threshold of uniform random sampling to generate an abnormal mask; the abnormal texture source image is sampled from an abnormal source image dataset that is independent of the input image distribution, and then random enhancement sampling, tone separation, sharpening, equalization, brightness change, color change and contrast enhancement are performed on the abnormal source image dataset; the enhanced texture image A is masked with the abnormal mask and compared with the training image I a Mixing, that is, generating abnormal patterns that deviate from the distribution; training image I a Defined as: Among them, I is the sample without abnormality in the data set, The mask M a The flip mask of , A is the enhanced texture image, and β is the opacity parameter; Module M302: construct a reconstruction sub-network, which consists of a diffusion probability model, an encoder, and a decoder; Among them, using the diffusion probability model, the training of the reconstruction subnetwork is transformed from the image pixel space to the low-dimensional potential representation space with the same perception; for the image in the image space, the encoder is used to learn the pixel-level reconstruction of the normal sample, the image is encoded into a potential space vector, and the diffusion process is performed, and then the diffusion model based on UNet is used to learn the semantic feature generation, and the potential space vector with noise added after the diffusion process is denoised. The training objective loss function L LDM It is expressed as: Among them, z and ∈ are noises that conform to the normal distribution, N(0,1) is the normal distribution, t is the number of time steps of random sampling, z t is the noise at time step t, ∈ θ is the model UNet that performs the denoising operation, and E is the expectation calculated under various conditions; The sampling process is carried out in the latent space, and the reconstructed sample vector is sampled from the learned distribution and decoded by the decoder into a reconstructed sample image in the image space; Module M303: Construct a discriminative subnetwork. The input of the discriminative subnetwork is the intermediate state image of the image space output by the reconstruction subnetwork, the reconstructed image and the abnormal sample synthesized by the normal image through the abnormal pattern generation module. The output of the network is the abnormal position mask, which is used to calculate the loss with the abnormal pattern mask used in the abnormal generation module to guide the update of the network.

9. The unsupervised surface anomaly detection system based on diffusion probability model according to claim 8, characterized in that: In the module M3, an unsupervised surface anomaly detection network based on a diffusion model is constructed, using noise embedding and channel interpolation; In the unsupervised scenario of anomaly detection, noise embedding is used to guide sample generation; Given a normal sample x, a synthetic abnormal sample x is obtained through the abnormal pattern generation module a , the synthetic abnormal sample is encoded into a latent vector by the encoder, which is regarded as the initial state c0, and then undergoes T-times iterative diffusion process, where the variance of the T-times added noise is β1…β T , then: Among them, t is the number of time steps of random sampling, c0, c1, c t-1 ,c t are the potential vectors of 0, 1, t-1, and t time steps respectively, and q is the conditional probability, that is, given c0, the probability of generating c t probability; Select a random time step t and get the diffusion-processed noise vector c noisy , put c noisy As a control condition, the diffusion probability model learns the target loss function L LDM becomes: Using interpolation channels, the discriminant subnetwork has recognition diversity during the reconstruction process and can distinguish real anomalies; The latent vector c and reconstructed sample z of the abnormal input image after being encoded by the encoder r Interpolate to get the intermediate state, the expression is: z inter =λc+(1-λ)z r , where λ is a weight parameter between 0 and 1.

10. The unsupervised surface anomaly detection system based on diffusion probability model according to claim 9, characterized in that: The module M4 comprises: Module M401: The optimizer uses momentum-based stochastic gradient descent Adam, the learning rate decay strategy uses linear decay, the encoder and decoder in the reconstruction sub-network use autoencoder models, and the network regularization method uses divergence regularization; Module M402: Select normal samples from the training set of the prepared anomaly detection dataset, pass through the abnormal pattern generation module, synthesize the corresponding abnormal samples, send the normal samples and the abnormal samples to the encoder of the reconstruction sub-network to be encoded as latent vectors, and after diffusion processing, obtain the noisy normal sample latent vector and the abnormal sample latent vector, which are sent to the denoising UNet network for denoising training. After denoising for a randomly selected number of time steps, output the latent vector of the reconstructed sample, perform channel interpolation on the latent vector of the input abnormal sample and the latent vector of the reconstructed sample to obtain the latent vector of the intermediate state, send it to the decoder to obtain the reconstructed image and the intermediate state image in the image space, and send them together with the abnormal samples as input to the discriminant sub-network, output the abnormal mask, calculate the loss with the abnormal pattern mask used in the abnormal generation module, calculate the gradient, and then guide the update of the parameters of the discriminant sub-network and the reconstruction sub-network to complete the network training.