A surface defect detection method based on a generative adversarial network

CN118864342BActive Publication Date: 2026-09-29CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310484962.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-28
Publication Date
2026-09-29
Estimated Expiration
2043-04-28

AI Technical Summary

Technical Problem

[0007]本发明的目的在于克服上述不足,提供了一种基于生成对抗网络的缺陷检测方法,基于无缺陷样本训练具有矢量量化器和判别器的生成网络模型,并设计了专门的缺陷分割模块和基于Transformer的精修方法,可得到待检测图像的有精确缺陷轮廓的缺陷分割图,该检测方法的步骤流程图可参阅附图1,解决现有无监督表面缺陷检测方法缺陷分割图不准确、面积较小的缺陷漏检率高的问题

Benefits of technology

[0118]1、本发明所提供的一种新的缺陷分割方法,使得缺陷分割图具有更精确的缺陷轮廓,从而有效去除重建残差中的噪声,明显提升缺陷分割质量,;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118864342B_ABST
    Figure CN118864342B_ABST
Patent Text Reader

Abstract

The application relates to the field of image processing and discloses a surface defect detection method based on a generative adversarial network. A surface defect detection model based on an unsupervised reconstruction idea is constructed, a vector quantization variational automatic encoder is introduced into the defect detection model, discrete feature representation of an input image is obtained, samples are reconstructed in an adversarial training manner, a plurality of loss functions are used according to a preset threshold value, so that the model can repair defect samples, a method based on Transformer fine repair is proposed to perform masking processing on some untrusted positions in a defect sample feature index sequence, the Transformer is used to perform multiple predictions on the masked positions, the masked positions are decoded into images and are subjected to average operation to achieve the purpose of fine repair of the images. Meanwhile, a new defect segmentation method for adaptively segmenting a difference feature map by using an OSTU algorithm is provided, and the purpose of accurate defect segmentation is achieved. Compared with the prior art, the application can obtain a more complete reconstructed image, has more excellent defect accurate segmentation effect, avoids manual labeling, solves the problem that the false detection rate and the missed detection rate are high in the detection process of model training by using a single convolutional neural network in an industrial scene to a certain extent, can realize accurate and rapid segmentation of various defects, and is more practical and has more application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision, specifically relating to a surface defect detection method based on generative adversarial networks. Background Technology

[0002] In industrial manufacturing, surface defect detection is a crucial step in the actual production process. Surface defects not only affect the appearance of products but also directly impact their functionality and performance, potentially leading to serious consequences and losses for businesses. Therefore, it is essential to conduct surface defect detection to effectively ensure the yield rate of products entering the market and maintain a positive corporate image.

[0003] Traditional visual inspection methods for identifying surface defects rely on human observation and experience. This approach is not only susceptible to subjective factors but also costly and inefficient, failing to meet the real-time inspection requirements of industrial production and hindering the development of industrial informatization. With the advancement of computer vision in industrial inspection, defect detection based on convolutional neural networks has been extensively studied, including object detection, semantic segmentation, detection methods based on generative adversarial networks and autoencoders, and methods based on pre-trained models. Among these, object detection models based on standard detection models such as SSD, YOLO, and Faster R-CNN typically offer good accuracy and speed, but they can only locate defects using bounding boxes and cannot extract accurate defect contours, thus failing to adequately meet market demands for certain specific detection tasks.

[0004] The paper "Automatic defect detection and segmentation of tunnel surface using modified Mask R-CNN" proposes a tunnel defect detection method based on the semantic segmentation model Mask R-CNN, adding a Path Enhancement Feature Pyramid Network (PAFPN) and an edge detection branch to improve network detection accuracy. The invention patent "A surface defect segmentation method based on two-stage incremental learning" (patent number: 202211600304.2) uses the UNet model as the basis for semantic segmentation, improving the learning rate through model parameter optimization while overcoming the catastrophic forgetting problem, making it better applicable to surface defect segmentation. While semantic segmentation models such as Mask R-CNN, UNet, and SegNet can locate defects and segment their actual contours, they are all based on supervised learning methods, requiring a large number of samples and pixel-level annotations during the training phase. In real-world industrial scenarios, the number of defect samples is usually small, and manual annotation is costly and susceptible to subjective factors.

[0005] Unsupervised defect detection methods have attracted attention in recent years because they do not require manual labeling of samples during the training phase. In the paper "A Defect Detection Method for Medical Glass Bottle Mouth Based on Convolutional Autoencoder", a defect detection model based on a convolutional autoencoder network with an encoder-decoder structure is established. Convolutional attention modules with spatial and channel attention are embedded in the encoder to enhance the network's feature extraction capabilities. Multi-scale structural similarity (MS-SSIM) and L1 loss are combined to improve image reconstruction, achieving accurate and efficient automated product quality inspection. In the paper "Automatic fabric defect detection with a multi-scale convolutional denoising autoencoder network model", a convolutional denoising autoencoder network is used to reconstruct image patches at multiple Gaussian pyramid layers. The detection results from each resolution channel are integrated, and the reconstruction residual of each image patch is used as an indicator for direct pixel-level prediction. The final detection result is generated by segmenting and integrating the reconstruction residual maps at each resolution level. In the patent "An Unsupervised Defect Detection Method for Photovoltaic Modules Based on an Improved GAN Algorithm" (Patent No.: 202010135948.3), an SSIM-GAN model is constructed. The encoder network generates the latent space corresponding to the input normal image, and structural similarity is used to describe the differences between images to determine whether a defect exists.

[0006] The aforementioned unsupervised defect detection methods primarily rely on the autoencoder model structure and reconstruction errors to determine whether an image to be detected has defects. When the target structure is complex, the reconstruction residuals of the model on the image to be detected are usually large, especially in areas that were originally without defects. Furthermore, directly constructing a defect contour segmentation map based on the reconstruction residuals is often affected by noise residuals, making it difficult to accurately detect small defects, and the resulting defect contour segmentation map is inaccurate and cannot meet the needs of practical applications. Summary of the Invention

[0007] The purpose of this invention is to overcome the above-mentioned shortcomings and provide a defect detection method based on generative adversarial networks. This method trains a generative network model with a vector quantizer and a discriminator based on defect-free samples, and designs a dedicated defect segmentation module and a Transformer-based refinement method. This yields a defect segmentation map with accurate defect contours for the image to be detected. A flowchart of the detection method's steps is provided in the appendix. Figure 1 This addresses the problems of inaccurate defect segmentation maps and high missed detection rates for small-area defects in existing unsupervised surface defect detection methods.

[0008] The disclosed defect detection method includes the following steps:

[0009] S1. Using normal, defect-free samples, train a generative adversarial network with the following structure:

[0010] (1) The network consists of an encoder, a vector quantizer, a decoder and a discriminator in sequence;

[0011] (2) The encoder and decoder contain an SE-ResNet module and a SAGAN-based self-attention module (AttnBlock). The SE-ResNet module adds the results of the input feature map x after passing through the Residual module and ShortCut operation respectively to obtain a new feature map, and then performs global pooling. Next, it performs two convolution operations with a kernel size of 1 and then uses the Sigmoid activation function to obtain the importance of each feature channel. Based on the channel importance, the features before pooling are combined. The self-attention module (Attn Block) takes the convolution feature map x as the input feature vector and performs linear transformation using a convolutional layer with a kernel size of 1. The output g(x), f(x), h(x) of the linear transformation correspond to the Q, K, V of the Transformer model and output the self-attention result.

[0012] (3) There is a vector quantizer between the encoder and the decoder. The vector quantizer uses a specific embedding space to match and replace the output features of the encoder. In the form of dictionary learning, each dimension of the encoder output features is sampled from the embedding space according to the probability and replaced, thereby obtaining the discrete feature map of the input image.

[0013] (4) After the decoder, a discriminator is connected. The discriminator uses the discriminator of the PatchGAN model, which outputs a matrix. Each element in the matrix is ​​an evaluation of a local region (receptive field) in the image.

[0014] S2. Preprocess the image to be detected according to the model's requirements for the input image;

[0015] S3. Define a Transformer-based image retouching method that resamples the feature map index after vector quantization.

[0016] S4. Using the generative adversarial network trained in step S1 and the refinement method invented in step S3, the image to be detected after processing in step S2 is reconstructed.

[0017] S5. Input the reconstructed image obtained in step S4 and the image to be detected in step S2 into the defect segmenter of the invention to obtain a defect segmentation map with accurate defect contours.

[0018] Preferably, the training process of the generative adversarial network described in step S1 can be found in the appendix. Figure 2 Specifically:

[0019] S11. Prepare normal, defect-free images;

[0020] S12. Construct the training dataset. This step is achieved through the following sub-steps:

[0021] S121. Resize the images in the training set: First, convert the original images to grayscale and resize them to 256×256. Then, randomly crop the images to a size of 224×224 to fit the model.

[0022] S122. Based on the size of the training dataset, select image augmentation methods such as random horizontal flipping, mix-up, and Mosaic to augment the data of the resized images;

[0023] S13. Construct a generative adversarial network model. This step is achieved through the following sub-steps:

[0024] S131. Construct the generator network structure; the generator is constructed based on a vector quantization variational autoencoder (VQ-VAE) and consists of an encoder, a vector quantizer, and a decoder connected in sequence. This step is implemented through the following sub-steps:

[0025] S1311. Construct the encoder network structure; the encoder includes four modules connected in sequence: SE-ResNet (Squeeze-and-Excitation-Residual Network), three downsampling modules, a self-attention module (Attn Block) based on the SAGAN network, and a mid-layer module; the downsampling module consists of two SE-ResNet modules, with a downsampling convolutional layer with a kernel size of 3 and a stride of 2 in between; the mid-layer module consists of two SE-ResNet modules, with a self-attention module in between; the SE-ResNet module combines the SE module and the Residual module, the SE module can suppress unimportant channels, and the Residual module ensures that gradient vanishing does not occur while building a deeper network; the self-attention module solves the problem of limited receptive field of convolution and also simplifies computation; the activation function of the encoder is the Swish activation function, calculated as follows:

[0026] f(x)=x·sigmoid(x) (1)

[0027] S1312. Construct a vector quantizer: map the continuous features of the encoded image to a discrete feature space;

[0028] (1) The input image is encoded and output as z e(x), its channel number is transformed into the embedding space through a convolution with a kernel size of 1. The number of vectors in the middle, K, is denoted as z. e ′(x), then for z e The channel dimension of ′(x) is subjected to GumbelSoftmax operation, specifically:

[0029]

[0030] Among them, g i =-log(-log(u) i )),u i ~U(0,1) is called Gumbel noise, and its purpose is to add randomness to increase the exploration degree; π i z represents the probability that the i-th vector in the embedding space is selected. i τ represents the sampling probability of the i-th vector in the embedding space. The smaller τ is (τ→0), the closer the entire softmax function is to argmax, that is, the closer the z vector is to the one-hot distribution.

[0031] (2) Select the index k of the largest component in the z vector, and then look up the table to find e k Vector as input z to the decoder q (x), that is:

[0032] z q (x)=e k , where k = argmax j z (3)

[0033] S1313. Constructing the decoder structure: The decoder includes, in sequence, an SE-ResNet, a mid-layer module (MidBlock), a self-attention module (Attn Block) based on the SAGAN network, three upsampling modules (UPsample), and a convolutional layer with a 3×3 convolutional kernel; the upsampling module uses two SE-ResNet modules, and a combination of nearest neighbor interpolation and convolution is used between them; the activation function of the decoder is the Swish function;

[0034] S132. Constructing the discriminator network: The discriminator consists of five sub-layers. The first four sub-layers extract features from the input image. Each sub-layer consists of a convolutional layer, a normalization layer, and an activation layer. The last sub-layer consists of a convolutional layer and an activation layer. Here, the convolutional layer reduces the number of channels in the feature map to 1. The output of the discriminator is an N×N matrix, where each value represents the evaluation of a region (receptive field) in the image. The activation function of the discriminator is LeakyReLU, calculated as follows:

[0035] f(x)=max(αx,x) (4)

[0036] In equation (4), α is a relatively small positive value, which is taken as 0.2 in this invention;

[0037] The discriminator normalization layer uses a group normalization function, calculated as follows:

[0038]

[0039] In equation (5) above, x i Let μ represent the i-th set of data. G This represents the mean of the i-th data set. This represents the variance of the i-th data set, where ∈ is a small positive value, typically 10. -6 γ and β represent learnable parameters;

[0040] S133, The generator and discriminator constitute a generative adversarial network module;

[0041] S14. Train the generative adversarial network model. This step is achieved through the following sub-steps:

[0042] S141. Input the image preprocessed in step S12 into the encoder;

[0043] S142, a feature map with encoder output dimensions of 28×28×C (C represents the number of feature channels);

[0044] S143. Replace the above feature map with a new feature map composed of vectors in the embedding space using a vector quantizer;

[0045] S144. Input the new feature map generated in step S143 into the decoder to reconstruct the input image;

[0046] S145. Input the preprocessed input image and the reconstructed image into the discriminator, and calculate the generator loss and the discriminator loss respectively. The generator loss is:

[0047] L G =λ1 L rec +λ2L ssim +λ3L KL +λ4L adv (6) In the above formula (6), L rec The reconstruction loss is defined as:

[0048] L rec =E x~p(x) [||xG(x)||1] (7)

[0049] In equation (6) above, L ssimAs a structural loss, considering that the SSIM metric measures the similarity between two images based on their brightness, contrast, and texture structure, an SSIM-related loss is introduced to improve the quality of the generated images and accelerate the convergence speed of the network. The loss is defined as follows:

[0050]

[0051] In equation (8) above, μ x σ represents the mean of x. x p represents the standard deviation of x. x μ represents the distribution of the real data. x ,μ G(x) These are the average gray levels of each pixel in the input image and the reconstructed image, calculated using a sliding window method, respectively, σ. xG(x) Let G(x) be the covariance of the contrast parameters between the original input image x and the reconstructed image G(x);

[0052] In equation (6) above, L adv This is an adversarial loss introduced to further improve the quality of the generated image, defined as:

[0053] L adv =E x~p(x) (D(G(x))-1) 2 (9)

[0054] In equation (6) above, L KL The KL divergence loss is a constraint between the posterior and prior distributions in a variational autoencoder, defined as:

[0055] L KL =D KL (q φ (z∣x)||p θ (z)) (10)

[0056] Where q φ (z|x) represents z e ′(x) is the probability distribution normalized along the channel dimension, and pθ(z) represents the sampling distribution of the vectors in the assumed embedding space, which follows a uniform distribution.

[0057] The discriminator loss is:

[0058] L D =E x~p(x) (D(x)-1) 2 +E x~p(x) (D(G(x))-0) 2 (11)

[0059] S146. Based on the generator and discriminator loss functions, use Adam as the optimizer for the generator and discriminator to update the weights of the generative adversarial network.

[0060] S15. Save the parameters and trained weights generated in step S14 of the adversarial network.

[0061] As a preferred embodiment, the specific process for preprocessing the test set images in step S2 is as follows:

[0062] S21. Resize the test set images: First, convert the original images to grayscale and resize them to 256×256;

[0063] S22. Crop the center of the image obtained in step S21 to a size of 224×224;

[0064] Preferably, the Transformer-based image retouching network structure described in step S3 uses the GeLU activation function. The flowchart for this module can be found in the appendix. Figure 4 The specific steps are as follows:

[0065] S31. Train the Transformer-based image retouching network module. This step is implemented through the following sub-steps:

[0066] S311. Serialize the embedded spatial index feature map obtained after vector quantization, randomly select 25% of the positions in the index sequence, and simultaneously mask the selected index with an 80% probability, randomly replace the current position index with a 10% probability, and keep it unchanged with a 10% probability.

[0067] S312. Use the index sequence processed in step S311 as the input of the Transformer;

[0068] S313. Use Transformer to predict the new index of the selected position;

[0069] S314. Calculate the cross-entropy loss function between the index before occlusion of the selected position and the new index after prediction by the Transformer, using the AdamW optimizer, where β1 = 0.9, β2 = 0.95, the learning rate is reduced from 0.0003 to 0.00005 after 100 epochs of cosine annealing, and the maximum training epoch is 200.

[0070] S32. Save the parameters and trained weights in the Transformer-based refined network module obtained in step S31.

[0071] S33. Using the training results of step S31, the image to be detected after preprocessing in step S2 is encoded and vector quantized to obtain a sampling probability map, and the index values ​​corresponding to the positions with a probability less than 0.5 are masked.

[0072] S34. Input the masked index sequence from step S33 into the Transformer and use the Transformer to predict the new index at the masked position.

[0073] S35. The new index sequence is input into the decoder and decoded into a reconstructed new image. The occluded index sequence is resampled 8 times, and steps S33 and S34 are performed sequentially to obtain 8 new images.

[0074] S36 and S35 perform an arithmetic average operation on the 8 images obtained, and finally obtain a refined image as the reconstructed image of the original image to be detected.

[0075] Preferably, step S4 involves reconstructing the input image to be detected, and the specific steps are as follows:

[0076] S41. Read the trained generative adversarial network from step S1;

[0077] S42. Read the Transformer-based image retouching module trained in step S3;

[0078] S43. Input the preprocessed image to be detected from step S2 into the improved network structure to obtain the reconstructed image to be detected.

[0079] Preferably, step S5 describes a novel defect segmenter proposed in this invention. A schematic diagram of the defect segmentation process can be found in the appendix. Figure 5 Specifically, it includes the following steps:

[0080] S51. Construct the SSIM similarity calculation formula to calculate the difference feature map between the reconstructed image and the image to be detected after preprocessing in step S2.

[0081] S511. Construct local SSIM exponents using a sliding window method, fix the SSIM window size at 5, traverse the entire image pixel by pixel, and use a Gaussian kernel with a standard deviation of 1.0 for weighted averaging.

[0082] S512, Let μ x ,μ G(x) The average gray levels of each pixel in the input image obtained using the sliding window method and the reconstructed image are used as the measure of brightness to calculate the brightness similarity L. (x,G(x)) The calculation formula is:

[0083]

[0084] S513, Calculate contrast similarity C (x,G(x)) The calculation formula is:

[0085]

[0086] In the above formula (13) It is based on the average brightness μ of each pixel within the window. x The standard deviation of each pixel within the calculated window is called the contrast parameter;

[0087] S514, Calculate structural similarity s (x,G(x)) The calculation formula is:

[0088]

[0089] In equation (14) above, σ xG(x) Let G(x) be the covariance of the contrast parameters between the original input image x and the reconstructed image G(x) within the window.

[0090] S515. Based on steps S512, S513, and S514, construct the SSIM metric to measure the similarity between the input image and the reconstructed image. The expression is:

[0091] SSIM = L (x,G(x)) C (x,g(x)) S (x,G(x)) (15)

[0092] Meanwhile, in this invention, the constants C2 and C3 in equation (15) satisfy a numerical relationship. The final local SSIM exponent expression is obtained as follows:

[0093]

[0094] S515. Perform reflection padding of 2 on the image obtained in step S2 and the reconstructed image in step S4. Then, the SSIM window calculates the local structural similarity of the two images in a sliding manner, and finally obtains the difference feature map of the same size as the test image.

[0095] S52. Use the OTU algorithm to perform adaptive threshold segmentation on the feature map obtained in step S51 to obtain a preliminary residual map. The specific steps are as follows:

[0096] S521. Find the foreground-background segmentation threshold t. This step can be achieved through the following sub-steps:

[0097] S5211. Count the number of each pixel in the reconstructed image in the entire image;

[0098] S5212. Calculate the probability distribution of each pixel in the entire image;

[0099] S5213. Traverse each grayscale layer and calculate the inter-class probability of foreground and background under the current grayscale value;

[0100] S5214. Establish the objective function: the inter-class variance g(t) when the segmentation threshold is t, expressed as:

[0101] g(t) = w0 × (u0 - u) 2 +w1×(u1-u) 2 = w0 × w1 × (u1 - u0) 2 (17)

[0102] In equation (17), w0 refers to the number of foreground pixels, w1 refers to the number of background pixels, u0 refers to the foreground weighted average, u1 refers to the background weighted average, and u = u0 × w0 + u1 × w1 refers to the weighted average of the entire image.

[0103] S5215. Find the t that maximizes g(t) in step S5214 and use it as the segmentation threshold between the foreground and the background.

[0104] S522. Values ​​less than the threshold t obtained in step S521 will be classified as background, and values ​​greater than or equal to the threshold t will be classified as candidate defects, thus obtaining a residual map with some noise.

[0105] S53. Traverse the maximum area of ​​the candidate defect region in each residual map and calculate the area value that makes the F1 index the highest as the optimal segmentation threshold K.

[0106] S531. Find the candidate defect with the largest area from the residual map in step S52, and count the maximum defect contour area of ​​each residual map.

[0107] S532. Take the area value that maximizes the F1 score from step S531 as the optimal threshold K for image segmentation, where the expression for the F1 score is:

[0108]

[0109] In equation (18) above, precision is the accuracy in the confusion matrix, and recall is the recall in the confusion matrix;

[0110] S54. Using the optimal segmentation threshold K obtained in step S53, the residual map in step S52 is binarized to obtain a more accurate defect segmentation map with most of the noise removed.

[0111] S541. If the area of ​​all candidate regions in the image in step S52 is less than the optimal segmentation threshold K obtained in step S53, the image will be judged as defect-free; if there is a defect contour area in the image that is greater than the optimal segmentation threshold K in step S53, the image will be judged as a defective image.

[0112] S542. If the detected image in step S541 is identified as a defect image, then the gray values ​​within the contours of the image that are less than the optimal segmentation threshold are all set to 0, resulting in a defect segmentation image with most of the noise removed and accurate defect location.

[0113] In summary, the distinguishing technical features disclosed in this invention include at least the following:

[0114] 1. This invention uses the OTU algorithm to perform adaptive threshold segmentation on the difference feature map, finds the area value that makes the highest F1 score as the optimal threshold K for judging whether there is a defect, and then binarizes the image that is judged to be defective to obtain a defect segmentation map with most of the noise removed and accurate defect contours. This provides a novel defect segmentation method that can effectively remove noise in the reconstruction residual and effectively detect defects with small areas.

[0115] 2. The model disclosed in this invention contains a vector quantizer, which can encode the input image through a sparse coding space to obtain a discrete feature space, which can play a regularization role, reduce overfitting during the training stage, improve the quality of reconstructed images, and lay the foundation for the image retouching module.

[0116] 3. This invention provides an image retouching method based on Transformer. After serializing the index feature map, the unconfident indices in the index sequence are masked, and the new index at the masked position is predicted by Transformer. This solves the problem of large reconstruction residuals caused by the defect positions not being trained during the training phase.

[0117] Compared with the prior art, the beneficial effects of the present invention are:

[0118] 1. The present invention provides a novel defect segmentation method that enables the defect segmentation image to have a more accurate defect contour, thereby effectively removing noise in the reconstruction residual and significantly improving the defect segmentation quality;

[0119] 2. Based on this, the proposed Transformer-based image retouching method improves the image quality reconstructed by the vector quantization variational autoencoder, effectively reduces the missed detection of defect areas, and makes up for the shortcomings of existing defect detection methods in terms of low accuracy in detecting defects of specific areas.

[0120] 3. The technology disclosed in this invention is an unsupervised defect detection method based on a generation method, which only requires normal samples without defects and is suitable for scenarios where samples lack manual annotation and the number of defective samples is small. Attached Figure Description

[0121] Figure 1 A flowchart illustrating the steps of a surface defect detection method based on generative adversarial networks provided by this invention.

[0122] Figure 2 Flowchart for training a generative adversarial network

[0123] Figure 3 Unsupervised defect detection model for the testing phase

[0124] Figure 4 Flowchart of the Transformer-based image retouching network module

[0125] Figure 5 This is a schematic diagram of the defect segmentation process of the present invention.

[0126] Figure 6 Examples of datasets used in embodiments of the present invention

[0127] Figure 7 Defect segmentation results for Experiment Category 6 dataset

[0128] Figure 8 Defect segmentation results for Experiment Category 7 dataset

[0129] Figure 9 Example figures showing the segmentation results of test category 6 on supervised and unsupervised algorithms.

[0130] Figure 10 Example figures showing the segmentation results of test category 6 on supervised and unsupervised algorithms.

[0131] Figure 11 Example of missed detection

[0132] Figure 12 The defect segmentation effect diagram after adding the "Refinement" module. Detailed Implementation

[0133] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0134] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0135] Example 1:

[0136] A dataset was selected to test the effectiveness of the published method. To simulate the scarcity of labeled data in real-world scenarios, the industrial defect detection dataset DAGM2007 was chosen, which contains ten categories of artificially generated texture defects. The first six categories each contain 1000 defect-free images and 150 defective images, while the last four categories each contain 2000 defect-free images and 300 defective images. To better represent real-world scenarios in industrial inspection, two representative categories, Class 6 and Class 7, were selected from this dataset for experiments. Class 6 has a more complex surface texture and larger defect areas, while Class 7 has smaller defects. Image examples are shown below. Figure 6 The column for Class 6 represents defect samples and defect labels for category six, and the column for Class 7 represents defect samples and defect labels for category seven. The datasets for the two categories are divided into test set, training set, and validation set, respectively. "N" represents "Normal" images without defects, and "P" represents "Paranormal" images with defects. Table 1 below lists the specific number of images and image categories for each division.

[0137] Table 1. Dataset partitioning for DAGM2007

[0138]

[0139] Recall, precision, accuracy, and F1 score—common metrics in classification problems—are used as evaluation metrics. True negatives (TP) represent the number of correctly classified defective images, false negatives (FN) represent the number of defective images classified as defect-free images, false positives (FP) represent the number of defect-free images classified as defective images, and true negatives (TN) represent the number of correctly classified defect-free images. Table 2 below lists the evaluation metrics for this embodiment.

[0140] Table 2 Evaluation Indicators for Example 1

[0141]

[0142] See Figure 1The provided flowchart of the surface defect detection method is shown. First, the input image is preprocessed: the size of the 512×512 grayscale image provided by the dataset is adjusted to 256×256. To improve the robustness of the model, the image is further randomly cropped to 224×224 and randomly horizontally flipped to enhance the image. Adam is used as the optimizer for the generator and discriminator with β1=0.5 and β2=0.9. The maximum training epoch is set to 200, and early stopping is performed based on the reconstruction loss of the validation set. The embedding space contains 64 vectors, each with a dimension of 64. The initial learning rate of the generator is set to 0.0002 and reduced to 0.00001 by cosine annealing for 100 epochs. The initial learning rate of the discriminator is set to 0.0001. In formula (3), λ1=1.0, λ2=0.5, and λ4 is increased from 0 to 0.0005 by cosine annealing for 50 epochs. λ3 follows the calculation formula:

[0143]

[0144] In equation (16) above, G L The numerator represents the weight parameters of the last layer of the decoder, and the numerator represents the reconstruction loss (including SSIM loss) relative to G. L The 2-norm of the gradient, where the denominator represents the adversarial loss relative to G. L The 2-norm of the gradient is set to 0.0001 to prevent the denominator from being zero.

[0145] See Figure 3 The provided diagram illustrates the unsupervised surface defect detection method. Experiments were conducted on preprocessed dataset images, presenting the results of various metrics for defect detection using the proposed surface defect detection method on categories 6 and 7 of the DAGM2007 dataset. Table 3 below lists the experimental results for each metric in this embodiment.

[0146] Table 3 Model Performance Indicators

[0147]

[0148] For the defect segmentation results of Category 6 defect samples, please refer to [link / reference]. Figure 7 For the defect segmentation results of the defect samples in Category 7, please refer to [link / reference]. Figure 8 Column (a) is the original image, column (b) is the reconstructed image, column (c) is the original label, column (d) is the pixel binarized image of the test image and the reconstructed image extracted using the OTU algorithm, and column (e) is the defect segmentation image after thresholding based on the maximum difference contour area. Figure 6 and Figure 7 The segmentation results shown demonstrate that the defect detection method disclosed in this invention can effectively segment large-area defects in images with complex backgrounds. Figure 8The segmentation results shown demonstrate that the defect detection method disclosed in this invention can effectively segment defects with small areas.

[0149] Example 2:

[0150] The disclosed method was compared with several other mainstream defect detection methods, including f-AnoGan, PatchCore, CycleGan, and Unet. f-AnoGan and PatchCore are unsupervised models, CycleGan is a semi-supervised model, and Unet is a supervised model. Furthermore, the defect segmentation maps for f-AnoGan and CycleGan are residual maps between the test image and the reconstructed image.

[0151] f-AnoGan and PacthCore are trained using the same dataset. In the f-AnoGan model, training samples are segmented into 56*56 images. During the testing phase, a 224*224 image is segmented into 16 56*56 images. Images are generated using a 100-dimensional random vector, and then the reconstructed images are sequentially stitched together to form a 224*224 image. The largest defect score among the 16 images is selected as the defect score for the test image. In the CycleGan and Unet models, due to the imbalance between the number of defective and non-defective samples, the original training and test sets of DAGM2007 from Example 1 are used for training and inference. Specifically, the sample distribution is as follows: Class 6 training set includes 492 non-defective samples and 83 defective samples, and its test set includes 508 non-defective samples and 67 defective samples; Class 7 training and test sets each include 1000 non-defective samples and 150 defective samples.

[0152] The evaluation metric is the AUC value, which is the area under the ROC curve, used to measure classifier performance. A value closer to 1 indicates better classifier performance; conversely, a value closer to 0 indicates worse performance. Table 4 below lists the experimental results comparing the AUC values ​​of various supervised and unsupervised algorithms.

[0153] Table 4 Comparison of AUC Indices for Each Model

[0154]

[0155] As shown in Table 4 above, the surface defect detection model provided by this invention significantly outperforms the unsupervised model f-AnoGan and the semi-supervised model CycleGan in defect classification tasks, improving the AUC index by 0.2409 and 0.1112 respectively; and is comparable to the performance of Unet based on supervised learning and the current unsupervised defect detection state-of-the-art model PatchCore.

[0156] For the performance of unsupervised and supervised algorithms on defect segmentation on the Class6 and Class7 datasets, please refer to [link / reference]. Figure 9 and Figure 10 Figure (1) shows an example of the segmentation effect of the unsupervised algorithm on the dataset, where (a) represents the original defect image, (b) represents the label of the defect image, (c) represents the segmentation result of f-AnoGan, (d) represents the segmentation result of PatchCore, and (e) represents the segmentation result of the present invention. Figure (2) shows an example of the segmentation effect of the supervised algorithm on the dataset. Where (a) represents the original defect image, (b) represents the label of the defect image, (c) represents the segmentation result of CycleGan, (d) represents the segmentation result of Unet, and (e) represents the segmentation result of the present invention. From the defect segmentation effect, the segmentation results of the present invention and PatchCore are the best. The f-AnoGan and CycleGan models are not ideal on the Class7 dataset with small defect areas, while Unet segments the largest possible region. Furthermore, while PatchCore performs well in both classification and segmentation, it uses a pre-trained model based on ImageNet for inference, which limits the size of the input image. In actual production, the samples collected may not conform to the input size of the pre-trained model, requiring the input image to be cropped, which increases the inference time during the testing phase.

[0157] Example 3:

[0158] Considering that the performance of defect detection based on image reconstruction strongly depends on the restoration effect of the test image, the better the restoration effect, the simpler and more effective the defect segmentation. For Class 6 images with large-area defects involved in Examples 1 and 2, the method disclosed in this invention, without fine-tuning operations, has a certain degree of missed detection. See the example image for missed detection. Figure 11 , where (a) represents the defect image, (b) represents the reconstructed image, (c) represents the defect annotation, (d) represents the SSIM residual map of (a) and (b), and (e) represents the defect segmentation result.

[0159] The image retouching method based on Transformer provided in this invention can effectively improve the retouching effect and obtain a better reconstructed image. See the principle section below. Figure 4During the training phase, 25% of the positions in the index sequence of the training set are randomly selected (784 * 0.25 = 196). Simultaneously, the selected index is masked with an 80% probability, the current position index is randomly replaced with a 10% probability, and it remains unchanged with a 10% probability. The processed index sequence is used as input to the Transformer, with the target being the original, unprocessed index sequence. The loss function is the cross-entropy loss function at the selected position. The optimizer uses AdamW, where β1 = 0.9, β2 = 0.95, and the learning rate is reduced from 0.0003 to 0.00005 after 100 epochs of cosine annealing. The maximum training epoch is 200.

[0160] During the testing phase, based on the sampling probability feature map obtained by vector quantization, the index values ​​corresponding to positions with a probability less than 0.5 are masked. The masked index sequence is then input into the Transformer. The 64-dimensional vector output at the masked position is annealed at a temperature τ of 0.8 and then normalized. The first 16 output values ​​with higher probabilities are retained. Then, eight sampling operations are performed based on the probability values ​​to obtain eight new feature index sequences. These eight feature sequences are decoded into images, and finally, the arithmetic mean of these eight images is performed to obtain a refined image.

[0161] The evaluation metric used is the DICE metric, commonly found in semantic segmentation, to assess the model's defect detection performance. Its calculation method is as follows: Where TP represents the number of correctly classified defective pixels, FN represents the number of defective pixels classified as defect-free pixels, and FP represents the number of defect-free pixels classified as defective pixels. The closer the dice index is to 1, the better the segmentation effect. Table 5 below lists the comparison of experimental results indices.

[0162] Table 5 Comparison of segmentation results before and after adding the refinement module.

[0163]

[0164] As shown in Table 5 above, the defect segmentation effect after Transformer refinement is significantly improved, the number of missed detection areas is significantly reduced, the defect segmentation DICE index is improved by about 0.3 percentage points, the defect classification index AUC is improved by 0.4 percentage points, and the recall rate reaches 100%, indicating that the Transformer model refinement method can effectively improve the defect detection effect of the model.

[0165] See experimental results Figure 12 In the figure, (a) is the defect image, (b) is its defect annotation, (c) is the segmentation result of the model disclosed in this invention without using Transformer for refinement, and (d) is the defect segmentation result after using Transformer for refinement.

[0166] The above description, in conjunction with the diagrams, illustrates specific embodiments of the present invention, but is not intended to limit the scope of protection of the present invention. Various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A surface defect detection method based on generative adversarial networks, characterized in that, Includes the following steps: S1. Using normal, defect-free samples, train a generative adversarial network with the following structure: (1) The network consists of an encoder, a vector quantizer, a decoder and a discriminator in sequence; (2) The encoder and decoder contain an SE-Res Net module and a SAGAN-based self-attention module. The SE-ResNet module adds the results of the input feature map x after passing through the Residual module and ShortCut operation respectively to obtain a new feature map, and then performs global pooling. Next, it performs two convolution operations with a kernel size of 1 and then uses the Sigmoid activation function to obtain the importance of each feature channel. Based on the channel importance, the features before pooling are combined. The self-attention module takes the convolution feature map x as the input feature vector and performs linear transformation using a convolutional layer with a kernel size of 1. The output g(x), f(x), h(x) of the linear transformation correspond to the Q, K, V of the Transformer model and output the self-attention result. (3) There is a vector quantizer between the encoder and the decoder. The vector quantizer uses a specific embedding space to match and replace the output features of the encoder. In the form of dictionary learning, each dimension of the encoder output features is sampled from the embedding space according to the probability and replaced, thereby obtaining the discrete feature map of the input image. (4) After the decoder, a discriminator is connected. The discriminator uses the discriminator of the PatchGAN model, which outputs a matrix in which each element is an evaluation of a local region in the image. S2. Preprocess the image to be detected according to the model's requirements for the input image; S3. Define a Transformer-based refinement method to resample the feature map index after vector quantization; Step S3 involves resampling the vector-quantized feature indexes using the Transformer method, specifically as follows: S31. Train the Transformer-based image retouching network module using normal samples. This step is implemented through the following sub-steps: S311. Serialize the embedded spatial index feature map obtained after vector quantization, randomly select 25% of the positions in the index sequence, and simultaneously mask the selected index with an 80% probability, randomly replace the current position index with a 10% probability, and keep it unchanged with a 10% probability. S312. Use the index sequence processed in step S311 as the input of the Transformer; S313. Use Transformer to predict the new index of the selected position; S314. Calculate the cross-entropy loss function between the index before occlusion of the selected position and the new index after Transformer prediction. S32. Save the parameters and trained weights of the Transformer-based refined network module obtained in step S31. S33. Using the training results of step S31, the image to be detected after preprocessing in step S2 is processed by an encoder and a vector quantizer to obtain a sampling probability map, and the index values ​​corresponding to the positions with a probability less than 0.5 are masked. S34. Input the masked index sequence from step S33 into the Transformer and use the Transformer to predict the new index at the masked position. S35. The new index sequence is input into the decoder and decoded into a reconstructed new image; S36. Perform steps S33, S34 and S35 sequentially on the masked index sequence, repeat 8 times, resample 8 times, and output 8 images. S37. The eight images obtained in step S36 are subjected to an arithmetic average operation to obtain a refined image as the reconstructed image of the original image to be detected. S4. Using the generative adversarial network trained in step S1 and the refinement method in step S3, reconstruct the image to be detected after processing in step S2. S5. Input the reconstructed image obtained in step S4 and the image to be detected in step S2 into the defect segmenter to obtain the defect segmentation map of the image to be detected. The specific processing steps of the defect segmenter described in step S5 are as follows: S51. Construct the SSIM similarity calculation formula to calculate the difference feature map between the reconstructed image and the preprocessed image to be detected in step S2. The SSIM similarity is calculated using the local SSIM index, and the calculation formula is as follows: μ x With μ G (x) represents the average gray level of each pixel in the input image and the reconstructed image, calculated using a sliding window method, respectively, and σ xG (x) is the covariance of the contrast parameters between the original input image x and the reconstructed image G(x), and C2 is a constant; S52. Adaptive threshold segmentation is performed on the feature map obtained in step S51 using the OTU algorithm. The steps include: (1) Let the segmentation threshold be t, and establish the objective function g(t)=w0×(u0-u) based on the variance between foreground and background classes. 2 +w1× (u1-u) 2 = w0 × w1 × (u1 - u0) 2 ; w0 refers to the number of foreground pixels, w1 refers to the number of background pixels, u0 refers to the foreground weighted average, u1 refers to the background weighted average, and u = u0 × w0 + u1 × w1 refers to the weighted average of the entire image. (2) Find the maximum value of the function g(t) as the threshold to distinguish between the foreground and the background, and construct the residual map; S53. Traverse the maximum area of ​​the candidate defect region in each residual map and calculate the area value that makes the defect classification F1 index the highest as the optimal segmentation threshold K. S54. Using the optimal segmentation threshold K obtained in step S53, the residual map in step S52 is binarized to obtain a defect segmentation map with accurate defect contours after removing most of the noise.

2. The surface defect detection method based on generative adversarial networks according to claim 1, characterized in that, The generative adversarial network described in step S1 has a vector quantizer with the following structure between the encoder and decoder: (1) The vector quantizer stores and updates the embedding space. It contains K D-dimensional vectors. Let ze(x) be the output of the input image after the encoder. The vector quantizer first transforms its channel count to K through a convolution operation with a kernel size of 1. Let the output of the convolution operation be ze′(x). Next, a Gumbel Softmax operation is performed on the channel dimension of ze′(x). Specifically, the sampling probability of the i-th vector in the embedding space E is... Where gi = -log (-log(ui)), ui ~ U(0,1) is Gumbel noise, πi represents the probability of the i-th vector being selected in the embedding space; τ>0 is a parameter, the smaller τ is, the closer the softmax function is to argmax, that is, the closer the z vector is to the one-hot distribution; (2) Select the index k with the largest component in the z vector, and e k z as input to the decoder q (x), i.e., z q (x)=e k , where k = argmaxj z; (3) z q (x) is used as the input to the decoder to obtain the discrete feature space of the normal sample.

Citation Information

Patent Citations

  • Photovoltaic module unsupervised defect detection method based on GAN improved algorithm

    CN111340791A

  • Surface defect segmentation method based on two-stage incremental learning

    CN115797309A

  • Defect target area image positioning segmentation method for rolling contact fatigue defect

    CN115239666A

  • Unsupervised defect detection method based on quantization auto-encoder

    CN115375604A