An industrial product surface defect detection method based on an attention generative adversarial network

By adopting an attention-based adversarial network method in the detection of surface defects of industrial products, combined with deep convolutional networks and self-attention mechanisms, the problem of lack of defect samples and insufficient feature extraction capabilities is solved, and higher detection accuracy and feature extraction capabilities are achieved.

CN115393333BActive Publication Date: 2025-06-10GUANGDONG UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211056663.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-31
Publication Date
2025-06-10
Estimated Expiration
2042-08-31

AI Technical Summary

Technical Problem

In the detection of surface defects of industrial products, the prior art lacks sufficient defect samples, resulting in poor detection results and insufficient feature extraction capabilities.

Method used

The method based on attention generation adversarial network is adopted, combined with deep convolutional network, generative adversarial network and self-attention mechanism, and surface defect detection of industrial products is carried out in the absence of defect samples. The method includes image preprocessing, model building and training, likelihood map generation in defect detection stages, image reconstruction, residual map creation, image fusion and thresholding.

Benefits of technology

It improves the accuracy of surface defect detection of industrial products, enhances the feature extraction capability of the model, and can effectively detect in the absence of sufficient defect samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393333B_ABST
    Figure CN115393333B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of defect detection methods, and proposes an industrial product surface defect detection method based on an attention generative adversarial network. The steps include: 1) image preprocessing; 2) building an attention generative adversarial network model; 3) model training stage; 4) defect detection stage. In the present invention, the problems of low detection accuracy and insufficient feature extraction in traditional defect detection methods in the absence of defect samples are broken, and an industrial product surface defect detection method based on an attention generative adversarial network is proposed, which effectively utilizes a deep convolutional network and a self-attention mechanism to improve the generative adversarial network and enhance the model's defect detection ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of defect detection methods, and particularly relates to an industrial product surface defect detection method based on an attention generative adversarial network. Background Art

[0002] In the technical field of defect detection methods, a well-trained deep neural network can detect whether there are defects on the surface of industrial products and show excellent performance. However, in the actual application of industrial product surface defect detection, there is a problem of lack of sufficient defect samples, and manual annotation of defect samples consumes a large amount of labor costs and is complicated to operate. At the same time, how to efficiently and accurately extract the features of the surface of industrial products is also an urgent problem to be solved at present.

[0003] To solve the above problems, an autoencoder can be used to learn without supervision. However, the autoencoder does not contribute much to the objective function, which may result in the defects retained in the reconstruction residuals being too weak to be recognized, and the detection cannot achieve a very perfect effect, affecting the model classification accuracy.

[0004] Therefore, the existing technology needs a defect detection method in the case of lack of defect samples to handle the problem of industrial product surface defect detection, and a method of adding a self-attention mechanism for feature extraction to improve the model feature extraction ability. Summary of the Invention

[0005] The purpose of the present invention is to provide an industrial product surface defect detection method based on an attention generative adversarial network, which can effectively utilize a deep convolutional network, a generative adversarial network, and a self-attention mechanism to complete the industrial product surface defect detection task in the scenario of lack of sufficient defect samples.

[0006] To achieve the above purpose, the present invention provides the following solution:

[0007] An industrial product surface defect detection method based on an attention generative adversarial network, comprising:

[0008] Step 1: Image preprocessing. Obtain an industrial product dataset, extract defect-free industrial product images, randomly extract multiple small pieces of each type and use them as training sets and validation sets respectively, and perform Gaussian denoising on each small piece and then normalize it.

[0009] Step 2: Build an attention generative adversarial network model.

[0010] Step 3: Model training phase. The industrial product images in the industrial product dataset are input into the attention-based generative adversarial network model constructed in Step 2 after being superimposed with Gaussian noise, and the discriminator D, the generator G, and the inverter T are trained respectively. When the number of training times reaches a certain number of iterations, the trained attention-based generative adversarial network model is obtained;

[0011] Step 4: Defect detection phase. After training, the models D, T, and G learned in Step 3 are prepared for industrial product surface detection. The detection process mainly includes likelihood map generation, image reconstruction, residual map creation, image fusion, and thresholding.

[0012] Furthermore, in the above-mentioned Step 1, the specific method of image preprocessing includes:

[0013] Step 1.1: Obtain the industrial product dataset. For each type of industrial product, multiple small patches are randomly extracted from 2 - 5 defect-free images with a size of 256*256 as training data. The size of each small patch is set to 32*32.

[0014] Step 1.2: Randomly divide the small patches into two groups, one group as the training set, and the other group as the validation set for evaluating the model performance.

[0015] Step 1.3: Each small patch is denoised with a 3*3 Gaussian filter and normalized in the interval [0, 1].

[0016] Furthermore, in the above-mentioned Step 2, the construction of the attention-based generative adversarial network model: The attention-based generative adversarial network model consists of a generator network, a discriminator network, and an inverter network. The generator network, the discriminator network, and the inverter network all adopt deep convolutional networks and incorporate self-attention mechanisms.

[0017] The input layer of the generator G is a 64-dimensional noise vector, followed by a 1024-dimensional fully connected (FC) layer. This FC layer is further reshaped into 2*2*256, where each dimension represents height (H), width (W), and depth respectively. Then all the subsequent connected layers are Deconv layers with gradually decreasing depths. Among them, the stride of all Deconv layers is 2*2, and each previous layer is upsampled. A self-attention module is inserted between layers. This module refers to SAGAN. Incorporating the self-attention mechanism can enable the model to better notice global features. The specific formula is:

[0018]

[0019]

[0020] y i =γO i +xi

[0021] Among them, β j,i represents the influence degree of the model on the i-th position when synthesizing the j-th region. The output of the attention layer is γ is a proportionality parameter, and x i is a feature map.

[0022] The output layer size is specifically set to 32*32. Batch normalization (BN) is performed on all layers except the output of the generator G, and the rectified linear unit (ReLU) activation is used.

[0023] The inverter T and the discriminator D are both constructed in reverse from G. The difference between them is the shape of the final output layer and the insertion position of the self-attention module. Except for the output layer, they all use the LeakyReLU activation.

[0024] Furthermore, in step 3, the specific method in the model training stage includes:

[0025] Step 3.1: Add Gaussian noise to the image generated by the model G to "break" the similarity between the real image and the fake image. The specific formula is as follows:

[0026]

[0027] In the formula, is the image after superimposing the noise, X is the defect-free image of the industrial product, and H represents the Gaussian noise that follows the normal distribution.

[0028] Step 3.2: Train the discriminator D. When the average likelihood of a batch of validation datasets reaches the threshold P during each optimization iteration, stop training and save the parameters of the discriminator D at this time. The specific formula is as follows:

[0029]

[0030] where n is the batch size. P is a threshold determined empirically.

[0031] Step 3.3: Train the generator G. Step 3.2 and the training of the generator G are continuously trained, and the network parameters are trained using the following loss function:

[0032]

[0033] where E represents the mathematical expectation of the real data x and the latent space data z; the G network is a generator, and P z(z) is the distribution of latent space data, generally a Gaussian distribution, to obtain a distribution Pg(x) of generated data. We hope that Pg(x) is very close to Pdata(x) to fit and approximate the true distribution. The D network is a discriminant function that needs to solve the traditional binary classification problem. Its responsibility is to effectively distinguish the true distribution and the generated distribution, that is, to measure the gap between Pg(x) and Pdata(x) and perform iterative training repeatedly.

[0034] Step 3.4: Train the inverter T. The inverter T defines an inverse mapping T(x): x → z relative to the generator G, mapping the normal image patch x back to the latent space. Specifically, the L2 norm loss function is used to optimize the inverter T, defined as:

[0035]

[0036] During the optimization process, the parameters of G are fixed, and only the parameters of E are updated.

[0037] Furthermore, in the step 4, the specific method in the defect detection stage includes:

[0038] Step 4.1: Extract a set of image patches from the given query image g, where the patch size is the same as that in the training set, which is 32*32. Given the patch x(i,j) located at (i,j) in g, the discriminator D estimates the probability that this patch belongs to the normal class. Then, the likelihood map s of g is constructed as:

[0039] S(i, j) = D(x(i, j))

[0040] Step 4.2: The generator G and the inverter T learned in step 3 are jointly used to reconstruct the input image patch x′(i, j), and the specific formula is as follows:

[0041] x′(i, j) = G(T(x(i, j)))

[0042] All the reconstructed patches x′(i, j) are reorganized to form a full-size reconstructed image f’.

[0043] Step 4.3: Perform a normalization operation on the reconstructed image f’, and the specific formula is as follows:

[0044]

[0045] where u(.) is the average value of the image, and σ(.) is the standard deviation of the image.

[0046] Step 4.4: Construct a residual map, and the specific formula is as follows:

[0047] r = |f - f r | ⊙ f err

[0048] Among them, ⊙ refers to element matrix multiplication, and f err is an error map composed of reconstruction errors calculated block by block.

[0049] Step 4.5: Fuse the likelihood map s constructed in Step 4.1 and the residual map r constructed in Step 4.4 to obtain a fused map u. The specific formula is as follows:

[0050] u(i, j) = [1(i, j) - s(i, j)] ⊙ r(i, j)

[0051] Among them, 1(i, j) - s(i, j) represents an inversion operation on the likelihood map s to make the background intensity close to zero, so that the fused map shows a cleaner background than the corresponding residual map.

[0052] Step 4.6: Perform threshold processing on the fused map. The specific formula is as follows:

[0053]

[0054] Among them, u 0 is a reference fused map obtained from defect-free samples, and t is a control constant determined empirically.

[0055] Advantages of the present invention:

[0056] The above-mentioned industrial product surface defect detection method based on an attention generative adversarial network improves the model performance by applying deep convolution to the generative adversarial network and fusing the self-attention mechanism with the generator, discriminator, and inverter of the model, thereby improving the detection accuracy of industrial product surface defects. Description of the Drawings

[0057] Figure 1 is a flowchart of an industrial product surface defect detection method based on an attention generative adversarial network of the present invention;

[0058] Figure 2 is a framework diagram of the model training process proposed by the present invention;

[0059] Figure 3 is a network structure diagram of the self-attention mechanism in the present invention. Detailed Embodiments

[0060] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below with reference to specific examples in conjunction with the accompanying drawings.

[0061] As Figure 1 shown, an industrial product surface defect detection method based on an attention generative adversarial network includes the following steps:

[0062] Step 1: Image preprocessing. Obtain the industrial product dataset. For each type of industrial product, randomly extract 5000 small patches from 2 - 5 defect-free images with a size of 256 * 256 as training data. The size of each small patch is set to 32 * 32.

[0063] Then, randomly divide the 5000 small patches into two groups, with 4500 patches as the training set and the other 500 patches as the validation set for evaluating the model performance. Denoise each small patch with a 3 * 3 Gaussian filter and normalize it in the range of [0, 1].

[0064] Step 2: Construction of the attention-based generative adversarial network model. As Figure 2 shown, the attention-based generative adversarial network model consists of a generator network, a discriminator network, and an inverter network. The generator network, discriminator network, and inverter network all adopt deep convolutional networks and incorporate self-attention mechanisms.

[0065] The input layer of the generator G is a 64-dimensional noise vector, followed by a 1024-dimensional fully connected (FC) layer. This FC layer is further reshaped into 2 * 2 * 256, where each dimension represents height (H), width (W), and depth respectively. Then all the subsequent connected layers are Deconv layers with gradually decreasing depths. Among them, the stride of all Deconv layers is 2 * 2, which upsamples each previous layer. Insert a self-attention module between the penultimate layer and the output layer. This module refers to SAGAN. As Figure 3 shown, adding the self-attention mechanism can enable the model to better notice global features. The specific formula is as follows:

[0066]

[0067]

[0068] y i = γo i + x i

[0069] where, β j,i represents the influence degree of the model on the i-th position when synthesizing the j-th region. The output of the attention layer is γ is the proportionality parameter, and x i is the feature map.

[0070] The size of the output layer is specifically set to 32 * 32. Perform batch normalization (BN) and use the rectified linear unit (ReLU) activation for all layers except the output of the generator G.

[0071] Both the inverter T and the discriminator D are constructed inversely from G. The difference between them lies in the shape of the final output layer. Also, the self-attention modules of the inverter T and the discriminator D are inserted between the penultimate layer and the output layer. Except for the output layer, they all use the LeakyReLU activation.

[0072] Step 3: Model training phase. Gaussian noise is added to the industrial product images in the industrial product dataset. The specific formula is as follows:

[0073]

[0074] where is the image after adding noise, X is the defect-free image of the industrial product, and H represents Gaussian noise that follows a normal distribution.

[0075] Then is input into the attention-based generative adversarial network model constructed in Step 2 for training, and the discriminator D, the generator G, and the inverter T are trained respectively. When the number of training times reaches a certain number of iterations, the trained attention-based generative adversarial network model is obtained;

[0076] Among them, when training the discriminator D, when the average likelihood of a batch of validation datasets reaches the threshold P at each optimization iteration, stop training and save the parameters of the discriminator D at this time. The specific formula is as follows:

[0077]

[0078] where n is the batch size. P is a threshold determined empirically.

[0079] Then train the generator G. The discriminator D and the generator G are continuously trained, and the network parameters are trained using the following loss function:

[0080]

[0081] where E represents the mathematical expectation of the real data x and the latent space data z; the G network is a generator, P z (z) is the latent space data distribution, usually a Gaussian distribution, to obtain a generated data distribution Pg(x). We hope that Pg(x) is very close to Pdata(x) to fit and approximate the real distribution. The D network is a discriminant function that needs to solve the traditional binary classification problem. Its responsibility is to effectively distinguish the real distribution and the generated distribution, that is, to measure the gap between Pg(x) and Pdata(x), and perform repeated iterative training.

[0082] When training the inverter T, the inverter T defines an inverse mapping T(x) relative to the generator G: x → z, which maps the normal image patch x back to the latent space. Specifically, the L2 norm loss function is used to optimize the inverter T, which is defined as:

[0083]

[0084] During the optimization process, the parameters of G are fixed and only the parameters of E are updated.

[0085] Step 4, Defect detection phase. After training, the models D, T, and G learned in Step 3 are prepared for industrial product surface detection.

[0086] First, a set of image patches are extracted from the given query image g, where the patch size is the same as that in the training set, which is 32*32. Given the patch x(i, j) located at (i, j) in g, the discriminator D estimates the probability that this patch belongs to the normal class. Then, the likelihood map s of g is constructed as:

[0087] s(i, j) = D(x(i, j))

[0088] Then, the learned generator G and inverter T are jointly used to reconstruct the input image patch x'(i, j), and the specific formula is as follows:

[0089] x'(i, j) = G(T(x(i, j)))

[0090] All the reconstructed patches x'(i, j) are reorganized to form a full-size reconstructed image f'. Then, a normalization operation is performed on the reconstructed image f', and the specific formula is as follows:

[0091]

[0092] where u(.) is the average value of the image and σ(.) is the standard deviation of the image.

[0093] Next, a residual map is constructed based on the difference between the original image and the reconstructed image, and the specific formula is as follows:

[0094] r = |f - f r | ⊙ f err

[0095] where ⊙ refers to element-wise matrix multiplication, and ferr is an error map composed of the reconstruction errors calculated block by block.

[0096] Then, the constructed likelihood map s and the constructed residual map r are fused to obtain a fused map u. The specific formula is as follows:

[0097] u(i, j) = [1(i, j) - s(i, j)] ⊙ r(i, j)

[0098] where 1(i, j) - s(i, j) represents an inversion operation on the likelihood map s to make the background intensity close to zero. This makes the fused map show a cleaner background than the corresponding residual map.

[0099] Finally, threshold processing is performed on the fused graph. The specific formula is as follows:

[0100]

[0101] Among them, uO is the reference fused graph obtained from defect-free samples, and t is a control constant determined empirically.

Claims

1. An industrial product surface defect detection method based on an attention generative adversarial network, characterized in that it includes the following steps: 1) Image preprocessing: Obtain an industrial product dataset, extract defect-free industrial product images, randomly extract multiple small patches of each type and use them as the training set and validation set respectively, and perform Gaussian denoising on each small patch and then normalize it; 2) Construction of an attention generative adversarial network model; The construction of the attention generative adversarial network model is specifically as follows: The attention generative adversarial network model consists of a generator network, a discriminator network, and an inverter network. The generator network, discriminator network, and inverter network all adopt deep convolutional networks and incorporate a self-attention mechanism; The input layer of the generator G is a 64-dimensional noise vector, followed by a 1024-dimensional fully connected (FC) layer, which is further reshaped into 2*2*256, where each dimension represents height (H), width (W), and depth respectively; Then all the connected layers are Deconv layers with gradually decreasing depth. Among them, the stride of all Deconv layers is 2*2, and each previous layer is upsampled; A self-attention module is inserted between layers. This module refers to SAGAN. Adding a self-attention mechanism can enable the model to better notice global features. The specific formula is as follows: y i = γo i + x i Among them, β j,i represents the influence degree of the model on the i-th position when synthesizing the j-th region. The output of the attention layer is γ is a proportionality parameter, and x i is a feature map; The output layer size is specifically set to 32*32. Batch normalization (BN) is performed on all layers except the output of the generator G and the rectified linear unit (ReLU) activation is used; The inverter T and the discriminator D are both constructed in reverse from G; The difference between them is the shape of the final output layer and the insertion position of the self-attention module. Except for the output layer, they all use LeakyReLU activation; 3) Model training stage: After superimposing Gaussian noise on the industrial product images in the industrial product dataset, input them into the attention generative adversarial network model constructed in step 2 for training. The discriminator D, generator G, and inverter T are trained respectively. When the number of training times reaches a certain number of iterations, a trained attention generative adversarial network model is obtained; 4) Defect detection stage: After training, the models D, T, and G learned in step 3 are prepared for industrial product surface detection. The detection process mainly includes likelihood map generation, image reconstruction, residual map creation, image fusion, and thresholding.

2. The industrial product surface defect detection method based on an attention generative adversarial network according to claim 1, characterized in that in step 1): The specific steps of the image preprocessing are as follows: 1) Obtain an industrial product dataset. For each type of industrial product, randomly extract multiple small patches from 2-5 defect-free images with a size of 256*256 as training data, and the size of each small patch is set to 32*32; 2) Randomly divide the small patches into two groups, one group as the training set and the other group as the validation set for evaluating the model performance; 3) Denoise each small patch with a 3*3 Gaussian filter and normalize it in the interval [0,1].

3. The industrial product surface defect detection method based on an attention generative adversarial network according to claim 1, characterized in that In step 3), the specific steps of the model training stage are as follows: 1) Add Gaussian noise to the image generated by model G. The specific formula is as follows: In the formula, is the image after adding the superimposed noise, X is the defect-free image of industrial products, and H represents Gaussian noise that follows a normal distribution; 2) Train discriminator D. When the average likelihood of a batch of validation datasets reaches the threshold P at each optimization iteration, stop training and save the parameters of discriminator D at this time. The specific formula is as follows: where n is the batch size and P is the threshold determined empirically; 3) Train generator G. Steps 3.2 and the training of generator G are continuously carried out. The network parameters are trained using the following loss function: where E represents the mathematical expectation of the real data x and the latent space data z; the G network is a generator, and P z (z) is the latent space data distribution, generally a Gaussian distribution, to obtain a generated data distribution Pg(x). We hope that Pg(x) is very close to Pdata(x) to fit and approximate the real distribution. The D network is a discriminant function that needs to solve the traditional binary classification problem. Its responsibility is to effectively distinguish the real distribution and the generated distribution, that is, to measure the gap between Pg(x) and Pdata(x), and perform repeated iterative training; 4) Train inverter T. Inverter T defines an inverse mapping T(x): x → z relative to generator G, mapping the normal image patch x back to the latent space. Specifically, the L2 norm loss function is used to optimize inverter T, defined as: During the optimization process, the parameters of G are fixed, and only the parameters of E are updated.

4. A method for detecting industrial product surface defects based on an attention generative adversarial network according to claim 1, characterized in that In step 4), the specific steps of the defect detection stage are as follows: 1) Extract a set of image patches from the given query image g. The size of the patches is the same as that in the training set, which is 32*32. Given the patch x(i,j) located at (i,j) in g, discriminator D estimates the probability that this patch belongs to the normal class. Then, the likelihood map s of g is constructed as: s(i,j) = D(x(i,j)) 2) The generator G and inverter T learned in step 3 are jointly used to reconstruct the input image patch x'(i,j). The specific formula is as follows: x'(i,j) = G(T(x(i,j))) All the reconstructed patches x'(i,j) are reorganized to form a full-size reconstructed image f'; 3) Perform a normalization operation on the reconstructed image f'. The specific formula is as follows: where u(.) is the average value of the image and σ(.) is the standard deviation of the image; 4) Construct a residual map. The specific formula is as follows: r = |f - f r | ⊙ f err where ⊙ denotes element-wise matrix multiplication, and f err is an error map composed of reconstruction errors calculated block by block; 5) Fuse the likelihood map s constructed in step 4.1 and the residual map r constructed in step 4.4 to obtain a fused map u. The specific formula is as follows: u(i,j) = [1(i,j) - s(i,j)] ⊙ r(i,j) where 1(i,j) - s(i,j) represents an inversion operation on the likelihood map s to make the background intensity close to zero, so that the fused map shows a cleaner background than the corresponding residual map; 6) Perform threshold processing on the fused map. The specific formula is as follows: where u 0 is a reference fusion graph obtained from defect-free samples, and t is a control constant determined empirically.

Citation Information

Patent Citations

  • Defect detection method based on generative adversarial network and attention

    CN114943694A