Small target flaw image detection method and system for high-end equipment product
By combining AttentionGAN and Pix2Pix models, an attention mechanism is introduced to conduct defect detection of high-end equipment products, generate real defect images and improve the defect repair effect, solving the problems of unreality and incomplete repair of defects in defect detection, and achieving higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202510873900.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing defect detection methods have unreal defect generation in high-end equipment products, and are difficult to fully cover all situations, and the image edge fits unnaturally. The defect repair process of traditional CycleGAN lacks targetedness, resulting in poor detection results.
The combination of AttentionGAN and Pix2Pix model is adopted to introduce an attention mechanism to weight the defect areas in the image, generate defective images through unsupervised learning, and repair defects through supervised training. The accuracy and robustness of defect detection are improved by using a multi-scale discriminator.
The generated defect images are more realistic, the defect repair effect is significantly improved, and the detection accuracy and robustness are significantly improved. It can detect unseen defects more accurately and reduce the complexity of defect image synthesis.
Smart Images

Figure CN120374626A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing. More specifically, the present invention relates to a small target defect image detection method and system for high-end equipment products. Background Art
[0002] In the industrial field, defect detection on the production line is a key link to ensure product quality and has become an indispensable part of the intelligent production line. Most traditional defect detection methods rely on manual screening, which is not only time-consuming and laborious, but also difficult to ensure the stability and consistency of product quality. In recent years, with the wide application of deep learning technology in the industrial field, the production process has gradually achieved automation and intelligence. Many deep learning-based automatic defect detection methods have been widely applied in actual production, not only effectively assisting manual screening, but also successfully replacing manual detection in many scenarios, greatly improving the detection efficiency and accuracy.
[0003] In industrial defect detection, it is crucial to ensure that the dataset contains a sufficient variety of defect samples. However, relying solely on sampling from the production line is both inefficient and insufficient. Therefore, image sample augmentation methods are needed. Although image enhancement techniques are highly versatile and easy to use, they also have some limitations. First, the number of image enhancement methods is limited, and the scale of sample augmentation is restricted. Second, image enhancement only modifies the surface and cannot introduce sufficient diversity at the image content level. The changes in the dataset are still limited to the original content. To solve these problems, samples can be generated through a generative adversarial network (GAN) to increase the randomness and diversity of the augmented samples.
[0004] With the gradual popularization of GAN technology, this technology has also begun to be gradually introduced into the field of defect detection. One of the common applications is to generate missing defect samples through GAN to solve the problems of insufficient samples and high acquisition costs. In addition, GAN can also be used to magnify defects in images to improve the detection effect. The common approach is to first generate small defect blocks through GAN and then paste them onto the original image. However, this method has the following problems: the design process is cumbersome, it is difficult to cover all situations comprehensively, and the fitting at the image edges is not natural, and the overall effect is not realistic enough.
[0005] In "Simple Generation of Defect Images Based on GAN and Its Application to Improve Defect Recognition" (Niu S, Li B, Wang X, et al., 2020), CycleGAN was used to generate defects, and its loop characteristics were utilized for image restoration. The defect regions were extracted by comparing the images before and after restoration. However, this method still has deficiencies: during the defect restoration process, the consistency loss and the adversarial loss are located in the forward and backward training stages respectively, lacking effective combination in the same training stage, resulting in a lack of pertinence in the training of the restoration stage, thus affecting the effect of defect extraction.
[0006] In view of this, the present invention proposes a method and system for detecting small target defect images of high-end equipment products to solve the above problems. Summary of the Invention
[0007] The purpose of the present invention is to provide a method and system for detecting small target defect images of high-end equipment products, aiming to reduce the complexity of defect image synthesis and solve the problem of detecting unseen defects. Compared with the traditional naive CycleGAN, the present invention can obtain better performance and is comparable to the supervised UNET in terms of effect; in addition, by introducing an attention mechanism to weight the defect regions in the image, the present invention can more accurately generate and detect unseen defects.
[0008] To achieve the above object, the present invention provides the following technical solution: A method for detecting small target defect images of high-end equipment products, including the following steps:
[0009] S1: Establish and train an AttentionGAN model, specifically including: Initialize a defect generator G, a defect restorer F, a defect-free image discriminator D, and a defective image discriminator H. Randomly select a defect-free image x and a defective image y to generate a fake defective image G(x), a fake defect-free image F(y), and an attention mask. Calculate the adversarial losses LossF1 and LossB1, update the parameters of the defect-free image discriminator D and the defective image discriminator H. Reconstruct the defect-free image F(G(x)) and the defective image G(F(y)), calculate the cycle consistency loss and the attention consistency loss, and update the parameters of the defect generator G and the defect restorer F until the training stop condition is reached;
[0010] S2: Establish and train the Pix2Pix model, specifically including: initializing the multi-scale discriminator A, randomly selecting a flawless image x to generate a fake defective image G(x), randomly selecting Gaussian noise z and generating a restored image F(G(x), z), calculating the adversarial loss LossF3 and updating the parameters of the multi-scale discriminator A, randomly selecting Gaussian noise z and generating a restored image F(G(x), z), calculating the consistency loss LossF4, and updating the parameters of the next defect repairer F until the training stop condition is reached;
[0011] S3: Defect detection, specifically including: randomly selecting a defective image y to generate an image F(y), pixel-by-pixel comparing the defective image y with the restored flawless image F(y) to obtain a binary image of the defect location and shape.
[0012] Further, in the said S1, the following steps are included:
[0013] a. Randomly select a flawless image x and a defective image y from the flawless image dataset X and the defective image dataset Y respectively;
[0014] b. Use the defect generator G to convert the flawless image x into a fake defective image G(x), and through the attention mask and , and at the same time use the defect repairer F to convert the defective image y into a fake flawless image F(y), and dynamically adjust the attention weights and ;
[0015] c. Calculate the defective true rate D(x) of the flawless image x and the true rate D(F(y)) of the converted fake flawless image F(y) through the flawless image discriminator D, and strengthen the attention to the flawless area through the attention mechanism, calculate the adversarial loss LossF1, and optimize and update the parameters of the flawless image discriminator D;
[0016] Calculate the defective true rate H(y) of the defective image y and the true rate H(G(x)) of the converted fake defective image G(x) through the defective image discriminator H, and enhance the discrimination ability of the defective area through the attention mechanism, calculate the adversarial loss LossB1, and optimize and update the parameters of the defective image discriminator H;
[0017] d. When training the generator, input the flawless image x and the defective image y; use the defect generator G to process the flawless image x, generate foreground and background attention masks through the parameter-sharing encoder and the attention module, and combine the content mask to generate a fake defective image G(x), and at the same time use the defect repairer F to repair the defective image through the generated attention masks and Focus on this defective area to generate a defect-free image F(y);
[0018] Then, calculate and optimize the following losses: the loss of the defect-free image discriminator D: calculate D(x) and D(F(y)) to obtain the adversarial loss of the defect-free image; the loss of the defective image discriminator H: calculate H(y) and H(G(x)) to obtain the defective image adversarial loss;
[0019] In addition, input the fake defect-free image F(y) into the generator G to generate a reconstructed defective image G(F(y)), and at the same time input the fake defective image G(x) into the restorer F to generate a reconstructed defect-free image F(G(x));
[0020] By jointly optimizing the adversarial loss, the consistency loss, and the attention consistency loss, improve the ability of the generator and the restorer to capture details in key areas and the accuracy of image reconstruction.
[0021] e. Calculate the difference between the generated defect-free image F(G(x)) and the original defect-free image x to obtain the consistency loss LossF2; at the same time, calculate the difference between the generated defective image G(F(y)) and the original defective image y to obtain the consistency loss LossB2;
[0022] In addition, use the attention consistency loss to constrain the generated attention mask to ensure that the attention mask can accurately focus on the key areas in the image; combine the adversarial losses LossF1, LossB1, the consistency losses LossF2, LossB2, and the attention consistency loss, and further improve the accuracy and detail restoration ability of the defect generator G and the restorer F in the process of generating and restoring images through joint optimization;
[0023] f. If the training stop condition is reached, enter S2; if the training stop condition is not reached, repeat steps a to e.
[0024] Furthermore, in the S2, the following steps are included:
[0025] g. Randomly select a fake defective image G(x) from the defective image library generated in the unsupervised stage, and form paired data (x, G(x)) with the original defect-free image x, and apply a random affine transformation to G(x);
[0026] h. Randomly select a Gaussian noise z, the noise dimension matches the encoder feature layer, superimpose the noise z in the intermediate feature layer of the encoder, and dynamically adjust the noise weight through a learnable attention gating mechanism:
[0027] ,
[0028] Among them, and are learnable parameters, is the Sigmoid function, and ⊙ represents element-wise multiplication;
[0029] j. Use U-Net++ as the restorer backbone network to enhance the multi-scale feature fusion ability through dense skip connections; embed a dynamic convolution module in the decoder layer to generate adaptive convolution kernel weights according to the noise z, and the formula is:
[0030] ,
[0031] Among them, represents the dynamically generated convolution kernel weight matrix, MLP is the multi-layer perceptron, Conv is the dynamic convolution operation, represents the feature map input to the dynamic convolution layer;
[0032] The restored image F(G(x),z)F(G(x),z) needs to satisfy to ensure pixel-level consistency;
[0033] k. Multi-scale discriminator A and adversarial loss optimization: The discriminator A uses a multi-scale PatchGAN to discriminate local texture consistency at three scales of 32×32, 64×64, and 128×128 respectively; the input is an image pair with channel concatenation (such as [F(G(x),z),G(x)]), and the matching probability at each scale is output ;
[0034] l. Improved loss function:
[0035] ,
[0036] In the formula, E represents the expected value of the joint distribution of the input variables x and z, and LPIPS represents the perceptual loss: it extracts multi-level features through a pre-trained VGG16 network and calculates the feature space distance, and the formula is:
[0037] ,
[0038] Among them, represents the restored image, represents the original flawless image, represents the feature map of the l-th layer of VGG16; hyperparameter: γ = 0.1, which is used to balance pixel-level and semantic-level consistency;
[0039] m. If the training stop condition is reached, go to step S3; if the training stop condition is not reached, repeat steps g to l.
[0040] Further, in the step S3, the following steps are included:
[0041] n. Given a randomly selected defective image y, use the defect repairer F to convert it into a fake defect-free image F(y);
[0042] o. Compare the defective image y with the repaired defect-free image F(y) pixel by pixel to generate a binary image, marking the positions and shapes of the defects.
[0043] Further, the calculation formulas for the adversarial losses LossF1 and LossB1 are as follows:
[0044] ,
[0045] ,
[0046] where Pdata represents the distribution of the dataset, and the optimization objectives of F and D are respectively .
[0047] Further, the optimization objective is:
[0048] ,
[0049] ,
[0050] where is a hyperparameter for balancing the adversarial losses LossF1, LossB1 and the consistency losses LossF2, LossB2.
[0051] Further, the calculation formulas for the consistency losses LossF2 and LossB2 are as follows:
[0052] ,
[0053] .
[0054] Further, the calculation formula for the adversarial loss LossF3 is as follows:
[0055] ,
[0056] where λ1 = 1.0, λ2 = 0.5, λ3 = 0.25 are scale weight coefficients for balancing the discriminative contributions at different resolutions.
[0057] Further, in the step S3, a binary image marking the positions and shapes of the defects is obtained by comparing y and F(y). Among them, if the absolute value of the difference between the corresponding pixels of y and F(y) is greater than the threshold, the pixel is marked as a defective area pixel.
[0058] The present invention further includes: a small target defect image detection system for high-end equipment products, including:
[0059] An unsupervised model establishment and training module, which is used to establish and train an AttentionGAN model. Specifically, it includes: initializing a defect generator G, a defect repairer F, a defect-free image discriminator D, and a defective image discriminator H, randomly selecting a defect-free image x and a defective image y to generate a fake defective image G(x), a fake defect-free image F(y), and an attention mask, calculating the adversarial loss LossF1 and LossB1, updating the parameters of the defect-free image discriminator D and the defective image discriminator H, reconstructing the defect-free image F(G(x)) and the defective image G(F(y)), calculating the cycle consistency loss and the attention consistency loss, and updating the parameters of the defect generator G and the defect repairer F until the training stop condition is reached;
[0060] A supervised model establishment and training module, which is used to establish and train a Pix2Pix model. Specifically, it includes: initializing a multi-scale discriminator A, randomly selecting a defect-free image x to generate a fake defective image G(x), randomly selecting Gaussian noise z and generating a repaired image F(G(x),z), calculating the adversarial loss LossF3 and updating the parameters of the multi-scale discriminator A, randomly selecting Gaussian noise z and generating a repaired image F(G(x),z), calculating the consistency loss LossF4, and updating the parameters of the next defect repairer F until the training stop condition is reached;
[0061] A defect detection module, which is used for defect detection. Specifically, it includes: randomly selecting a defective image y to generate an image F(y), and pixel-by-pixel comparing the defective image y with the repaired defect-free image F(y) to obtain a binary image of the defect position and shape.
[0062] The technical effects and advantages of the small target defect image detection method and system for high-end equipment products of the present invention:
[0063] 1. The present invention combines the training strategies of AttentionGAN and Pix2Pix. By introducing the attention mechanism, the generator can more accurately focus on the key defect areas in the image, the generated defect images are more realistic, and the repairer can also better focus on the defect areas that need to be emphasized during the repair process; the defect generator generates defective images through unsupervised learning, and the repairer conducts supervised training through these generated defective images, so as to achieve more accurate defect repair and detection;
[0064] 2. The present invention introduces an attention mechanism in the repair stage. This mechanism enhances the repair effect of defects during the repair process by weighting important regions in the image. Specifically, the attention mechanism network dynamically adjusts the attention to different regions when processing the image, ensuring that the network pays more attention to the regions related to defects during repair. This improvement effectively enhances the quality of defect repair, and compared with traditional methods, it can significantly improve the accuracy and robustness of defect detection.
[0065] 3. The present invention can directly generate complete defect samples by adopting the AttentionGAN network architecture, avoiding the artificial synthesis process, and at the same time, the generated samples are more realistic. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is the overall flowchart of the present invention;
[0067] Figure 2 is the schematic structural diagram of the unsupervised learning stage of the present invention;
[0068] Figure 3 is the schematic structural diagram of the supervised learning stage of the present invention;
[0069] Figure 4 is the schematic structural diagram of the defect detection of the present invention;
[0070] Figure 5 are the defect detection result diagrams of the metal nut images (a), (b), and (c) in MVTec AD; the detection result image sequence of each image is the original image and the result of the present invention;
[0071] Figure 6 are the defect detection result diagrams of the transistor images (a), (b), and (c) in MVTec AD; the detection result image sequence of each image is the original image and the result of the present invention;
[0072] Figure 7 is the flowchart of the present invention;
[0073] Figure 8 is the schematic structural diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0075] Embodiment 1
[0076] Please refer to Figures 1 to 7 shown in the figure. The small target defect image detection method for high-end equipment products described in this embodiment includes the following steps:
[0077] S1: Establish and train an AttentionGAN model, specifically including: initializing a defect generator G, a defect repairer F, a defect-free image discriminator D, and a defective image discriminator H, randomly selecting a defect-free image x and a defective image y to generate a fake defective image G(x), a fake defect-free image F(y), and an attention mask, calculating the adversarial loss LossF1 and LossB1, updating the parameters of the defect-free image discriminator D and the defective image discriminator H, reconstructing the defect-free image F(G(x)) and the defective image G(F(y)), calculating the cycle consistency loss and the attention consistency loss, and updating the parameters of the defect generator G and the defect repairer F until the training stop condition is reached;
[0078] S2: Establish and train a Pix2Pix model, specifically including: initializing a multi-scale discriminator A, randomly selecting a defect-free image x to generate a fake defective image G(x), randomly selecting Gaussian noise z and generating a repaired image F(G(x),z), calculating the adversarial loss LossF3 and updating the parameters of the multi-scale discriminator A, randomly selecting Gaussian noise z and generating a repaired image F(G(x),z), calculating the consistency loss LossF4, and updating the parameters of the next defect repairer F until the training stop condition is reached;
[0079] S3: Defect detection, specifically including: randomly selecting a defective image y to generate an image F(y), and pixel-by-pixel comparing the defective image y with the repaired defect-free image F(y) to obtain a binary image of the defect position and shape.
[0080] Furthermore, in the said S1, it includes the following steps:
[0081] a. Randomly select a defect-free image x and a defective image y from the defect-free image dataset X and the defective image dataset Y respectively;
[0082] b. Use the defect generator G to convert the defect-free image x into a fake defective image G(x), and through the attention mask and , and at the same time use the defect repairer F to convert the defective image y into a fake defect-free image F(y), and dynamically adjust the attention weights and ;
[0083] c. Calculate the true defect rate D(x) of the defect-free image x and the true rate D(F(y)) of the transformed false defect-free image F(y) through the defect-free image discriminator D, and strengthen the attention to the defect-free area through the attention mechanism, calculate the adversarial loss LossF1, and optimize and update the parameters of the defect-free image discriminator D;
[0084] Calculate the true defect rate H(y) of the defective image y and the true rate H(G(x)) of the transformed false defective image G(x) through the defective image discriminator H, and enhance the discrimination ability of the defective area through the attention mechanism, calculate the adversarial loss LossB1, and optimize and update the parameters of the defective image discriminator H;
[0085] d. When training the generator, input the defect-free image x and the defective image y; use the defect generator G to process the defect-free image x, generate foreground and background attention masks through the parameter-sharing encoder and attention module, and generate a false defective image G(x) in combination with the content mask. At the same time, use the defect repairer F to repair the defective image through the generated attention mask and focus on this defective area to generate a defect-free image F(y);
[0086] Then, calculate and optimize the following losses: the loss of the defect-free image discriminator D: calculate D(x) and D(F(y)) to obtain the adversarial loss of the defect-free image; the loss of the defective image discriminator H: calculate H(y) and H(G(x)) to obtain the adversarial loss of the defective image;
[0087] In addition, input the false defect-free image F(y) into the generator G to generate a reconstructed defective image G(F(y)), and at the same time input the false defective image G(x) into the repairer F to generate a reconstructed defect-free image F(G(x));
[0088] By jointly optimizing the adversarial loss, the consistency loss, and the attention consistency loss, the accuracy of the generator and the repairer in capturing details of key areas and image reconstruction is improved.
[0089] e. Calculate the difference between the generated defect-free image F(G(x)) and the original defect-free image x to obtain the consistency loss LossF2; at the same time, calculate the difference between the generated defective image G(F(y)) and the original defective image y to obtain the consistency loss LossB2;
[0090] In addition, the generated attention mask is constrained using the attention consistency loss to ensure that the attention mask can accurately focus on the key regions in the image; by combining the adversarial losses LossF1, LossB1, the consistency losses LossF2, LossB2, and the attention consistency loss, the precision and detail restoration ability of the defect generator G and the restorer F in the process of generating and restoring images are further improved through joint optimization;
[0091] f. If the training stop condition is reached, go to step two S2; if the training stop condition is not reached, repeat steps a to e;
[0092] It should be noted that S1 is the unsupervised learning stage. By establishing and training the AttentionGAN model, as Figure 2 shown, the AttentionGAN model defines two generators G and F, which are responsible for defect generation and defect restoration respectively, and an attention module is introduced in the generator and the restorer to dynamically adjust the attention to the defect regions in the image; the model realizes image conversion through the forward mapping flow (X→Y→X) and the backward mapping flow (Y→X→Y), and conducts joint optimization by combining the adversarial loss, the consistency loss, and the attention consistency loss, so as to generate more realistic defect images and achieve more accurate defect restoration and detection.
[0093] Furthermore, in the said S2, the following steps are included:
[0094] g. Randomly select a fake defective image G(x) from the defective image library generated in the unsupervised stage, and form paired data (x, G(x)) with the original defect-free image x, and apply a random affine transformation (rotation ±5°, translation ±10%) to G(x) to simulate the position and angle changes of objects in the industrial scenario and improve the robustness of the model to the target pose;
[0095] h. Randomly select a Gaussian noise z, the noise dimension matches the encoder feature layer, and superimpose the noise z in the intermediate feature layer of the encoder, and dynamically adjust the noise weight through a learnable attention gating mechanism:
[0096] ,
[0097] where and are learnable parameters, is the Sigmoid function, and ⊙ represents element-wise multiplication;
[0098] j. Use U-Net++ as the backbone network of the restorer to enhance the multi-scale feature fusion ability through dense skip connections; embed a dynamic convolution module in the decoder layer to generate an adaptive convolution kernel weight according to the noise z, and the formula is:
[0099] ,
[0100] Among them, represents the dynamically generated convolutional kernel weight matrix, which is generated by a multi-layer perceptron (MLP) based on the input Gaussian noise z. MLP is the multi-layer perceptron, and Conv is the dynamic convolution operation; represents the feature map input to the dynamic convolution layer, that is, the input feature of the current layer of the decoder;
[0101] The repaired image F(G(x), z) needs to satisfy , ensuring pixel-level consistency, where is the threshold for controlling the pixel-level difference between the generated image and the original image;
[0102] k. Multi-scale discriminator A and adversarial loss optimization: The discriminator A uses a multi-scale PatchGAN to discriminate local texture consistency at three scales of 32×32, 64×64, and 128×128 respectively; the input is an image pair with channel concatenation (such as [F(G(x), z), G(x)]), and the matching probability at each scale is output ;
[0103] l. Improved loss function:
[0104] ,
[0105] In the formula, E represents the expected value of the joint distribution of the input variables x and the noise z, and LPIPS represents the perceptual loss: it extracts multi-level features through a pre-trained VGG16 network and calculates the feature space distance. The formula is:
[0106] ,
[0107] Among them, represents the repaired image, represents the original flawless image, represents the feature map of the l-th layer of VGG16; Hyperparameters: γ = 0.1, which is used to balance pixel-level and semantic-level consistency, , , respectively represent the number of channels, height, and width of the feature map of the th layer.
[0108] m. If the training stop condition is reached, go to step S3; if the training stop condition is not reached, repeat steps g to l.
[0109] It should be noted that in the supervised learning stage S2, based on the defect generator G trained in the unsupervised stage S1, the defect-free image x is converted into a highly realistic defective image G(x), forming paired data (x, G(x)). To improve the adaptability of the model to complex interferences in industrial scenarios, a feature space noise injection strategy is introduced in this stage: multi-dimensional Gaussian noise z is superimposed on the intermediate feature layer of the encoder, and the contribution weight of the noise to the feature map is dynamically adjusted through a learnable attention gating mechanism to avoid directly damaging the spatial structure of the input image; the restorer F adopts the U-Net++ network architecture, and a dynamic convolution module is embedded in the decoder to adaptively adjust the convolution kernel weight according to the noise z to enhance the noise robustness; the generated restored image F(G(x),z) is constrained by the multi-scale discriminator A: the discriminator A discriminates the local texture consistency at three scales of 32×32, 64×64, and 128×128. The input is an image pair spliced by channels, and the multi-scale adversarial loss is used to optimize the game process between the restorer and the discriminator;
[0110] The goal of the restorer F is to minimize the joint loss function, including the adversarial loss ( )and the improved consistency loss , where, ( )combines the pixel-level L1 loss and the LPIPS perceptual loss, extracts multi-level features through the pre-trained VGG network, and constrains the semantic consistency; to balance the optimization objectives of different training stages, during the training process, if the F1 value of the validation set has not improved for 10 consecutive rounds, or the total number of rounds reaches 200, the early stopping mechanism is triggered; finally, the model parameters with the highest F1 value of the validation set are selected to ensure the optimal performance of the restorer F in noise environments and small target defect detection;
[0111] Therefore, in this stage, through the diverse defect samples generated based on the unsupervised learning S1, combined with the noise robustness enhancement design and the multi-scale conditional generative adversarial network (CGAN) optimization strategy, the generalization ability and detection accuracy of the restorer F are further improved.
[0112] Furthermore, in the S3, the following steps are included:
[0113] n. Given a random defective image y, use the defect restorer F to convert it into a fake defect-free image F(y);
[0114] o. Compare the defective image y and the restored defect-free image F(y) pixel by pixel to generate a binary image, marking the positions and shapes of the defects;
[0115] It should be noted that after the model training is completed, F in the model can be used for defect detection. For example Figure 4As shown, for the defective image y, it is repaired by F to obtain F(y); then y and F(y) are compared pixel by pixel. If the absolute value of the difference between the corresponding pixels of y and F(y) is greater than the threshold, they are marked as pixels in the defective area. Finally, the binary image D&T(y, F(y)) is obtained, where D&T is the operation of comparing the pixel-by-pixel value with the threshold; this binary image represents the position and shape of the defects in the image.
[0116] Further, in step c, the calculation formulas for the adversarial losses LossF1 and LossB1 are as follows:
[0117] ,
[0118] ,
[0119] where Pdata represents the distribution of the dataset, and the optimization objectives of F and D are respectively .
[0120] Further, the optimization objective is:
[0121] ,
[0122] , where is a hyperparameter for balancing the adversarial losses LossF1, LossB1 and the consistency losses LossF2, LossB2.
[0123] Further, in step e, the calculation formulas for the consistency losses LossF2 and LossB2 are as follows:
[0124] ,
[0125] .
[0126] Further, in step k, the calculation formula for the adversarial loss LossF3 is as follows:
[0127] ,
[0128] where λ1 = 1.0, λ2 = 0.5, λ3 = 0.25 are scale weight coefficients for balancing the discriminative contributions at different resolutions.
[0129] Further, in S3, a binary image marking the position and shape of the defects is obtained by comparing y and F(y). Among them, if the absolute value of the difference between the corresponding pixels of y and F(y) is greater than the threshold, they are marked as pixels in the defective area.
[0130] Experimental Example 1: This experimental example uses the MVTec AD dataset, which contains 5,354 high-resolution images of 5 textures and 10 objects from different fields. There are a total of 73 types of defects, and pixel-level accurate labels are provided for each defective image.
[0131] In this experimental example, the commonly used IOU (Intersection over Union), pixel-level accuracy, pixel-level recall, and pixel-level F1 value in semantic segmentation are used as evaluation metrics; as shown in the following formula, TP, FP, and FN represent the number of correctly detected pixels, incorrectly detected pixels, and undetected defect pixels respectively; in subsequent experiments, these values are averaged over the entire test set.
[0132] IOU = TP / (TP + FP + FN),
[0133] Pre = TP / (TP + FP),
[0134] Rec = TP / (TP + FN),
[0135] F1 = 2 × Pre × Rec / (Pre + Rec),
[0136] The results of each evaluation metric of the method of the present invention, CycleGAN, UNET, and CFE are shown in Table 1.
[0137] Table 1 Results of evaluation metrics on MVTec AD
[0138] Method IOU(%) Pre(%) Rec(%) F1(%) CFE - 87.54 85.63 85.12 UNET 45.73 89.25 58.33 68.86 CycleGAN 27.18 60.23 46.40 48.45 Ours 36.71 84.12 48.61 57.26
[0139] It can be obtained from the experiment that the semi-supervised method proposed by the present invention, which combines the AttentionGAN and Pix2Pix training strategies, reduces the complexity of defective image synthesis and can improve the defective detection effect. It does not require labeling real data for training and also has good adaptability to unseen defects.
[0140] Example 2
[0141] Please refer to Figure 8 As shown, the small target defective image detection system for high-end equipment products described in this embodiment includes:
[0142] An unsupervised model building and training module, which is used to build and train the AttentionGAN model, specifically including: initializing a flaw generator G, a flaw repairer F, a flawless image discriminator D, and a flawed image discriminator H, randomly selecting a flawless image x and a flawed image y to generate a fake flawed image G(x), a fake flawless image F(y), and an attention mask, calculating adversarial losses LossF1 and LossB1, updating the parameters of the flawless image discriminator D and the flawed image discriminator H, reconstructing the flawless image F(G(x)) and the flawed image G(F(y)), calculating the cycle consistency loss and the attention consistency loss, and updating the parameters of the flaw generator G and the flaw repairer F until the training stop condition is reached;
[0143] A supervised model building and training module, which is used to build and train the Pix2Pix model, specifically including: initializing a multi-scale discriminator A, randomly selecting a flawless image x to generate a fake flawed image G(x), randomly selecting Gaussian noise z and generating a repaired image F(G(x),z), calculating the adversarial loss LossF3 and updating the parameters of the multi-scale discriminator A, randomly selecting Gaussian noise z and generating a repaired image F(G(x),z), calculating the consistency loss LossF4, and updating the parameters of the next flaw repairer F until the training stop condition is reached;
[0144] A flaw detection module, which is used for flaw detection, specifically including: randomly selecting a flawed image y to generate an image F(y), and pixel-by-pixel comparing the flawed image y with the repaired flawless image F(y) to obtain a binary image of the flaw position and shape.
[0145] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0146] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only one type, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0147] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.
[0148] Finally: The above description is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should all be included within the protection scope of the present invention.
Claims
1. A method for detecting small target defect images of high-end equipment products, characterized in that It includes the following steps: S1: Establish and train an AttentionGAN model, specifically including: initializing a defect generator G, a defect repairer F, a defect-free image discriminator D, and a defective image discriminator H, randomly selecting a defect-free image x and a defective image y to generate a fake defective image G(x), a fake defect-free image F(y), and an attention mask, calculating the adversarial losses LossF1 and LossB1, updating the parameters of the defect-free image discriminator D and the defective image discriminator H, reconstructing the defect-free image F(G(x)) and the defective image G(F(y)), calculating the cycle consistency loss and the attention consistency loss, and updating the parameters of the defect generator G and the defect repairer F until the training stop condition is reached; S2: Establish and train a Pix2Pix model, specifically including: initializing a multi-scale discriminator A, randomly selecting a defect-free image x to generate a fake defective image G(x), randomly selecting Gaussian noise z and generating a repaired image F(G(x), z), calculating the adversarial loss LossF3 and updating the parameters of the multi-scale discriminator A, randomly selecting Gaussian noise z and generating a repaired image F(G(x), z), calculating the consistency loss LossF4, and updating the parameters of the next defect repairer F until the training stop condition is reached; S3: Defect detection, specifically including: randomly selecting a defective image y to generate an image F(y), pixel-by-pixel comparing the defective image y with the repaired defect-free image F(y) to obtain a binary image of the defect location and shape.
2. The method for detecting small target defect images of high-end equipment products according to claim 1, characterized in that In the said S1, it includes the following steps: a. Randomly select a defect-free image x and a defective image y from the defect-free image dataset X and the defective image dataset Y respectively; b. Use the defect generator G to convert the defect-free image x into a fake defective image G(x), and pass through the attention masks and , while using the defect remover F to convert the defective image y into a fake defect-free image F(y), and dynamically adjust the attention weights and ; c. Calculate the defective true rate D(x) of the defect-free image x and the true rate D(F(y)) of the transformed fake defect-free image F(y) through the defect-free image discriminator D, and strengthen the attention to the defect-free area through the attention mechanism, calculate the adversarial loss LossF1, and optimize and update the parameters of the defect-free image discriminator D; Calculate the defective true rate H(y) of the defective image y and the true rate H(G(x)) of the transformed fake defective image G(x) through the defective image discriminator H, and enhance the discriminability of the defective area through the attention mechanism, calculate the adversarial loss LossB1, and optimize and update the parameters of the defective image discriminator H; d. When training the generator, input the flawless image x and the flawed image y; use the flaw generator G to process the flawless image x, generate the foreground and background attention masks through the parameter - shared encoder and the attention module, and generate the fake flawed image G(x) by combining the content mask. At the same time, use the flaw fixer F to repair the flawed image by the generated attention masks and focus on this flawed area to generate the flawless image F(y); Then, calculate and optimize the following losses: the loss of the defect-free image discriminator D: calculate D(x) and D(F(y)) to obtain the adversarial loss of the defect-free image; the loss of the defective image discriminator H: calculate H(y) and H(G(x)) to obtain the adversarial loss of the defective image; In addition, input the fake defect-free image F(y) into the generator G to generate a reconstructed defective image G(F(y)), and at the same time input the fake defective image G(x) into the repairer F to generate a reconstructed defect-free image F(G(x)); By jointly optimizing the adversarial loss, consistency loss, and attention consistency loss, the detail capture of key regions by the generator and restorer and the accuracy of image reconstruction are improved; e. Calculate the difference between the generated flawless image F(G(x)) and the original flawless image x to obtain the consistency loss LossF2; at the same time, calculate the difference between the generated defective image G(F(y)) and the original defective image y to obtain the consistency loss LossB2; In addition, the generated attention mask is constrained by the attention consistency loss to ensure that the attention mask can accurately focus on the key regions in the image; combining the adversarial losses LossF1, LossB1, the consistency losses LossF2, LossB2, and the attention consistency loss, the accuracy and detail restoration ability of the defect generator G and restorer F in the process of generating and restoring images are further improved through joint optimization; f. If the training stop condition is reached, enter S2; if the training stop condition is not reached, repeat steps a to e.
3. The method for detecting small target defect images of high-end equipment products according to claim 1, wherein In S2, the following steps are included: g. Randomly select a fake defective image G(x) from the defective image library generated in the unsupervised stage, and form paired data (x, G(x)) with the original flawless image x, and apply a random affine transformation to G(x); h. Randomly select a Gaussian noise z, the noise dimension matches the encoder feature layer, and the noise z is superimposed in the intermediate feature layer of the encoder, and the noise weight is dynamically adjusted through a learnable attention gating mechanism: , wherein, and are learnable parameters, is the Sigmoid function, and ⊙ represents element-wise multiplication; j. Use U-Net++ as the restorer backbone network to enhance the multi-scale feature fusion ability through dense skip connections; embed a dynamic convolution module in the decoder layer to generate an adaptive convolution kernel weight according to the noise z, and the formula is: , Among them, represents the dynamically generated convolutional kernel weight matrix, MLP is the multi-layer perceptron, and Conv is the dynamic convolution operation. represents the feature map input to the dynamic convolution layer; The repaired image F(G(x), z) should satisfy to ensure pixel-level consistency; k. Multi-scale Discriminator A and Adversarial Loss Optimization: Discriminator A uses a multi-scale PatchGAN to discriminate local texture consistency at three scales of 32×32, 64×64, and 128×128 respectively; the input is an image pair concatenated by channels, and the matching probabilities at each scale are output ; l. Improved loss function: , In the formula, E represents the expected value of the joint distribution of the input variables x and z, and LPIPS represents the perceptual loss: it extracts multi-level features through a pre-trained VGG16 network and calculates the feature space distance, and the formula is: , Among them, represents the repaired image, represents the original flawless image, represents the feature map of the l-th layer of VGG16; hyperparameter: γ = 0.1, which is used to balance pixel-level and semantic-level consistency; m. If the training stop condition is reached, enter step S3; if the training stop condition is not reached, repeat steps g to l.
4. The method for detecting small target defect images of high-end equipment products according to claim 1, characterized in that In S3, the following steps are included: n. Given a random defective image y, use the defect restorer F to convert it into a fake flawless image F(y); o. Compare the defective image y and the restored flawless image F(y) pixel by pixel to generate a binary image, marking the position and shape of the defects.
5. The method for detecting small target defect images of high-end equipment products according to claim 1, wherein, The calculation formulas of the adversarial losses LossF1 and LossB1 are as follows: , , Among them, Pdata represents the distribution of the dataset, and the optimization objectives of F and D are respectively .
6. The method for detecting small target defect images of high-end equipment products according to claim 5, wherein The optimization objective is: , , Among them, is a hyperparameter for balancing the adversarial losses LossF1, LossB1 and the consistency losses LossF2, LossB2.
7. The small target defect image detection method for high-end equipment products according to claim 1, wherein The calculation formulas of the consistency losses LossF2 and LossB2 are as follows: , 。 8. The method for detecting small target defect images of high-end equipment products according to claim 1, characterized in that, The calculation formula of the adversarial loss LossF3 is as follows: , Among them, λ1 = 1.0, λ2 = 0.5, λ3 = 0.25 are scale weight coefficients, which are used to balance the discriminative contributions of different resolutions.
9. The small target defect image detection method for high-end equipment products according to claim 1, characterized in that In S3, a binary image marking the position and shape of the defects is obtained by comparing y and F(y), where if the absolute value of the pixel difference between y and F(y) is greater than the threshold, it is marked as a pixel in the defect area.
10. A small target defect image detection system for high-end equipment products, which is used to implement the small target defect image detection method for high-end equipment products according to any one of claims 1 to 9, characterized in that, Include: An unsupervised model building and training module, which is used to build and train the AttentionGAN model, specifically including: initializing the defect generator G, the defect repairer F, the defect-free image discriminator D, and the defective image discriminator H, randomly selecting a defect-free image x and a defective image y to generate a fake defective image G(x), a fake defect-free image F(y), and an attention mask, calculating the adversarial losses LossF1 and LossB1, updating the parameters of the defect-free image discriminator D and the defective image discriminator H, reconstructing the defect-free image F(G(x)) and the defective image G(F(y)), calculating the cycle consistency loss and the attention consistency loss, and updating the parameters of the defect generator G and the defect repairer F until the training stop condition is reached; A supervised model building and training module, which is used to build and train the Pix2Pix model, specifically including: initializing the multi-scale discriminator A, randomly selecting a defect-free image x to generate a fake defective image G(x), randomly selecting Gaussian noise z and generating a repaired image F(G(x),z), calculating the adversarial loss LossF3 and updating the parameters of the multi-scale discriminator A, randomly selecting Gaussian noise z and generating a repaired image F(G(x),z), calculating the consistency loss LossF4, and updating the parameters of the next defect repairer F until the training stop condition is reached; A defect detection module, which is used for defect detection, specifically including: randomly selecting a defective image y to generate an image F(y), and pixel-by-pixel comparing the defective image y with the repaired defect-free image F(y) to obtain a binary image of the defect position and shape.
Citation Information
Patent Citations
Defect detection method based on semi-supervised learning
CN113808035A
Defect detection method and device, electronic equipment and computer readable storage medium
CN114596242A
Fuzzy invoice image restoration method and device based on generative adversarial network
CN119515738A
Robust placenta implantation detection method based on semi-supervised multi-scale generative adversarial network
CN119785066A
Methods, apparatuses, systems and computer-readable mediums for correcting echo planar imaging artifacts
US20250005715A1
Cited By
Inclined basal plane laser cladding cross section size prediction method and system
CN121010637A
A method and system for predicting the cross-sectional size of laser cladding on an inclined base surface
CN121010637B