Industrial defect image generation method and device, electronic equipment and storage medium

By preprocessing and edge detection of industrial images with a multi-loss function optimization and micro-tuning, the method addresses the low-quality image generation issue in diffusion models, achieving high-quality and diverse defect images for improved industrial vision detection.

CN120318097APending Publication Date: 2025-07-15BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510378011.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

When generating industrial defect images, the existing diffusion models have rough image details and difficult training, making it difficult to meet the data needs of deep learning in the field of industrial vision detection.

Method used

By optimizing the VAE module of the diffusion model multiple loss function and combining preset fine-tuning algorithms, high-quality and diverse industrial defect images are generated.

Benefits of technology

The reconstruction accuracy and detection accuracy of industrial images have been improved, and the data needs of deep learning in the field of industrial vision detection have been met. The quality and diversity of generated images have been significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318097A_ABST
    Figure CN120318097A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial defect image generation method and device, electronic equipment and a storage medium. The generation method comprises the following steps: acquiring an industrial image, processing the industrial image, and generating a preprocessed image and an edge image; training a VAE module of the diffusion model through the preprocessed image, the edge image and preset loss functions, wherein the preset loss functions comprise a reconstruction loss function, a perception loss function and an edge reconstruction loss function; performing fine tuning training on the diffusion model through a preset fine tuning algorithm to obtain an image generation model; and generating an industrial defect image through the image generation model. According to the method, the VAE module in the diffusion model is jointly optimized through multiple loss functions, so that the coding capability of the VAE module on industrial image details is improved, the image generation model obtained by performing fine tuning training on the diffusion model is ensured, and high-quality and diversified industrial defect images can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of industrial vision inspection, and particularly to a method, device, electronic device and storage medium for generating industrial defect images. Background Art

[0002] In the field of industrial vision inspection, the generality and performance of deep learning are significantly higher than those of traditional digital image generation methods. Therefore, using deep learning to solve the problems of detection, classification, and segmentation in industrial scenarios has become a research hotspot in the field of industrial vision inspection. Deep learning is based on a large amount of data training. Although a large amount of data can be obtained on industrial production lines, most of the data is repetitive, and defective data rarely appears. Therefore, the data demand of deep learning and the difficulty of obtaining defective data on production lines have become a pair of contradictions.

[0003] In related technologies, defective data can be generated through diffusion models such as StableDiffusion to meet the needs of deep learning. Since the number of parameters of the diffusion model is large, it is very difficult to train from scratch. Therefore, the parameters of the diffusion model will be fine-tuned for specific scenarios to reduce the training difficulty of the diffusion model. However, the details of the images generated by the fine-tuned diffusion model are rough (artifacts, deformations). Therefore, how to improve the image quality generated by the diffusion model has become an urgent problem to be solved at present. Summary of the Invention

[0004] In view of the above problems, the present application provides a method, device, electronic device and storage medium for generating industrial defect images, which can effectively improve the problem of poor image quality generated by the diffusion model.

[0005] The method for generating industrial defect images of the present application includes:

[0006] Obtain an industrial image;

[0007] Process the industrial image to generate a preprocessed image and an edge image;

[0008] Train the VAE module of the diffusion model through the preprocessed image, the edge image and a preset loss function, where the preset loss function includes a reconstruction loss function, a perceptual loss function and an edge reconstruction loss function;

[0009] Fine-tune and train the diffusion model through a preset fine-tuning algorithm to obtain an image generation model; and

[0010] Generate the industrial defect image through the image generation model.

[0011] In some embodiments, training the VAE module of the diffusion model through the preprocessed image, the edge image and a preset loss function includes:

[0012] Input the preprocessed image into the VAE module of the diffusion model for encoding and decoding to generate an output image;

[0013] According to the output image, the preprocessed image, and the edge image, correct the parameters of the VAE module through a preset loss function to generate a trained VAE module.

[0014] In some embodiments, the generation method further includes:

[0015] Obtain at least two types of defect template images;

[0016] Respectively perform feature extraction on each type of defect template image and the industrial defect image through a feature extraction model to generate defect template features and features to be recognized. The feature extraction model is trained by a convolutional neural network;

[0017] Calculate the similarity between the features to be recognized and each type of template feature;

[0018] Determine the type and quality of the industrial defect image according to the similarity.

[0019] In some embodiments, determining the type and quality of the industrial defect image according to the similarity includes:

[0020] Use the category corresponding to the maximum similarity value as the category of the industrial defect image;

[0021] Determine the quality of the industrial defect image according to the maximum similarity value.

[0022] In some embodiments, the training method of the feature extraction model includes one of CosFaceLoss, arcfaceloss, or a masked autoencoder.

[0023] In some embodiments, processing the industrial image to generate a preprocessed image and an edge image includes:

[0024] Preprocess the industrial image to generate the preprocessed image;

[0025] Perform binarization processing on the preprocessed image to generate the edge image.

[0026] In some embodiments, the calculation expression of the preset loss function includes:

[0027] Loss 总 = L rec + L lpip + L rec-edge

[0028] Among them, Loss 总 is a preset loss function, and l rec is a reconstruction loss function, and l lpip is a perceptual loss function, and L rec-edge is an edge reconstruction loss function.

[0029] The generating device according to the embodiment of the present application includes:

[0030] An acquisition module, configured to acquire industrial images;

[0031] A processing module, configured to process the industrial image to generate a preprocessed image and an edge image;

[0032] A first training module, configured to train the VAE module of the diffusion model through the preprocessed image, the edge image, and a preset loss function, where the preset loss function includes a reconstruction loss function, a perceptual loss function, and an edge reconstruction loss function;

[0033] A second training module, configured to perform fine-tuning training on the diffusion model through a preset fine-tuning algorithm to obtain an image generation model; and

[0034] A generation module, configured to generate the industrial defect image through the image generation model.

[0035] The present application also provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and when the computer program is executed by the processor, the method for generating an industrial defect image as described above is implemented.

[0036] The present application also provides a non-volatile computer-readable storage medium including a computer program, and when the computer program is executed by a processor, the processor is caused to execute the method for generating an industrial defect image as described above.

[0037] In the method for generating an industrial defect image, the generating device, the electronic device, and the computer-readable storage medium according to the embodiment of the present application, by optimizing the loss function of the VAE module in the diffusion model, and adopting a multi-loss function to jointly optimize the VAE module, the encoding and decoding capabilities of the optimized VAE module are enhanced, and the reconstruction effect is good, which can improve the reconstruction accuracy of industrial images. Furthermore, by combining a preset fine-tuning algorithm to perform fine-tuning training on the diffusion model, an image generation model is obtained. In this way, the image generation model can generate high-quality and diverse industrial defect images, thereby meeting the data requirements of deep learning in the field of industrial vision detection and improving the detection accuracy in the field of industrial vision detection. Description of the Drawings

[0038] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0039] Figure 1 It is a schematic flowchart of a method for generating industrial defect images according to some embodiments of the present application.

[0040] Figure 2 It is a schematic diagram of modules of a device for generating industrial defect images according to some embodiments of the present application.

[0041] Figure 3 It is a simplified schematic diagram of a VAE module according to some embodiments of the present application.

[0042] Figures 4-6 It is a schematic flowchart of a method for generating industrial defect images according to some embodiments of the present application.

[0043] Figure 7 It is another schematic diagram of modules of a device for generating industrial defect images according to some embodiments of the present application.

[0044] Figure 8 It is a schematic flowchart of a method for generating industrial defect images according to some embodiments of the present application. Detailed Embodiments

[0045] The following describes in detail the embodiments of the present application. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.

[0046] Please refer to Figure 1 , an embodiment of the present application provides a method for generating industrial defect images. The method for generating industrial defect images includes the steps of:

[0047] 01. Obtain industrial images;

[0048] 02. Process the industrial images to generate a preprocessed image and an edge image;

[0049] 03. Train the VAE module of the diffusion model with the preprocessed image, the edge image, and a preset loss function. The preset loss function includes a reconstruction loss function, a perceptual loss function, and an edge reconstruction loss function;

[0050] 04. Fine-tune and train the diffusion model with a preset fine-tuning algorithm to obtain an image generation model; and

[0051] 05. Generate industrial defect images with the image generation model.

[0052] Please further refer to Figure 2, an embodiment of the present application provides a generating device 100 for industrial defect images. The generating device 100 includes an acquisition module 110, a processing module 120, a first training module 130, a second training module 140, and a generating module 150.

[0053] Step 01 can be implemented by the acquisition module 110, step 02 can be implemented by the processing module 120, step 03 can be implemented by the first training module 130, step 04 can be implemented by the second training module 140, and step 05 can be implemented by the generating module 150.

[0054] Or rather, the acquisition module 110 can be used to acquire industrial images; the processing module 120 can be used for 02, process the industrial images to generate a preprocessed image and an edge image; the first training module 130 can be used to train the VAE module of the diffusion model through the preprocessed image, the edge image, and a preset loss function, and the preset loss function includes a reconstruction loss function, a perceptual loss function, and an edge reconstruction loss function. The second training module 140 can be used to fine-tune and train the diffusion model through a preset fine-tuning algorithm to obtain an image generation model, and the generating module 150 can be used to generate industrial defect images through the image generation model.

[0055] The present application also provides an electronic device, including a processor and a memory. A computer program is stored in the memory. When the computer program is executed by the processor, the above-mentioned method for generating industrial defect images is implemented. That is, the processor can be used to acquire industrial images; process the industrial images to generate a preprocessed image and an edge image; train the VAE module of the diffusion model through the preprocessed image, the edge image, and a preset loss function, and the preset loss function includes a reconstruction loss function, a perceptual loss function, and an edge reconstruction loss function; fine-tune and train the diffusion model through a preset fine-tuning algorithm to obtain an image generation model; and generate industrial defect images through the image generation model.

[0056] In the methods for generating industrial defect images, the generating device 100, and the electronic device of these embodiments, by optimizing the loss function of the VAE module in the diffusion model, that is, using multiple loss functions (reconstruction loss function, perceptual loss function, and edge reconstruction loss function) to jointly optimize the VAE module, the encoding and decoding capabilities of the optimized VAE module are enhanced, and the reconstruction effect is good, which can improve the reconstruction accuracy of industrial images. Furthermore, by combining a preset fine-tuning algorithm to fine-tune and train the diffusion model, an image generation model is obtained. In this way, the image generation model can generate high-quality and diverse industrial defect images, thereby meeting the data needs in the field of industrial vision detection and improving the detection accuracy in the field of industrial vision detection.

[0057] In some examples, the generating device 100 may be part of an electronic device. Or rather, the electronic device includes the generating device 100. As hardware, the generating device 100 can be independent or added as an additional peripheral component to the electronic device. The generating device 100 can also be integrated into the electronic device. For example, the generating device 100 can be integrated onto a processor. As software, the code segment corresponding to the generating device 100 can be stored in a memory and implemented by a processor to perform the foregoing functions. Or rather, the generating device 100 includes a computer program, or rather, the foregoing computer program includes the generating device 100.

[0058] The electronic device can be a mobile phone, a tablet, a computer (personal computer, tablet computer), etc. This embodiment can be described by taking the electronic device as a computer as an example. That is to say, the method for generating industrial defect images and the generating device 100 are applied to but not limited to a computer. The generating device can be pre-installed hardware or software on the computer and can execute the generating method when started and running on the computer. For example, the generating device 100 can be a low-level software code segment of the computer or a part of the operating system.

[0059] It should be noted that industrial images refer to images involved and used in related fields such as industrial production, manufacturing, and detection. For example, in the display field, industrial images may include but are not limited to raw material detection images (such as glass substrate images, liquid crystal material images), process images during manufacturing (such as lithography process images, etching process images, coating process images), product detection images (such as appearance detection images, display performance detection images, electrical performance detection images), etc. It can be understood that in order to improve the training effect, clear and high-quality industrial images in the industry can be collected in step 01.

[0060] In step 02, the preprocessed image is a training image obtained by preprocessing the industrial image, and the edge image is an image generated by performing an edge extraction operation on the industrial image. Both the preprocessed image and the edge image are used to implement the training of the VAE module of the diffusion model. That is, the preprocessed image and the edge image serve as the training data for the VAE module in the diffusion model.

[0061] The diffusion model is a type of generative model for generating data, and its purpose is to learn to generate pictures from pure noise. In this embodiment, the diffusion model can be a StableDiffusion model (such as StableDiffusion1.5, StableDiffusionXL).

[0062] Those skilled in the art can understand that the StableDiffusion model is a free and open-source AI image generator launched by StabilityAI, which can generate high-resolution and realistic images according to the input text description. The StableDiffusion model may include several parts such as a TextEncoder module, a U-Net module, and a Variational Autoencoder (VAE module).

[0063] Among them, the TextEncoder in the StableDiffusion model is responsible for converting the input text prompt into a high-dimensional vector representation (text embedding) and providing conditional information for the U-Net; the U-Net serves as the core noise prediction network in StableDiffusion, receiving the text embedding output by the text encoder and the latent representation with noise, and gradually predicting and removing the noise through multiple inferences to optimize the latent representation; the VAE module is used to compress the image from the pixel space to the latent space through an encoder, enabling the diffusion process to occur in the latent space, and can also reconstruct the image from the latent space to the pixel space through a decoder, and uses its own structural characteristics to denoise and optimize the details of the image to improve the running efficiency of the StableDiffusion model; the simplification process of the VAE module is as Figure 3 shown. The encoder can encode an image from the RGB space, i.e., the pixel space, to the latent space representation through the encoder, and the decoder reconstructs the latent space representation to the image RGB.

[0064] Due to the large number of parameters in the StableDiffusion model (about 1B parameters for StableDiffusion 1.5, and nearly 3B parameters for StableDiffusion XL), it is extremely difficult to train from scratch. Therefore, it is necessary to fine-tune for specific scenarios and specific data to reduce the training difficulty of the StableDiffusion model. However, the images generated by the StableDiffusion model after fine-tuning have poor details, such as artifacts and deformations in the densely arranged wires, that is, there are problems with inaccurate details.

[0065] In response to this, in step 03, this embodiment improves the training of the VAE module and jointly optimizes the VAE module through multiple loss functions (including a reconstruction loss function, a perceptual loss function, and an edge reconstruction loss function). Among them, the edge reconstruction loss function is designed to highlight details. In this way, the VAE module can adapt to the encoding and decoding of industrial scenarios and handle details better.

[0066] Specifically, the preprocessed image is input into the VAE module to generate an output image. Then, based on the output image, the preprocessed image, the edge image, as well as the reconstruction loss function, the perceptual loss function, and the edge reconstruction loss function, the reconstruction loss value, the perceptual loss value, and the edge reconstruction loss value are calculated. Subsequently, the parameters of the VAE module are optimized according to the reconstruction loss value, the perceptual loss value, and the edge reconstruction loss value. Finally, the output image output by the VAE module is obtained with the minimum reconstruction loss value, perceptual loss value, and edge reconstruction loss value, that is, the training of the VAE module is completed, and the trained VAE module is generated. In some examples, by testing the VAE module trained with a preset loss function, it is found that the reconstruction effect of the output image generated by the VAE module is better, and the PSNR can be increased by 1.2 db.

[0067] Further, the preset loss function can be the reconstruction loss function L rec 、perceptual loss function L lpip and edge reconstruction loss function L rec-edge sum, and the calculation expression of the preset loss function can include:

[0068] Loss 总 =L rec +L lpip +l rec-edge

[0069] where Loss 总 is the preset loss function, l rec is the reconstruction loss function, l lpip perceptual loss function, l rec-edge is the edge reconstruction loss function.

[0070] The calculation expression of the reconstruction loss function is:

[0071]

[0072] where l rec is the reconstruction loss function, x is the original image, is the reconstructed output image, H represents the height of the image, W represents the width of the image, and C represents the number of channels of the image.

[0073] The calculation expression of the perceptual loss function is:

[0074]

[0075] where LPIPS(x,x0) is the perceptual loss function, x is the original image, x0 is the reconstructed output image, l is the l-th layer of the pre-trained network. H l , W l are the height and width of the feature map of the l-th layer, is the activation value at the position (h, w) of the original image in the feature map of the l-th layer, is the activation value at the position (h, w) of the output image in the feature map of the l-th layer.

[0076] The calculation expression of the edge reconstruction loss function is:

[0077]

[0078] where Loss rec-edge is the edge reconstruction loss function, x is the original image, is the reconstructed output image, and x edge is the image edge mask extracted by edge detection.

[0079] In step 04, the preset fine-tuning algorithm may include but is not limited to Low-Rank Adaptation (LoRA) or DreamBooth. That is, the diffusion model can be fine-tuned through LoRA or DreamBooth to obtain an image generation model.

[0080] It should be noted that LoRA is a class of techniques aimed at reducing the complexity of large models by approximating their high-dimensional structures with low-dimensional structures. In the diffusion model, LoRA adjusts the attention layer and Linear layer in the Transformer. A pre-trained weight matrix is W0 ∈ R p×q , and its weight update is constrained to be:

[0081] W0 + ΔW = W0 + BA

[0082] where B ∈ R p×r , A ∈ R r×q , and the rank r << min(p, q). For the input X, its output h can be expressed as:

[0083] h = (W0 + ΔW)X = W0X + BAX

[0084] At initialization, A is initialized with a random Gaussian, and B is set to 0 because at the start of training, ΔW = BA = 0.

[0085] DreamBooth directly fine-tunes all the parameters of the diffusion model, aiming to embed new knowledge related to its own definition while retaining the original knowledge in the diffusion model.

[0086] In this way, by fine-tuning the diffusion model through the preset fine-tuning algorithm, the diffusion model can quickly adapt to new training requirements, while greatly reducing the number of parameters to be trained, improving the training efficiency. In addition, the obtained image generation model can generate high-quality and diverse industrial defect images.

[0087] Please refer to Figure 4 , in some embodiments, step 02 includes:

[0088] 021, preprocess the industrial image to generate a preprocessed image;

[0089] 022, perform binarization on the preprocessed image to generate an edge image.

[0090] Please further combine with Figure 2 , in some embodiments, sub-steps 021 and 022 can be implemented by the processing module 120. Or rather, the processing module 120 can also be used to preprocess the industrial image to generate a preprocessed image, and perform binarization on the preprocessed image to generate an edge image.

[0091] In some embodiments, the processor can be used to preprocess the industrial image to generate a preprocessed image, and perform binarization on the preprocessed image to generate an edge image.

[0092] The preprocessing can include but is not limited to duplicate removal processing, outlier processing, image denoising, grayscale processing, image enhancement processing, geometric correction processing, etc. It can be understood that the preprocessed image obtained by preprocessing the industrial image can improve the quality of the preprocessed image, enhance the training effect of the VAE module, and improve the training efficiency.

[0093] Binarization processing is to convert a grayscale image into an image containing only two pixel values, usually black and white. The grayscale value of each pixel point is set to 0 (black) or 255 (white), thus forming an obvious contrast effect. This processing method can greatly reduce the data volume, making the subsequent edge image analysis and processing more efficient. It can be understood that due to the poor adaptability of the general pre-trained model of the VAE module to the industrial scenario and the insufficient edge reconstruction ability. Therefore, when training the VAE module by generating the preprocessed image and the edge image, the training effect of the VAE module can be improved, enabling the VAE module to adapt to the encoding and decoding of the industrial scenario and handle details better.

[0094] Please refer to Figure 5 , in some embodiments, step 03 includes:

[0095] 031, input the preprocessed image into the VAE module of the diffusion model for encoding and decoding to generate an output image;

[0096] 032, correct the parameters of the VAE module according to the output image, the preprocessed image, and the edge image through a preset loss function to generate a trained VAE module.

[0097] Please further refer to Figure 2, in some embodiments, sub-steps 031 and 032 can be implemented by the first training module 130. Or rather, the first training module 130 can also be used to encode and decode the preprocessed image into the VAE module of the diffusion model to generate an output image; the parameters of the VAE module are corrected according to the output image, the preprocessed image, and the edge image through a preset loss function to generate a trained VAE module.

[0098] In some embodiments, the processor can be used to encode and decode the preprocessed image into the VAE module of the diffusion model to generate an output image; the parameters of the VAE module are corrected according to the output image, the preprocessed image, and the edge image through a preset loss function to generate a trained VAE module.

[0099] In sub-step 032, the output image, the preprocessed image, and the edge image can be respectively input into the reconstruction loss function L rec , the perceptual loss function L lpip and the edge reconstruction loss function L rec-edge to obtain the reconstruction loss value, the perceptual loss value, and the edge reconstruction loss value, and then the parameters of the VAE module are corrected according to the reconstruction loss value, the perceptual loss value, and the edge reconstruction loss value to obtain a trained VAE module.

[0100] In this way, the VAE module can be adapted to the encoding and decoding of industrial scenarios, and the details can be better processed, improving the quality of the reconstructed image output.

[0101] In the related art, the quality and defect diversity of industrial defect images are judged by the human eye. In the case of whole-image generation, sometimes defects may not appear. Judging by the human eye is time-consuming, and there will be more types and diversities of future defects.

[0102] Therefore, in view of this, please refer to Figure 6 , in some embodiments, the method for generating industrial defect images further includes:

[0103] 06. Obtain at least two types of defect template images;

[0104] 07. Respectively extract features from each type of defect template image and industrial defect image through a feature extraction model to generate defect template features and features to be recognized. The feature extraction model is trained by a convolutional neural network;

[0105] 08. Calculate the similarity between the features to be recognized and each template feature;

[0106] 09. Determine the type and quality of the industrial defect image according to the similarity.

[0107] Please further refer to Figure 7, in some embodiments, the generating device 100 further includes a feature extraction module 160, a calculation module 170, and a determination module 180. Among them, step 05 can be implemented by the acquisition module 110, step 06 can be implemented by the feature extraction module 160, step 07 can be implemented by the calculation module 170, and step 08 can be implemented by the determination module 180. Or rather, the acquisition module 110 can also be used to acquire at least two types of defect template images; the feature extraction module 160 can be used to respectively perform feature extraction on each type of defect template image and industrial defect image through a feature extraction model to generate defect template features and features to be recognized, and the feature extraction model is trained by a convolutional neural network; the calculation module 170 can be used to calculate the similarity between the features to be recognized and each type of template feature; the determination module 180 can be used to determine the type and quality of the industrial defect image according to the similarity.

[0108] In some embodiments, the processor can be used to acquire at least two types of defect template images; respectively perform feature extraction on each type of defect template image and industrial defect image through a feature extraction model to generate defect template features and features to be recognized, and the feature extraction model is trained by a convolutional neural network; calculate the similarity between the features to be recognized and each type of template feature; determine the type and quality of the industrial defect image according to the similarity.

[0109] It should be noted that the defect template image is used to evaluate the generation effect of the industrial defect image generated by the image generation model. In step 06, the acquired industrial images with defects can be used, and then the industrial images can be classified according to the defects to obtain multiple types of defect template images. The specific type of the defect template image is not limited. For example, it can be divided into 2 types, 3 types, 5 types, 8 types, 10 types, 15 types, 20 types, 30 types, or 50 types, etc.

[0110] In step 07, the feature extraction model can be a deep learning model trained by a convolutional neural network. Among them, the backbone network of the convolutional neural network can adopt a residual network (Residual Network, Resnet) or a densely connected convolutional network (Densely Connected Convolutional Networks, Densenet), etc. It can be understood that Resnet introduces residual connections in the neural network, allowing the network to learn the difference between the input and the output, that is, the residual, so that the deep network can more effectively propagate the gradient during training, solve the problems of gradient disappearance and degradation, and thus can train a deeper network to improve the model performance. DenseNet is a convolutional neural network architecture that establishes dense connections between layers, enabling each layer to directly receive the feature maps of all previous layers as input, thereby realizing efficient feature reuse, effectively alleviating the problem of gradient disappearance, and improving the model performance.

[0111] The training method of the feature extraction model includes one of face recognition methods (such as CosFaceLoss, Arcfaceloss) or Masked Autoencoders (MAE). That is, the convolutional neural network can be trained using Cosface, Rrcfaceloss or MAE to obtain the feature extraction model.

[0112] For the classified defect template images, a fixed number of samples (M) can be sampled as templates, and each type of defect template image is sequentially input into the feature extraction model to obtain the defect template features of each type of defect template image, whose dimension is N (N is generally 2048, 1024, 512). In this way, the total size of the image template features of this category is M * N. For K categories, the total size of all templates is K * M * N. The defect template features can be denoted as: Features template For industrial defect images, they are input into the feature extraction model to generate the features to be recognized, features g 。

[0113] In step 08, the features to be recognized, features g and the defect template features, Features template can be subjected to normalization operations such as standardization or normalization, and then the cosine similarity between the features to be recognized, features g and each type of defect template feature, Features template is calculated. In this way, when calculating the similarity, it is ensured that the distance calculation is fair, and the stability and accuracy of the algorithm are improved. The calculation expression of the similarity can be:

[0114] similar = features g * Features template

[0115] where similar is the similarity, features g is the feature to be recognized, and Features template is the defect template feature.

[0116] After obtaining the similarities between the industrial defect images and various defect template images, it can be judged whether the generated defects belong to a certain category in the template. After the similarity judgment, if the consistency between the generated industrial defect image and the real image to be generated is poor, it means that the generation quality is not good. If the consistency is very high, it means that the generated industrial defect image meets the requirements.

[0117] Thus, in this embodiment, the feature extraction model extracts features from the defect template image and the industrial defect image generated by the image generation model, and calculates the similarity based on the extracted features, thereby determining the quality of the industrial defect image. The automatic screening of industrial defect images is realized. On the one hand, compared with the related technologies in which the quality and defect diversity of images are judged by human eyes, manual intervention can be reduced. On the other hand, it supports dynamically adding new defect categories. After the feature extraction model is trained, features can be directly extracted and used for new defects, with good applicability and improved model iteration efficiency.

[0118] Please refer to Figure 8 , in some embodiments, step 09 includes:

[0119] 091, using the category corresponding to the maximum similarity as the category of the industrial defect image;

[0120] 092, determining the quality of the industrial defect image according to the maximum similarity.

[0121] Please further refer to Figure 2 , in some embodiments, sub-steps 091 and 092 can be implemented by the determination module 180, or rather, the determination module 180 can also be used to use the category corresponding to the maximum similarity as the category of the industrial defect image and determine the quality of the industrial defect image according to the maximum similarity.

[0122] In some embodiments, the processor can be used to use the category corresponding to the maximum similarity as the category of the industrial defect image and determine the quality of the industrial defect image according to the maximum similarity.

[0123] It should be noted that in sub-step 091, the maximum similarity needs to exceed a certain threshold to determine whether the generated industrial defect image belongs to a certain category in the template. For example, only when the maximum similarity exceeds 0.5, it is determined that the industrial defect image belongs to one of the template images.

[0124] For example, in some examples, the similarity between industrial defect image 1 and the defect template image is 0.82, with a relatively high similarity, while the similarity between industrial defect image 2 and the defect template image is 0.48, with a relatively low similarity. Then it is considered that industrial defect image 1 and the defect template image belong to the same category, while industrial defect image 2 and the defect template image do not belong to the same category. In addition, industrial defect image 1 and industrial defect image 2 can be replaced with different generated industrial defect images.

[0125] In sub-step 092, after determining the maximum similarity, if the maximum similarity is very high, it indicates that the consistency of the industrial defect image is poor and the quality of the generated industrial defect image is not good. If the consistency is very high, it indicates that the quality of the generated industrial defect image meets the requirements.

[0126] The present application provides a non-volatile computer-readable storage medium including a computer program. When the computer program is executed by a processor, the processor is caused to execute the above-mentioned method for generating industrial defect images.

[0127] In the computer-readable storage medium according to the embodiments of the present application, by optimizing the loss function of the VAE module in the diffusion model, that is, using multiple loss functions (reconstruction loss function, perceptual loss function, and edge reconstruction loss function) to jointly optimize the VAE module, the encoding and decoding capabilities of the optimized VAE module are enhanced, and the reconstruction effect is good, which can improve the reconstruction accuracy of industrial images. Furthermore, by combining a preset fine-tuning algorithm to fine-tune and train the diffusion model, an image generation model can be obtained. In this way, the image generation model can generate high-quality and diverse industrial defect images, thereby meeting the data requirements for deep learning in the field of industrial vision detection and improving the detection accuracy in the field of industrial vision detection.

[0128] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid-state disk (SSD)), etc.

[0129] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0130] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.

[0131] In addition, the functional modules in each embodiment of this application can be integrated into a combined module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0132] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A method for generating industrial defect images, characterized in that, The generation method includes: Obtain an industrial image; Process the industrial image to generate a preprocessed image and an edge image; Train the VAE module of the diffusion model through the preprocessed image, the edge image, and a preset loss function, where the preset loss function includes a reconstruction loss function, a perceptual loss function, and an edge reconstruction loss function; Fine-tune and train the diffusion model through a preset fine-tuning algorithm to obtain an image generation model; and Generate the industrial defect image through the image generation model.

2. The generation method according to claim 1, wherein Training the VAE module of the diffusion model through the preprocessed image, the edge image, and a preset loss function includes: Input the preprocessed image into the VAE module of the diffusion model for encoding and decoding to generate an output image; Correct the parameters of the VAE module according to the output image, the preprocessed image, and the edge image through the preset loss function to generate a trained VAE module.

3. The generation method according to claim 1, wherein The generation method further includes: Obtain at least two types of defect template images; Extract features from each type of defect template image and the industrial defect image respectively through a feature extraction model to generate defect template features and features to be recognized, where the feature extraction model is trained by a convolutional neural network; Calculate the similarity between the features to be recognized and each type of template feature; Determine the type and quality of the industrial defect image according to the similarity.

4. The generation method according to claim 3, wherein Determining the type and quality of the industrial defect image according to the similarity includes: Taking the category corresponding to the maximum similarity value as the category of the industrial defect image; Determine the quality of the industrial defect image according to the maximum similarity value.

5. The generation method according to claim 3, characterized in that The training method of the feature extraction model includes one of CosFaceLoss, arcfaceloss, or a masked autoencoder.

6. The generating method according to claim 3, wherein Processing the industrial image to generate a preprocessed image and an edge image includes: Preprocess the industrial image to generate the preprocessed image; Perform binarization processing on the preprocessed image to generate the edge image.

7. The generation method according to claim 1, wherein The calculation expression of the preset loss function includes: Loss 总 = L rec + L lpip + L rec-edge Among them, Loss 总 is a preset loss function, L rec is a reconstruction loss function, L lpip is a perceptual loss function, L rec-edge is an edge reconstruction loss function.

8. An apparatus for generating industrial defect images, characterized in that The generation device may include: An acquisition module for acquiring an industrial image; A processing module for processing the industrial image to generate a preprocessed image and an edge image; A first training module for training the VAE module of the diffusion model through the preprocessed image, the edge image, and a preset loss function, where the preset loss function includes a reconstruction loss function, a perceptual loss function, and an edge reconstruction loss function; A second training module for fine-tuning and training the diffusion model through a preset fine-tuning algorithm to obtain an image generation model; and A generation module for generating the industrial defect image through the image generation model.

9. An electronic device, characterized in that, It includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the industrial defect image generation method according to any one of claims 1-7.

10. A non-volatile computer-readable storage medium comprising a computer program, characterized in that, When the computer program is executed by the processor, the processor executes the industrial defect image generation method according to any one of claims 1-7.