Defect generation model, defect removal model training method and device
Patent Information
- Application Number
- CN202611007650.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明实施例提供一种缺陷生成模型、缺陷去除模型的训练方法及装置,用于解决现有的扩散生成模型泛化能力不足,训练效率低的问题
[0050]第七方面,本发明实施例提供了一种计算机程序产品,包括计算机指令,所述计算机指令被处理器执行时实现如上述第一方面所述的缺陷生成模型的训练方法的步骤,或者,所述计算机指令被处理器执行时实现如上述第二方面所述的缺陷图像生成方法的步骤,或者所述计算机指令被处理器执行时实现如上述第三方面所述的缺陷去除模型的训练方法的步骤,或者所述计算机指令被处理器执行时实现如上述第四方面所述的缺陷去除方法的步骤。
Smart Images

Figure CN122821071A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to a training method and apparatus for a defect generation model and a defect removal model. Background Technology
[0002] In today's wave of intelligent transformation in industrial manufacturing, product quality inspection, as a crucial link in ensuring production quality, is facing the dual challenges of data scarcity and model training efficiency. To improve the accuracy and generalization ability of defect detection, leveraging generative models to generate large amounts of data to supplement the dataset has become a widely adopted core technology in the industry.
[0003] Among numerous generative modeling techniques, applications based on Generative Adversarial Networks (GANs) and diffusion generative models are particularly prominent. Compared to GANs, diffusion generative models, with their unique denoising diffusion mechanism, can generate more natural and diverse image samples. This advantage makes them stand out in defective data generation tasks, making them a key focus for researchers and engineers.
[0004] Current diffusion generation models are usually obtained by collecting defect images and fine-tuning the basic diffusion model. However, the generalization ability of the trained diffusion generation model is insufficient. It can only generate defect images of a specified type of defect. When a new type of defect appears, the diffusion generation model needs to be trained again using the new type of defect. The training efficiency of the diffusion generation model is low. Summary of the Invention
[0005] This invention provides a training method and apparatus for a defect generation model and a defect removal model, which solves the problems of insufficient generalization ability and low training efficiency of existing diffusion generation models.
[0006] To solve the above-mentioned technical problems, the present invention is implemented as follows:
[0007] In a first aspect, embodiments of the present invention provide a training method for a defect generation model, comprising:
[0008] Acquire a reference defect image and a real defect image, wherein the product defects in the reference defect image and the real defect image are of the same type;
[0009] The reference defect image and the real defect image are stitched together to obtain a first stitched image;
[0010] A first mask image is generated, the size of which is the same as the size of the first stitched image. The first mask image includes a mask, the position and size of which are determined by the position and size of the product defect in the actual defect image in the first stitched image.
[0011] A first fused image is obtained based on the first mask image and the first stitched image. The first fused image includes the reference defect image and the first mask occlusion image. The first mask occlusion image is generated from the real defect image in the first stitched image, and the product defect is occluded by the mask.
[0012] The diffusion model is trained based on the first fused image, wherein the diffusion model generates a defect generation image based on the first fused image; the diffusion model is fine-tuned based on the defect generation image and the real defect image to obtain a defect generation model.
[0013] Optionally, the diffusion model generates a defect image based on the first fused image, including:
[0014] The diffusion model refers to the product defects in the reference defect image in the first fused image, predicts the product defects in the area of the first masked image that is occluded by the mask, and generates a defect generation image based on the predicted product defects and the first masked image.
[0015] Optionally, obtaining the reference defect image includes:
[0016] Select an initial reference defect image with a specified type of product defect from the dataset, and crop the area where the product defect is located in the initial reference defect image according to the location of the product defect marked in the initial reference defect image. Then merge the area where the product defect is located into the background image to generate the reference defect image.
[0017] Optionally, merging the area where the product defect is located into the background image includes at least one of the following:
[0018] Place the area where the product defect is located in the center of the background image;
[0019] If the ratio of the area where the product defect is located to the background image is lower than a preset ratio threshold, the area where the product defect is located is enlarged so that the ratio of the area where the product defect is located to the background image is higher than the preset ratio threshold, and the enlarged area where the product defect is located is merged into the background image.
[0020] Optionally, the area where the mask is located in the first mask image is black, and the other areas are white.
[0021] Optionally, obtaining the first fused image based on the first mask image and the first stitched image includes:
[0022] The first mask image and the first stitched image are merged to obtain the initial first fused image;
[0023] Random noise is added to the region where the mask is located in the initial first fused image to obtain the first fused image.
[0024] Secondly, embodiments of the present invention provide a method for generating defect images, including:
[0025] Acquire a reference defect image and a good product image, wherein the reference defect image contains a product defect of a specified type;
[0026] The reference defect image and the good product image are stitched together to obtain a second stitched image;
[0027] A second mask image is generated, the size of which is the same as the size of the second stitched image, and the second mask image includes a mask;
[0028] A second fused image is obtained based on the second mask image and the second stitched image. The second fused image includes the reference defect image and the second mask occlusion image. The second mask occlusion image is generated from the good product image in the second stitched image, and a specified position is occluded by the mask.
[0029] The second fused image is input into the defect generation model, wherein the defect generation model generates the specified type of product defect in the mask area on the good product image based on the product defect in the reference defect image in the second fused image, thereby obtaining a defect generation image;
[0030] The defect generation model is trained using the method described in the first aspect above.
[0031] Optionally, the position and size of the mask are determined randomly or specified by the user.
[0032] Thirdly, embodiments of the present invention provide a training method for a defect removal model, comprising:
[0033] Obtain a reference good product image and a real defect image, wherein the real defect image contains a product defect and the product in the real defect image is the same as the product in the reference good product image;
[0034] The reference good product image and the actual defect image are stitched together to obtain a third stitched image;
[0035] A third mask image is obtained, the size of which is the same as that of the third stitched image. The third mask image includes a mask, and the area where the mask is located corresponds to the location of the real defect image in the third stitched image.
[0036] Based on the third mask image and random noise, a fourth mask image is obtained, wherein the region where the mask is located in the fourth mask image is the random noise;
[0037] A third fused image is obtained based on the fourth mask image and the third stitched image. The third fused image includes the reference good product image and the half-occluded image. The half-occluded image is formed by fusing the real defect image and the random noise according to a preset ratio.
[0038] The diffusion model is trained based on the third fused image, wherein the diffusion model predicts the product structure of the partially occluded image based on the reference good product image in the third fused image, and generates a defect removal image; the diffusion model is fine-tuned based on the defect removal image and the reference good product image to obtain the defect removal model.
[0039] Optionally, in the third mask image, the area where the mask is located is black, and the other areas are white.
[0040] Fourthly, embodiments of the present invention provide a defect removal method, comprising:
[0041] Obtain a reference good product image and a defective product image, wherein the product in the defective product image is the same as the product in the reference good product image;
[0042] The reference good product image and the defective image are stitched together to obtain a fourth stitched image;
[0043] A fifth mask image is obtained, the size of which is the same as that of the fourth stitched image. The fifth mask image includes a mask, and the area where the mask is located corresponds to the location of the defect image in the fourth stitched image.
[0044] Based on the fifth mask image and random noise, a sixth mask image is obtained, wherein the region where the mask is located in the sixth mask image is the random noise;
[0045] A fourth fused image is obtained based on the sixth mask image and the fourth stitched image. The fourth fused image includes the reference good product image and the half-occluded image. The half-occluded image is formed by fusing the defect image and the random noise according to a preset ratio.
[0046] The fourth fused image is input into the defect removal model, wherein the defect removal model predicts the defect image based on the reference good product image in the fourth fused image to generate a defect-removed image;
[0047] The defect removal model is trained using the method described in the third aspect above.
[0048] Fifthly, embodiments of the present invention provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the steps of the training method for the defect generation model as described in the first aspect above; or, when the program is executed by the processor, it implements the steps of the defect image generation method as described in the second aspect above; or, when the program is executed by the processor, it implements the steps of the training method for the defect removal model as described in the third aspect above; or, when the program is executed by the processor, it implements the steps of the defect removal method as described in the fourth aspect above.
[0049] In a sixth aspect, embodiments of the present invention provide a computer-readable storage medium storing a computer program, wherein when executed by a processor, the computer program implements the steps of the training method for the defect generation model as described in the first aspect above, or, when executed by a processor, the computer program implements the steps of the defect image generation method as described in the second aspect above, or, when executed by a processor, the computer program implements the steps of the training method for the defect removal model as described in the third aspect above, or, when executed by a processor, the computer program implements the steps of the defect removal method as described in the fourth aspect above.
[0050] In a seventh aspect, embodiments of the present invention provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the training method for the defect generation model as described in the first aspect above, or, when executed by a processor, implement the steps of the defect image generation method as described in the second aspect above, or, when executed by a processor, implement the steps of the training method for the defect removal model as described in the third aspect above, or, when executed by a processor, implement the steps of the defect removal method as described in the fourth aspect above.
[0051] In this embodiment of the invention, during model training, a reference image is stitched together along the spatial dimension of the training image and used as the input to the model to be trained. The diffusion model is then fine-tuned locally, enabling it to transfer any input defect to the mask region of a good product image based on the reference image. Alternatively, the diffusion model can remove local or global defects from a defective image based on the reference image to transform it into a good product image. Even if a new type of defect appears, as long as the reference image is available, there is no need to retrain the diffusion model, thus improving the model's generalization ability and training efficiency. Attached Figure Description
[0052] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0053] Figure 1 This is one of the flowcharts illustrating the training method of the defect generation model according to an embodiment of the present invention;
[0054] Figure 2 This is a second schematic flowchart of the training method for the defect generation model according to an embodiment of the present invention;
[0055] Figure 3 This is a schematic diagram of a method for generating a first fused image according to an embodiment of the present invention;
[0056] Figure 4 This is a schematic diagram of the structure of the training device for the defect generation model according to an embodiment of the present invention;
[0057] Figure 5 This is one of the flowcharts illustrating the defect image generation method according to an embodiment of the present invention;
[0058] Figure 6 This is a second schematic flowchart of the defect image generation method according to an embodiment of the present invention;
[0059] Figure 7 This is one of the flowcharts illustrating the training method of the defect removal model according to an embodiment of the present invention;
[0060] Figure 8 This is a second schematic flowchart of the training method for the defect removal model according to an embodiment of the present invention;
[0061] Figure 9 This is a schematic diagram of the structure of the training device for the defect removal model in an embodiment of the present invention;
[0062] Figure 10 This is a schematic flowchart of the defect removal method according to an embodiment of the present invention;
[0063] Figure 11 This is a schematic diagram of the structure of the training device for the defect generation model according to an embodiment of the present invention;
[0064] Figure 12 This is a schematic diagram of the defect image generation device according to an embodiment of the present invention;
[0065] Figure 13 This is a schematic diagram of the structure of the training device for the defect removal model according to an embodiment of the present invention;
[0066] Figure 14This is a schematic diagram of the defect removal device according to an embodiment of the present invention;
[0067] Figure 15 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0068] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0069] Please refer to Figure 1 and Figure 2 This invention provides a training method for a defect generation model, comprising:
[0070] Step S11: Obtain a reference defect image and a real defect image, wherein the product defects in the reference defect image and the real defect image are of the same type;
[0071] In this embodiment of the invention, the product may include, for example, a display screen, a semiconductor wafer, a printed circuit board (PCB), a metal casting, textiles, optical components, etc. The types of defects in the product may include, for example, cracks, scratches, holes, dirt, etc.
[0072] For example, please refer to Figure 2 The product defects in both the reference defect image P1 and the actual defect image P2 are dirt.
[0073] In this embodiment of the invention, the types of product defects in the reference defect image and the real defect image are the same, but the shape, size, etc. of the product defects may be different.
[0074] In this embodiment of the invention, different types of product defects and real defect images of different products can be collected to train the model.
[0075] Step S12: Stitch the reference defect image and the real defect image together to obtain a first stitched image;
[0076] In this embodiment of the invention, the reference defect image and the actual defect image can be stitched together horizontally. For example, the reference defect image can be on the left and the actual defect image on the right, or the reference defect image can be on the right and the actual defect image on the left. Of course, it is also possible to stitch the reference defect image and the actual defect image vertically.
[0077] For example, please refer to Figure 2The reference defect image P1 and the real defect image P2 are stitched together left and right to obtain the first stitched image P3.
[0078] Step S13: Generate a first mask image. The size of the first mask image is the same as the size of the first stitched image. The first mask image includes a mask. The position and size of the mask are determined by the position and size of the product defect in the real defect image in the first stitched image.
[0079] Please refer to Figure 2 Based on the position and size of the product defect in the real defect image in the first stitched image P3, the position and size of the mask in the first mask image are determined so that after the first mask image and the first stitched image are fused, the mask can exactly cover the product defect in the real defect image in the first stitched image.
[0080] Step S14: Obtain a first fused image based on the first mask image and the first stitched image. The first fused image includes the reference defect image and the first mask occlusion image. The first mask occlusion image is generated from the real defect image in the first stitched image, and the product defect is occluded by the mask.
[0081] In this embodiment of the invention, the areas outside the mask in the first mask-covered image are all areas without defects. Since the product defects are covered by the mask, they can be considered as good product images.
[0082] Please refer to Figure 2 The first mask image P4 and the first stitched image P3 are fused to obtain the first fused image P5. The left side of the first fused image P5 is a reference defect image, and the right side is the first mask occlusion image. The first mask occlusion image corresponds to the real defect image in the first stitched image P3, and the product defect is occluded by the mask (that is, except for the product defect being occluded by the mask, the other areas of the first mask occlusion image are the same as the real defect image).
[0083] Step S15: Train the diffusion model based on the first fused image, wherein the diffusion model generates a defect generation image based on the first fused image; fine-tune the diffusion model based on the defect generation image and the real defect image to obtain a defect generation model.
[0084] In some embodiments, optionally, the diffusion model generating a defect image based on the first fused image includes: the diffusion model referring to a product defect in a reference defect image in the first fused image, predicting product defects in areas of the first masked image that are occluded by the mask, and generating a defect image based on the predicted product defects and the first masked image.
[0085] Please refer to Figure 2 The diffusion model takes a first fused image P5 as input and outputs a stitched image P6 of a reference defect image and a defect-generated image. A loss is calculated based on the defect-generated image and the real defect image P2. The diffusion model is then fine-tuned based on this loss to obtain the defect generation model.
[0086] In this embodiment of the invention, the loss can be the mean squared error (MSE) loss.
[0087] The diffusion model in this embodiment of the invention can be of various types, such as the FLUX.1-Fill-dev diffusion model, which has context learning capability and learns and references the first fused image.
[0088] In this embodiment of the invention, by stitching a reference image (reference defect image) along the spatial dimension of the training image (real defect image) as the input to the diffusion model to be trained, the diffusion model is locally fine-tuned so that the diffusion model can transfer any input defect to the mask area of the good product image based on the reference image (reference defect image). Even if a new type of defect appears, as long as the reference image (reference defect image) is available, there is no need to retrain the diffusion model, which improves the generalization ability of the model and the efficiency of model training.
[0089] In this embodiment of the invention, optionally, obtaining the reference defect image includes:
[0090] Step S111: Select an initial reference defect image with a specified type of product defect from the dataset, and crop the area where the product defect is located in the initial reference defect image according to the location of the product defect marked in the initial reference defect image, and merge the area where the product defect is located into the background image to generate the reference defect image.
[0091] In some embodiments, the background image may optionally be a pure white background image, thereby facilitating processing.
[0092] Please refer to Figure 2 Select a defect image P0 containing a product defect of a specified type from the dataset, and crop the area containing the product defect in the defect image P2 according to the location of the product defect marked in the defect image P2, and merge the area containing the product defect into the background image to generate the reference defect image P1.
[0093] In some embodiments, optionally, incorporating the area where the product defect is located into the background image includes at least one of the following:
[0094] The area containing the product defect is placed in the center of the background image to facilitate diffusion model recognition.
[0095] If the ratio of the area containing the product defect to the background image is lower than a preset ratio threshold, the area containing the product defect is enlarged so that the ratio of the area containing the product defect to the background image is higher than the preset ratio threshold. The enlarged area containing the product defect is then merged into the background image to facilitate recognition by the diffusion model. The preset ratio threshold can be, for example, 10%.
[0096] In some embodiments, alternatively, please refer to Figure 2 In the first mask image P4, the area containing the mask is black, and the other areas are white. Of course, in other embodiments of the present invention, it is not excluded that the area containing the mask in the first mask image may be white, and the other areas black.
[0097] In some embodiments, optionally, obtaining the first fused image based on the first mask image and the first stitched image includes:
[0098] Step S141: Merge the first mask image and the first stitched image to obtain an initial first fused image;
[0099] Step S142: Add random noise to the region where the mask is located in the initial first fused image to obtain the first fused image.
[0100] Please refer to Figure 3 The first mask image P4 and the first stitched image P3 are merged to obtain an initial first fused image P5'; random noise is added to the area where the mask is located in the initial first fused image P5' to obtain the first fused image P5.
[0101] Please refer to Figure 4 , Figure 4This is a schematic diagram of the training device for the defect generation model in an embodiment of the present invention. The device includes: a diffusion model and a VAE (Variational Autoencoder) encoder, a VAE decoder, and a text encoder connected to the diffusion model. The VAE encoder fuses a first stitched image P3 with a first mask image P4, converting them into a latent vector representation through encoding. The VAE decoder converts the latent variables back into image space, thereby generating a defect generation image. The text encoder encodes prompt text, which may include, for example, prompting the diffusion model to refer to a reference defect image in the first fused image, predict product defects in areas of the first masked image that are obscured by the mask, and generate a defect generation image based on the predicted product defects and the first masked image. Figure 4 The diffusion model in the model can be a Transformer structure, which includes multiple blocks (modules), such as downsampling blocks, intermediate blocks, and upsampling blocks.
[0102] In this embodiment of the invention, the process of predicting product defects in the area of the first mask occlusion image in the first fused image that is occluded by the mask is a denoising process for the random noise in the mask area of the first mask occlusion image.
[0103] In this embodiment of the invention, when training the defect generation model, only a single defect sample is needed for a type of product defect, which reduces the cost and complexity of model training.
[0104] Please refer to Figure 5 and Figure 6 This invention also provides a method for generating defect images, comprising:
[0105] Step S21: Obtain a reference defect image and a good product image, wherein the reference defect image contains a product defect of a specified type;
[0106] In this embodiment of the invention, the product may include, for example, a display screen, a semiconductor wafer, a printed circuit board, a metal casting, textiles, optical components, etc. The types of defects in the product may include, for example, cracks, scratches, holes, dirt, etc.
[0107] For example, please refer to Figure 6 The product defect in the reference defect image P7 is dirt.
[0108] In this embodiment of the invention, the reference defect image may be generated by cropping or enlarging a product defect in a defect image and placing it on a white background image.
[0109] Step S22: Stitch the reference defect image and the good product image together to obtain a second stitched image;
[0110] In this embodiment of the invention, the reference defect image and the good product image can be stitched together horizontally. For example, the reference defect image can be on the left and the good product image on the right, or the reference defect image can be on the right and the good product image on the left. Of course, it is also possible to stitch the reference defect image and the good product image vertically. It should be noted that the image stitching method is the same as the stitching method used during model training.
[0111] For example, please refer to Figure 6 The defective image P7 and the good image P8 are stitched together left and right to obtain the second stitched image P9.
[0112] Step S23: Generate a second mask image, the size of which is the same as the size of the second stitched image, and the second mask image includes a mask;
[0113] In this embodiment of the invention, the position and size of the mask represent the position and size of the product defects subsequently generated on the good product image, and can be determined randomly or specified by the user.
[0114] Step S24: Obtain a second fused image based on the second mask image and the second stitched image. The second fused image includes the reference defect image and the second mask occlusion image. The second mask occlusion image is generated from the good product image in the second stitched image, and a specified position is occluded by the mask.
[0115] Please refer to Figure 6 The second mask image P10 and the second stitched image P9 are fused to obtain the second fused image P11. The left side of the second fused image P11 is a reference defect image, and the right side is the second mask occlusion image. The second mask occlusion image corresponds to the good product image in the second stitched image P9, and a specified position is occluded by the mask. The specified position is the position of the mask in the second mask image.
[0116] Step S25: Input the second fused image into the defect generation model, wherein the defect generation model generates the specified type of product defect based on the product defect in the reference defect image in the second fused image, in the mask area on the good product image, to obtain a defect generation image;
[0117] The defect generation model is trained using the defect generation model training method described in any of the above embodiments.
[0118] Please refer to Figure 6The input to the defect generation model is the second fused image P11, and the output is a stitched image P12 of the reference defect image and the defect generation image.
[0119] In this embodiment of the invention, it is supported to generate new product defects in a single step, even if the new product defect is not used to train the defect generation model, demonstrating extremely high generalization ability. That is, it achieves structural reconstruction of defect images even without defective samples.
[0120] Please refer to Figure 7 and Figure 8 This invention provides a training method for a defect removal model, comprising:
[0121] Step S31: Obtain a reference good product image and a real defect image, wherein the real defect image contains a product defect and the product in the real defect image is the same as the product in the reference good product image;
[0122] In some embodiments, optionally, the product in the real defect image is the same as the product in the reference good product image, and the structure and background of the product are the same.
[0123] In this embodiment of the invention, the product may include, for example, a display screen, a semiconductor wafer, a printed circuit board, a metal casting, textiles, optical components, etc. The types of defects in the product may include, for example, cracks, scratches, holes, dirt, etc.
[0124] For example, please refer to Figure 8 The product defect in the actual defect image P14 is dirt.
[0125] Step S32: The reference good product image and the real defect image are stitched together to obtain a third stitched image;
[0126] In this embodiment of the invention, the reference good product image and the actual defective image can be stitched together horizontally. For example, the reference good product image can be on the left and the actual defective image on the right, or the reference good product image can be on the right and the actual defective image on the left. Of course, it is also possible to stitch the reference good product image and the actual defective image vertically.
[0127] For example, please refer to Figure 8 The reference good product image P13 and the real defective image P14 are stitched together left and right to obtain the third stitched image P15.
[0128] Step S33: Obtain a third mask image, the size of which is the same as the size of the third stitched image. The third mask image includes a mask, and the area where the mask is located corresponds to the location of the real defect image in the third stitched image.
[0129] In this embodiment of the invention, optional details may be found, please refer to [the relevant documentation]. Figure 8 In the third mask image P16, the area containing the mask is black, and the other areas are white. Of course, in other embodiments of the present invention, it is not excluded that the area containing the mask is white and the other areas are black.
[0130] Step S34: Based on the third mask image and random noise, obtain a fourth mask image, wherein the region where the mask is located in the fourth mask image is the random noise;
[0131] Step S35: Obtain a third fused image based on the fourth mask image and the third stitched image. The third fused image includes the reference good product image and the half-occluded image. The half-occluded image is formed by fusing the real defect image and the random noise according to a preset ratio.
[0132] Optionally, the preset ratio is, for example, 6:4, where 6 and 4 are weighting ratios, such as half-occluded image = true defect image. 0.6+ random noise 0.4 indicates that the higher the weight of random noise, the more pronounced the noise. The real defect image and the random noise are fused according to a preset ratio, rather than completely obscuring the real defect image with random noise. The purpose is to allow the diffusion model to obtain the basic background structure information in the real defect image.
[0133] Please refer to Figure 8 Random noise is added to the mask region of the third mask image P16 to obtain the fourth mask image P17.
[0134] Please refer to Figure 8 The fourth mask image P7 and the third stitched image P15 are fused to obtain the third fused image P18. The left side of the third fused image P18 is a reference good product image, and the right side is a half-occluded image. The half-occluded image corresponds to the real defect image in the third stitched image P15. The half-occluded image is formed by fusing the real defect image P14 and the random noise according to a preset ratio, so that some details in the half-occluded image can be recognized by the diffusion model.
[0135] Step S36: Train the diffusion model based on the third fused image. The diffusion model predicts the product structure of the partially occluded image based on the reference good product image in the third fused image and generates a defect removal image. Fine-tune the diffusion model based on the defect removal image and the reference good product image to obtain the defect removal model.
[0136] Please refer to Figure 8The diffusion model takes a third fused image P18 as input and outputs a stitched image P19 composed of a reference good image and a defect-removed image. A loss is calculated based on the defect-removed image and the reference good image, and the diffusion model is fine-tuned based on this loss to obtain the defect-removal model.
[0137] In this embodiment of the invention, during model training, a reference image (reference good product image) is stitched along the spatial dimension from the training image (real defect image) and used as the input to the model to be trained. The diffusion model is locally fine-tuned so that the diffusion model can remove local or global defects in the defect image (real defect image) based on the reference image (reference good product image) to transform it into a good product image. Even if a new type of defect appears, as long as the reference image is available, there is no need to retrain the diffusion model, which improves the generalization ability of the model and the efficiency of model training.
[0138] Please refer to Figure 9 , Figure 9 This is a schematic diagram of the training device for the defect removal model in an embodiment of the present invention. The device includes: a diffusion model and a VAE (Variational Autoencoder) encoder, a VAE decoder, and a text encoder connected to the diffusion model. The VAE encoder is used to fuse a third stitched image with a fourth mask image, converting them into a latent vector representation through encoding. The VAE decoder is used to convert the latent variables back into image space, thereby generating a defect-removed image. The text encoder is used to encode prompt text, which may include, for example, prompting the diffusion model to predict the product structure of the partially occluded image based on the reference good product image in the third fused image, and generate the defect-removed image. Figure 9 The diffusion model in the model can be a Transformer structure, which includes multiple blocks (modules), such as downsampling blocks, intermediate blocks, and upsampling blocks.
[0139] In this embodiment of the invention, it is not necessary to mark the location of product defects in the real defect image, but the product structure can be reconstructed and the product curves can be removed to achieve a good product image.
[0140] Please refer to Figure 10 This invention also provides a defect removal method, comprising:
[0141] Step S41: Obtain a reference good product image and a defective product image, wherein the product in the defective product image is the same as the product in the reference good product image;
[0142] Step S42: The reference good product image and the defective image are stitched together to obtain a fourth stitched image;
[0143] Step S43: Obtain a fifth mask image, the size of which is the same as that of the fourth stitched image. The fifth mask image includes a mask, and the area where the mask is located corresponds to the location of the defect image in the fourth stitched image.
[0144] Step S44: Based on the fifth mask image and random noise, obtain a sixth mask image, wherein the region where the mask is located in the sixth mask image is the random noise;
[0145] Step S45: Obtain a fourth fused image based on the sixth mask image and the fourth stitched image. The fourth fused image includes the reference good product image and the half-occluded image. The half-occluded image is formed by fusing the defect image and the random noise according to a preset ratio.
[0146] Step S46: Input the fourth fused image into the defect removal model, wherein the defect removal model predicts the defect image based on the reference good product image in the fourth fused image to generate a defect-removed image;
[0147] The defect removal model is trained using the defect removal model training method described in any of the above embodiments.
[0148] In this embodiment of the invention, it is possible to generate a new good product image at once, even if the new good product image is not used to train the defect removal model, thus exhibiting extremely high generalization ability.
[0149] The effectiveness of the defect generation model in the embodiments of the present invention will be verified based on the experimental results below.
[0150] To address the unique challenges of industrial defect image generation (such as the special nature of defect features and the robustness requirements of detection models), a three-layer progressive verification framework is constructed, comprising semantic consistency verification, generation pass rate verification, and detection model adaptability verification.
[0151] Semantic consistency was verified by vectorizing both generated and real images using image vectorization feature extraction models. To make the results more convincing, two multimodal vectorization models (CLIP-Vit-Large and Gme_qwen2_VL) were selected. Twenty images were randomly selected from real and generated images with the same product defect. Feature vectors were extracted using the models, and the average vectors were calculated. The cosine similarity of the average vectors of the two images was then determined. The results are shown in Table 1, indicating that the generated images achieved good semantic consistency.
[0152] Table 1. Results of Semantic Consistency Comparison
[0153]
[0154] The pass rate is determined by generating a certain number of images of the same product defect and judging by human visual inspection whether the defects are complete and reasonable. As shown in Table 2, most product defects achieve a good pass rate, while defect three, being a global defect, is relatively more difficult to generate.
[0155] Table 2 Defect Generation Pass Rate
[0156]
[0157] To further verify the effectiveness of the generated images, the target detection model was trained using generated images and real images with different proportions and quantities of three types of defects. The target detection model used was YOLOv8, and the model results were the optimal model obtained during the iteration process. Using 1800 real images, 360 real images, and a combination of 360 real images and 1400 generated images, respectively, it was found that 80% of the generated images significantly improved the model's recall and precision. Although there was still a gap compared to the results of training with the same number of real images, the recall was improved by 15.61% and the precision by 7.3%.
[0158] Table 3 Test results of the object detection model based on generated images
[0159]
[0160] Please refer to Figure 11 This invention also provides a training device 100 for a defect generation model, comprising:
[0161] The first acquisition module 101 is used to acquire a reference defect image and a real defect image, wherein the product defects in the reference defect image and the real defect image are of the same type.
[0162] The stitching module 102 is used to stitch the reference defect image and the real defect image together to obtain a first stitched image;
[0163] The second acquisition module 103 is used to generate a first mask image, the size of which is the same as the size of the first stitched image. The first mask image includes a mask, and the position and size of the mask are determined by the position and size of the product defect in the real defect image in the first stitched image.
[0164] The fusion module 104 is used to obtain a first fused image based on the first mask image and the first stitched image. The first fused image includes the reference defect image and the first mask occlusion image. The first mask occlusion image is generated from the real defect image in the first stitched image, and the product defect is occluded by the mask.
[0165] Training module 105 is used to train a diffusion model based on the first fused image, wherein the diffusion model generates a defect generation image based on the first fused image; and the diffusion model is fine-tuned based on the defect generation image and the real defect image to obtain a defect generation model.
[0166] Optionally, the diffusion model generating a defect image based on the first fused image includes: the diffusion model refers to a product defect in a reference defect image in the first fused image, predicts the product defect in the area of the first mask occlusion image that is occluded by the mask, and generates a defect image based on the predicted product defect and the first mask occlusion image.
[0167] Optionally, obtaining the reference defect image includes:
[0168] Select an initial reference defect image with a specified type of product defect from the dataset, and crop the area where the product defect is located in the initial reference defect image according to the location of the product defect marked in the initial reference defect image. Then merge the area where the product defect is located into the background image to generate the reference defect image.
[0169] Optionally, the background image is a pure white background image.
[0170] Optionally, merging the area where the product defect is located into the background image includes at least one of the following:
[0171] Place the area where the product defect is located in the center of the background image;
[0172] If the ratio of the area where the product defect is located to the background image is lower than a preset ratio threshold, the area where the product defect is located is enlarged so that the ratio of the area where the product defect is located to the background image is higher than the preset ratio threshold, and the enlarged area where the product defect is located is merged into the background image.
[0173] Optionally, the area where the mask is located in the first mask image is black, and the other areas are white.
[0174] Optionally, obtaining the first fused image based on the first mask image and the first stitched image includes:
[0175] The first mask image and the first stitched image are merged to obtain the initial first fused image;
[0176] Random noise is added to the region where the mask is located in the initial first fused image to obtain the first fused image.
[0177] Please refer to Figure 12This invention also provides a defect image generation apparatus 200, comprising:
[0178] The acquisition module 201 is used to acquire a reference defect image and a good product image, wherein the reference defect image contains a product defect of a specified type;
[0179] The stitching module 202 is used to stitch the reference defect image and the good product image together to obtain a second stitched image;
[0180] The mask generation module 203 is used to generate a second mask image, the size of which is the same as the size of the second stitched image, and the second mask image includes a mask;
[0181] Optionally, the position and size of the mask are determined randomly or specified by the user.
[0182] The fusion module 204 is used to obtain a second fused image based on the second mask image and the second stitched image. The second fused image includes the reference defect image and the second mask occlusion image. The second mask occlusion image is generated from the good product image in the second stitched image, and a specified position is occluded by the mask.
[0183] The defect generation module 205 is used to input the second fused image into the defect generation model, wherein the defect generation model generates the specified type of product defect in the mask area on the good product image based on the product defect in the reference defect image in the second fused image, thereby obtaining a defect generation image;
[0184] The defect generation model is trained using the defect generation model training method described in any of the above embodiments.
[0185] Please refer to Figure 13 This invention also provides a training device 300 for a defect removal model, comprising:
[0186] The first acquisition module 301 is used to acquire a reference good product image and a real defect image, wherein the real defect image contains product defects and the product in the real defect image is the same as the product in the reference good product image.
[0187] Optionally, the product in the real defect image is the same as the product in the reference good product image, and the product structure and background are the same.
[0188] The stitching module 302 is used to stitch the reference good product image and the real defect image to obtain a third stitched image;
[0189] The second acquisition module 303 is used to acquire a third mask image, the size of which is the same as the size of the third stitched image. The third mask image includes a mask, and the area where the mask is located corresponds to the location of the real defect image in the third stitched image.
[0190] The third acquisition module 304 is used to obtain a fourth mask image based on the third mask image and random noise, wherein the region where the mask is located in the fourth mask image is the random noise;
[0191] The fusion module 305 is used to obtain a third fused image based on the fourth mask image and the third stitched image. The third fused image includes the reference good product image and the half-occluded image. The half-occluded image is formed by fusing the real defect image and the random noise according to a preset ratio.
[0192] Training module 306 is used to train the diffusion model based on the third fused image, wherein the diffusion model predicts the product structure of the partially occluded image based on the reference good product image in the third fused image and generates a defect removal image; the diffusion model is fine-tuned based on the defect removal image and the reference good product image to obtain the defect removal model.
[0193] Optionally, in the third mask image, the area where the mask is located is black, and the other areas are white.
[0194] Please refer to Figure 14 This invention also provides a defect removal device 400, comprising:
[0195] The first acquisition module 401 is used to acquire a reference good product image and a defective image, wherein the product in the defective image is the same as the product in the reference good product image;
[0196] Optionally, the product in the defective image is the same as the product in the reference good product image, and the product structure and background are the same.
[0197] The stitching module 402 is used to stitch the reference good product image and the defective image together to obtain a fourth stitched image;
[0198] The second acquisition module 403 is used to acquire a fifth mask image, the size of which is the same as the size of the fourth stitched image. The fifth mask image includes a mask, and the area where the mask is located corresponds to the location of the defect image in the fourth stitched image.
[0199] The third acquisition module 404 is used to obtain a sixth mask image based on the fifth mask image and random noise, wherein the region where the mask is located in the sixth mask image is the random noise;
[0200] The fusion module 405 is used to obtain a fourth fused image based on the sixth mask image and the fourth stitched image. The fourth fused image includes the reference good product image and the half-occluded image. The half-occluded image is formed by fusing the defect image and the random noise according to a preset ratio.
[0201] The defect removal module 406 is used to input the fourth fused image into the defect removal model, wherein the defect removal model predicts the defect image based on the reference good product image in the fourth fused image to generate a defect-removed image;
[0202] The defect removal model is trained using the defect removal model training method described in any of the above embodiments.
[0203] Please refer to Figure 15 The present invention also provides an electronic device 500, including a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the computer program is executed by the processor 501, it implements the various processes of the above method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0204] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described method embodiments and achieves the same technical effects. To avoid repetition, further details are omitted here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0205] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described... Figure 1 The various processes of the method embodiments shown can achieve the same technical effect, and will not be described again here to avoid repetition.
[0206] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0207] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0208] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.
Claims
1. A training method for a defect generation model, characterized in that, include: Acquire a reference defect image and a real defect image, wherein the product defects in the reference defect image and the real defect image are of the same type; The reference defect image and the real defect image are stitched together to obtain a first stitched image; A first mask image is generated, the size of which is the same as the size of the first stitched image. The first mask image includes a mask, the position and size of which are determined by the position and size of the product defect in the actual defect image in the first stitched image. A first fused image is obtained based on the first mask image and the first stitched image. The first fused image includes the reference defect image and the first mask occlusion image. The first mask occlusion image is generated from the real defect image in the first stitched image, and the product defect is occluded by the mask. The diffusion model is trained based on the first fused image, wherein the diffusion model generates a defect generation image based on the first fused image; The diffusion model is fine-tuned based on the generated defect image and the real defect image to obtain the defect generation model.
2. The method according to claim 1, characterized in that, The diffusion model generates a defect-generating image based on the first fused image, including: The diffusion model refers to the product defects in the reference defect image in the first fused image, predicts the product defects in the area of the first masked image that is occluded by the mask, and generates a defect generation image based on the predicted product defects and the first masked image.
3. The method according to claim 1, characterized in that, The acquisition of the reference defect image includes: Select an initial reference defect image with a specified type of product defect from the dataset, and crop the area where the product defect is located in the initial reference defect image according to the location of the product defect marked in the initial reference defect image. Then merge the area where the product defect is located into the background image to generate the reference defect image.
4. The method according to claim 3, characterized in that, Incorporating the area containing the product defect into the background image includes at least one of the following: Place the area where the product defect is located in the center of the background image; If the ratio of the area where the product defect is located to the background image is lower than a preset ratio threshold, the area where the product defect is located is enlarged so that the ratio of the area where the product defect is located to the background image is higher than the preset ratio threshold, and the enlarged area where the product defect is located is merged into the background image.
5. The method according to claim 1, characterized in that, The process of obtaining the first fused image based on the first mask image and the first stitched image includes: The first mask image and the first stitched image are merged to obtain the initial first fused image; Random noise is added to the region where the mask is located in the initial first fused image to obtain the first fused image.
6. A method for generating defect images, characterized in that, include: Acquire a reference defect image and a good product image, wherein the reference defect image contains a product defect of a specified type; The reference defect image and the good product image are stitched together to obtain a second stitched image; A second mask image is generated, the size of which is the same as the size of the second stitched image, and the second mask image includes a mask; A second fused image is obtained based on the second mask image and the second stitched image. The second fused image includes the reference defect image and the second mask occlusion image. The second mask occlusion image is generated from the good product image in the second stitched image, and a specified position is occluded by the mask. The second fused image is input into the defect generation model, wherein the defect generation model generates the specified type of product defect in the mask area on the good product image based on the product defect in the reference defect image in the second fused image, thereby obtaining a defect generation image; The defect generation model is trained using the method described in any one of claims 1-5.
7. The method according to claim 6, characterized in that, The position and size of the mask are determined randomly or specified by the user.
8. A training method for a defect removal model, characterized in that, include: Obtain a reference good product image and a real defect image, wherein the real defect image contains a product defect and the product in the real defect image is the same as the product in the reference good product image; The reference good product image and the actual defect image are stitched together to obtain a third stitched image; A third mask image is obtained, the size of which is the same as that of the third stitched image. The third mask image includes a mask, and the area where the mask is located corresponds to the location of the real defect image in the third stitched image. Based on the third mask image and random noise, a fourth mask image is obtained, wherein the region where the mask is located in the fourth mask image is the random noise; A third fused image is obtained based on the fourth mask image and the third stitched image. The third fused image includes the reference good product image and the half-occluded image. The half-occluded image is formed by fusing the real defect image and the random noise according to a preset ratio. The diffusion model is trained based on the third fused image, wherein the diffusion model predicts the product structure of the partially occluded image based on the reference good product image in the third fused image, and generates a defect removal image; The diffusion model is fine-tuned based on the defect-removed image and the reference good product image to obtain the defect removal model.
9. The method according to claim 8, characterized in that, In the third mask image, the area where the mask is located is black, and the other areas are white.
10. A defect removal method, characterized in that, include: Obtain a reference good product image and a defective product image, wherein the product in the defective product image is the same as the product in the reference good product image; The reference good product image and the defective image are stitched together to obtain a fourth stitched image; A fifth mask image is obtained, the size of which is the same as that of the fourth stitched image. The fifth mask image includes a mask, and the area where the mask is located corresponds to the location of the defect image in the fourth stitched image. Based on the fifth mask image and random noise, a sixth mask image is obtained, wherein the region where the mask is located in the sixth mask image is the random noise; A fourth fused image is obtained based on the sixth mask image and the fourth stitched image. The fourth fused image includes the reference good product image and the half-occluded image. The half-occluded image is formed by fusing the defect image and the random noise according to a preset ratio. The fourth fused image is input into the defect removal model, wherein the defect removal model predicts the defect image based on the reference good product image in the fourth fused image to generate a defect-removed image; The defect removal model is trained using the method described in claim 7 or 8.
11. A training device for a defect generation model, characterized in that, include: The first acquisition module is used to acquire a reference defect image and a real defect image, wherein the product defects in the reference defect image and the real defect image are of the same type. The stitching module is used to stitch the reference defect image and the real defect image together to obtain a first stitched image; The second acquisition module is used to generate a first mask image, the size of which is the same as the size of the first stitched image. The first mask image includes a mask, and the position and size of the mask are determined by the position and size of the product defect in the real defect image in the first stitched image. The fusion module is used to obtain a first fused image based on the first mask image and the first stitched image. The first fused image includes the reference defect image and the first mask occlusion image. The first mask occlusion image is generated from the real defect image in the first stitched image, and the product defect is occluded by the mask. A training module is used to train a diffusion model based on the first fused image, wherein the diffusion model generates a defect generation image based on the first fused image; The diffusion model is fine-tuned based on the generated defect image and the real defect image to obtain the defect generation model.
12. A defect image generation apparatus, characterized in that, include: The acquisition module is used to acquire a reference defect image and a good product image, wherein the reference defect image contains a product defect of a specified type; The stitching module is used to stitch the reference defect image and the good product image together to obtain a second stitched image; A mask generation module is used to generate a second mask image, the size of which is the same as the size of the second stitched image, and the second mask image includes a mask; A fusion module is used to obtain a second fused image based on the second mask image and the second stitched image. The second fused image includes the reference defect image and the second mask occlusion image. The second mask occlusion image is generated from the good product image in the second stitched image, and a specified position is occluded by the mask. A defect generation module is used to input the second fused image into a defect generation model, wherein the defect generation model generates the specified type of product defect in the mask area on the good product image based on the product defect in the reference defect image in the second fused image, thereby obtaining a defect generation image; The defect generation model is trained using the method described in any one of claims 1-5.
13. A training device for a defect removal model, characterized in that, include: The first acquisition module is used to acquire a reference good product image and a real defect image, wherein the real defect image contains product defects and the product in the real defect image is the same as the product in the reference good product image. The stitching module is used to stitch the reference good product image and the real defect image together to obtain a third stitched image; The second acquisition module is used to obtain a third mask image, the size of which is the same as the size of the third stitched image. The third mask image includes a mask, and the area where the mask is located corresponds to the location of the real defect image in the third stitched image. The third acquisition module is used to obtain a fourth mask image based on the third mask image and random noise, wherein the region where the mask is located in the fourth mask image is the random noise; The fusion module is used to obtain a third fused image based on the fourth mask image and the third stitched image. The third fused image includes the reference good product image and the half-occluded image. The half-occluded image is formed by fusing the real defect image and the random noise according to a preset ratio. The training module is used to train the diffusion model based on the third fused image, wherein the diffusion model predicts the product structure of the partially occluded image based on the reference good product image in the third fused image and generates a defect removal image; The diffusion model is fine-tuned based on the defect-removed image and the reference good product image to obtain the defect removal model.
14. A defect removal device, characterized in that, include: The first acquisition module is used to acquire a reference good product image and a defective product image, wherein the product in the defective product image is the same as the product in the reference good product image; The stitching module is used to stitch the reference good product image and the defective image together to obtain a fourth stitched image; The second acquisition module is used to obtain a fifth mask image, the size of which is the same as the size of the fourth stitched image. The fifth mask image includes a mask, and the area where the mask is located corresponds to the location of the defect image in the fourth stitched image. The third acquisition module is used to obtain a sixth mask image based on the fifth mask image and random noise, wherein the region where the mask is located in the sixth mask image is the random noise; The fusion module is used to obtain a fourth fused image based on the sixth mask image and the fourth stitched image. The fourth fused image includes the reference good product image and the half-occluded image. The half-occluded image is formed by fusing the defect image and the random noise according to a preset ratio. The defect removal module is used to input the fourth fused image into the defect removal model, wherein the defect removal model predicts the defect image based on the reference good product image in the fourth fused image to generate a defect-removed image; The defect removal model is trained using the method described in claim 8 or 9.
15. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when executed by the processor, the program implements the steps of a training method for a defect generation model as described in any one of claims 1 to 5; or, when executed by the processor, the program implements the steps of a defect image generation method as described in claim 6 or 7; when executed by the processor, the program implements the steps of a training method for a defect removal model as described in claim 8 or 9; or, when executed by the processor, the program implements the steps of a defect removal method as described in claim 10.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the training method for the defect generation model as described in any one of claims 1 to 5; or, when executed by a processor, implements the steps of the defect image generation method as described in claim 6 or 7; or, when executed by a processor, implements the steps of the training method for the defect removal model as described in claim 8 or 9; or, when executed by a processor, implements the steps of the defect removal method as described in claim 10.
17. A computer program product, characterized in that, The method includes computer instructions that, when executed by a processor, implement the steps of a training method for a defect generation model as described in any one of claims 1 to 5; or, when executed by a processor, implement the steps of a defect image generation method as described in claim 6 or 7; or, when executed by a processor, implement the steps of a training method for a defect removal model as described in claim 8 or 9; or, when executed by a processor, implement the steps of a defect removal method as described in claim 10.