A method, device, equipment and medium for generating a workpiece defect image

The diffusion model combines CLIP picture encoder and variational autoencoder to generate high-quality defect images, which solves the problems of low resolution and unnatural defect generation of GAN model, and achieves high-resolution and high-detail defect image generation.

CN119559276BActive Publication Date: 2025-07-25ZHIZI ENGINE (BEIJING) TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411509846.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-28
Publication Date
2025-07-25
Estimated Expiration
2044-10-28

AI Technical Summary

Technical Problem

The defect image generated by traditional GAN models has low resolution and is unnatural, difficult to train, and the defect generation intensity is difficult to control, and requires a lot of manual annotation work.

Method used

The diffusion model is used to combine CLIP picture encoder, multi-layer perceptron and variational autoencoder to generate high-quality defect images through cropping, data enhancement, mask shape enhancement and diffusion processing, and the LORA module is added to improve model adaptability.

Benefits of technology

The generated defect images have high resolution and clear details. They can flexibly control the position and shape of defects, adapt to different resolutions, and produce natural and smoothly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559276B_ABST
    Figure CN119559276B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method, apparatus, device, and medium for generating workpiece defect images. First, the input high-resolution defect image is cropped to obtain a slice containing the generated defect area, and then the image containing the defect area is cropped from it. After performing data augmentation operations, it is sequentially input into a CLIP image encoder and a multi-layer perceptron to obtain feature vectors, and then a mask shape enhancement process is performed on them to generate an image with a defect mask. The image with the defect mask is spliced with Gaussian noise and then input into a diffusion model together with the feature vectors for training, and then diffused to obtain a generated image. The generated image is pasted onto the input high-resolution defect image and then input into a variational autoencoder for decoding to obtain the original-resolution defect image. Adding LoRA makes this method more adaptable to the generation in the defect field, restoring the low-dimensional feature vectors to generate high-quality and high-resolution defect images, ensuring that the generated images have high clarity and details.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image generation, and more particularly, to a method, device, equipment and medium for generating workpiece defect images. Background Art

[0002] In traditional GAN-based methods, the generated defect images have low resolution and are not natural. This is because when the GAN model generates high-resolution and detail-rich images, it needs to process more pixels and more complex features, which increases the training difficulty. If the images in the training dataset have low resolution or insufficient diversity, the generated images will also be affected. Traditional GAN architectures may not be able to effectively capture the subtle features in the images, resulting in the generated images looking less realistic.

[0003] In addition, in order to make the generated defects more accurately conform to the real situation, a finely defined mask is usually required to indicate which areas should be modified. And providing a fine mask usually requires a large amount of manual annotation work, increasing the cost of data preparation.

[0004] Most methods cannot control the intensity of defect generation. Since the intensity of defect generation is usually determined by the internal parameters of the model, and these parameters are often difficult to directly adjust. At the same time, the specific needs of users for the defect intensity are not considered in the model design.

[0005] Therefore, a solution needs to be studied to solve the above technical problems. Summary of the Invention

[0006] The present application provides a method, device, equipment and medium for generating workpiece defect images based on a diffusion model. By inputting a picture into the diffusion model, the goal of generating high-quality and long-duration videos from a single picture is achieved, and the naturalness of face generation and the smoothness of video actions are improved.

[0007] To achieve the above object, the technical solutions adopted in the embodiments of the present application are as follows:

[0008] In a first aspect, an embodiment of the present application provides a method for generating workpiece defect images. The method is applied to a paint by example model, and the paint by example model includes a diffusion model with a lora module, a CLIP image encoder, a multi-layer perceptron, and a variational autoencoder, and includes:

[0009] Cropping the input high-resolution defect image to obtain a slice containing the generated defect area;

[0010] Cropping the slice containing the generated defect area to obtain a picture containing the defect area;

[0011] After performing data augmentation on the image containing the defective area, it is sequentially input into the CLIP image encoder and the multi-layer perceptron to obtain a feature vector;

[0012] Perform mask shape enhancement processing on the image containing the defective area to generate an image with a defective mask;

[0013] Stitch the image with the defective mask and Gaussian noise, and then input them together with the feature vector into the diffusion model for training to obtain a trained diffusion model. Use the trained diffusion model to perform diffusion to obtain a generated image;

[0014] Paste the generated image onto the input high-resolution defective image, and then input it into the variational autoencoder for decoding to obtain the original-resolution defective image.

[0015] In a possible implementation manner, the step of performing data augmentation on the image containing the defective area and then sequentially inputting it into the CLIP image encoder and the multi-layer perceptron to obtain a feature vector includes:

[0016] Input the image containing the defective area into the CLIP image encoder for encoding to obtain an image feature representation;

[0017] Input the image feature vector into the multi-layer perceptron for non-linear transformation to obtain a feature vector.

[0018] In a possible implementation manner, the step of performing mask shape enhancement processing on the image containing the defective area to generate an image with a defective mask includes:

[0019] Multiply the image containing the defective area with the mask to generate an image with a defective mask.

[0020] In a possible implementation manner, randomly select a defect-free area or a defective area image as the ground truth for training.

[0021] In a possible implementation manner, the method further includes:

[0022] If the slice containing the generated defective area does not meet the processing requirements, perform scaling.

[0023] In a second aspect, an embodiment of the present application further provides a workpiece defective image generation device, and the device includes:

[0024] A first cropping module, configured to crop the input high-resolution defective image to obtain a slice containing the generated defective area;

[0025] A second cropping module, configured to crop the slice containing the generated defective area to obtain an image containing the defective area;

[0026] A data augmentation module, which is used to perform data augmentation operations on pictures containing defect regions and then input them into a CLIP picture encoder and a multi-layer perceptron in sequence to obtain feature vectors;

[0027] A mask shape augmentation module, which is used to perform mask shape augmentation processing on pictures containing defect regions to generate images with defect masks;

[0028] A training and diffusion module, which is used to splice the image with the defect mask and Gaussian noise and then input them together with the feature vectors into a diffusion model for training to obtain a trained diffusion model, and use the trained diffusion model to perform diffusion to obtain generated pictures;

[0029] A decoding module, which is used to paste the generated picture onto the input high-resolution defect image and then input it into a variational autoencoder for decoding to obtain the original-resolution defect picture.

[0030] In a possible implementation manner, the data augmentation module is further used for:

[0031] Input the picture containing the defect region into the CLIP picture encoder for encoding to obtain a picture feature representation;

[0032] Input the picture feature vector into the multi-layer perceptron for non-linear conversion to obtain a feature vector.

[0033] In a possible implementation manner, the mask shape augmentation module is further used for:

[0034] Multiply the picture containing the defect region with the mask to generate an image with a defect mask.

[0035] In a third aspect, the present application also proposes a computer device, which includes a processor and a memory. A computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the workpiece defect image generation method as described in any item of the first aspect.

[0036] In a fourth aspect, the present application also proposes a computer-readable storage medium. A computer program is stored in the storage medium, and the computer program is loaded and executed by a processor to implement the workpiece defect image generation method as described in any item of the first aspect.

[0037] The main solution of the present application and its various further alternative solutions can be freely combined to form multiple solutions, all of which are solutions that can be adopted and claimed by the present application; and in the present application, (each non-conflicting alternative) alternatives can be freely combined with each other and with other alternatives. Those skilled in the art can understand that there are various combinations according to the prior art and common general knowledge after understanding the solution of the present application, all of which are the technical solutions to be protected by the present application, and will not be enumerated here.

[0038] Compared with the prior art, embodiments of the present application provide a method, apparatus, device, and medium for generating workpiece defect images. First, the input high-resolution defect image is cropped to obtain a slice containing the generated defect area, and then the image containing the defect area is cropped from it. After performing data augmentation operations, it is sequentially input into a CLIP image encoder and a multi-layer perceptron to obtain feature vectors, and then a mask shape enhancement process is performed on them to generate an image with a defect mask. The image with the defect mask is spliced with Gaussian noise and then input into a diffusion model together with the feature vectors for training, and then diffused to obtain a generated image. The generated image is pasted onto the input high-resolution defect image and then input into a variational autoencoder for decoding to obtain the original-resolution defect image. Adding LoRA makes the method more adaptable to the generation in the defect field, restoring the low-dimensional feature vectors to generate high-quality and high-resolution defect images, ensuring that the generated images have high clarity and details. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0040] Figure 1 The flowchart of a method for generating workpiece defect images proposed by an embodiment of the present application is shown.

[0041] Figure 2 The schematic diagram of the paintby example model proposed by an embodiment of the present application is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.

[0043] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0044] It should be noted that, without conflict, the features in the embodiments of the present application can be combined with each other.

[0045] Please refer to Figure 1 and Figure 2 , Figure 1 which shows a schematic flowchart of a method for generating a workpiece defect image proposed in an embodiment of the present application. Figure 2 which shows a schematic diagram of the paint by example model proposed in an embodiment of the present application. This method is applied to the paint by example model, and the paint by example model includes a diffusion model with a LORA module, a CLIP image encoder, a multi-layer perceptron, and a variational autoencoder.

[0046] The diffusion model is a deep learning model for generating high-quality images. It restores the image by gradually removing the noise added to the original image, thereby generating a new image.

[0047] The LORA (Low-Rank Adaptation) module can optimize the model performance by adding a small number of trainable parameter layers without modifying the parameters of the base model, enabling the model to better adapt to specific tasks or datasets.

[0048] The CLIP (Contrastive Language–Image Pre-training) image encoder is a multi-modal pre-training model that can encode an image into a feature vector for extracting high-level features from the input image. These features are then fed into a multi-layer perceptron for further processing.

[0049] The multi-layer perceptron (MLP) is a feed-forward neural network used to further process the feature vector obtained from the CLIP encoder, helping the model better understand and utilize these features, thereby improving the quality of the generated image.

[0050] The variational autoencoder (VAE) is a generative model for image reconstruction. It reconstructs the image generated by the diffusion model into a high-resolution defect image by learning a latent space distribution and mapping it back to the image space during the decoding process, thereby generating high-quality outputs.

[0051] The paint by example model in the embodiments of the present application combines a diffusion model with a LORA module, a CLIP image encoder, a multi-layer perceptron, and a variational autoencoder to execute the defect image generation method. It can not only generate realistic defect images but also significantly improve the adaptability and generation effect of the model in specific fields through the introduction of the LORA module.

[0052] The method for generating workpiece defect images includes:

[0053] Step S1: Crop the input high-resolution defect image to obtain a slice containing the generated defect area.

[0054] Use existing detection algorithms or manual marking to determine the defect areas in the image. Generate a bounding box for each defect area. Crop the image according to the bounding box to obtain a slice (512×512) containing the defect area, and the size of the input original image is 512×512 pixels.

[0055] If the slice containing the generated defect area does not meet the processing requirements, perform scaling.

[0056] After cropping, if the size of the cropped image does not meet the requirements of subsequent processing, it needs to be scaled to meet specific size requirements.

[0057] Step S2: Crop the slice containing the generated defect area to obtain a picture containing the defect area.

[0058] Further crop the already cropped slice to obtain small images containing individual defect areas. Based on the preliminary cropping, for each slice containing multiple defect areas, crop again to ensure that each small image contains only a single defect area, and finally obtain a series of small images, each image containing only one defect area (picture containing the defect area).

[0059] Step S3: After performing data augmentation operations on the pictures containing the defect areas, input them into the CLIP picture encoder and multi-layer perceptron in sequence to obtain feature vectors.

[0060] Input the pictures containing the defect areas into the CLIP picture encoder to extract the semantic features of the images. Further process the feature vectors output by the CLIP encoder through a multi-layer perceptron (MLP) to generate more representative feature vectors, and input the processed feature vectors into the diffusion model as part of the conditional input.

[0061] Step S3 includes:

[0062] Input the pictures containing the defect areas into the CLIP picture encoder for encoding to obtain picture feature representations;

[0063] Input the picture feature vectors into the multi-layer perceptron for non-linear transformation to obtain feature vectors.

[0064] Input the image containing the defect area into the CLIP image encoder. The CLIP image encoder encodes the image to generate a vector representing the features of the image (image feature representation). Input the image feature vector obtained from the CLIP image encoder into a multi-layer perceptron. The multi-layer perceptron performs a non-linear transformation on the feature vector through a multi-layer neural network to generate the final feature vector.

[0065] Step S4: Perform mask shape enhancement processing on the image containing the defect area to generate an image with a defect mask.

[0066] The image containing the defect area refers to replacing the area where the defect needs to be generated (or containing the defect) with an irregular-shaped black layer on the original image, or covering the area where the defect needs to be generated with an irregular-shaped black layer, defining one or more areas containing defects, and generating different-shaped masks m through operations such as random distortion and magnification to specify the areas that need to be processed by the model to generate defects.

[0067] Multiply the image containing the defect area with the mask to generate an image with a defect mask.

[0068] The mask m is a binary layer with the same size as the input image. In the mask, the positions with a value of 0 represent the areas where defects need to be generated, and these positions will be replaced with a black layer in the final image. In the mask, the positions with a value of 1 represent the parts that do not need to be replaced, and these positions remain unchanged in the final image. Apply the mask m to the image containing the defect area to generate an image with black irregular-shaped occlusions, obtaining an image with the defect area occluded by a black layer. Concatenate the processed image with Gaussian noise to form the input data.

[0069] Step S5: Concatenate the image with the defect mask with Gaussian noise and input them together with the feature vector into the diffusion model for training to obtain the trained diffusion model. Use the trained diffusion model for diffusion to generate images.

[0070] During the training process, randomly selected defect-free area or defect area images are used as the ground truth for training. During the generation process, the generation intensity of the defect can be controlled by adjusting the strength, so as to achieve more flexible image generation.

[0071] In the generation stage, given a highly noisy initial state, the model gradually performs denoising operations until it reaches the final state. The entire diffusion process is a recursive denoising loop, and each iteration makes the output image closer to the target. Starting from a highly noisy image, the model predicts and removes a portion of the noise to obtain a new image, repeating the previous steps until the final state is reached. When a satisfactory result is obtained, it is placed back into the corresponding position in the original resolution image, and the output quality is further optimized through a variational autoencoder (VAE) to ensure that the generated defective image naturally blends into the background, with a natural and smooth overall visual effect.

[0072] The image after masking the defective area is spliced with noise to form a new input. The spliced image and the feature vector are input into the diffusion model together. A LoRA module is added to the diffusion model, enabling the model to better adapt to tasks in specific domains, such as defect generation. The model generates the repaired defective area.

[0073] Step S6: Paste the generated image after the input high-resolution defective image and then input it into the variational autoencoder for decoding to obtain the original resolution defective image.

[0074] Put the generated defective area back into the original high-resolution image. Input the pasted image into the VAE for decoding, and finally obtain the original resolution defective image.

[0075] In summary, this application can crop the defective area from a high-resolution defective image, generate diverse defect shapes, extract feature vectors using CLIP and MLP, and combine the LoRA-adapted diffusion model to generate defective images. Finally, a high-quality original resolution defective image is obtained through VAE decoding.

[0076] Compared with the prior art, the embodiments of this application have the following beneficial effects:

[0077] First, by providing different types of masks, users can flexibly control the position and shape of the defects, thereby generating more diverse defective images, enabling the model to generate results that better meet the requirements according to the user's guidance, and enhancing the diversity and controllability of the output.

[0078] Second, it can generate defective images at different image resolutions, and the generated defective images maintain high quality, without obvious blurring or distortion due to an increase in resolution. Even at high resolutions, the generated images still maintain high details and clarity, and the quality does not degrade due to compression algorithms.

[0079] Third, the strength of defect generation can be adjusted by adjusting the strength parameter.

[0080] The following presents a possible implementation of a workpiece defect image generation device, which is used to execute each execution step and corresponding technical effect of the workpiece defect image generation method shown in the above embodiments and possible implementations. The device includes:

[0081] A first cropping module, configured to crop the input high-resolution defect image to obtain a slice containing the generated defect area;

[0082] A second cropping module, configured to crop the slice containing the generated defect area to obtain a picture containing the defect area;

[0083] A data augmentation module, configured to perform data augmentation operations on the picture containing the defect area and then input it into a CLIP image encoder and a multi-layer perceptron in sequence to obtain a feature vector;

[0084] A mask shape augmentation module, configured to perform mask shape augmentation processing on the picture containing the defect area to generate an image with a defect mask;

[0085] A training and diffusion module, configured to splice the image with the defect mask and Gaussian noise and then input them together with the feature vector into a diffusion model for training to obtain a trained diffusion model, and use the trained diffusion model for diffusion to obtain a generated picture;

[0086] A decoding module, configured to paste the generated picture onto the input high-resolution defect image and then input it into a variational autoencoder for decoding to obtain a defect picture with the original resolution.

[0087] In a possible implementation, the data augmentation module is further configured to:

[0088] Input the picture containing the defect area into the CLIP image encoder for encoding to obtain a picture feature representation;

[0089] Input the picture feature vector into the multi-layer perceptron for non-linear transformation to obtain a feature vector.

[0090] In a possible implementation, the mask shape augmentation module is further configured to:

[0091] Multiply the picture containing the defect area with the mask and then generate an image with a defect mask.

[0092] This preferred embodiment provides a computer device, which can implement the steps in any embodiment of the workpiece defect image generation method provided in the embodiments of the present application. Therefore, the beneficial effects of the workpiece defect image generation method provided in the embodiments of the present application can be achieved. For details, see the previous embodiments and will not be elaborated here.

[0093] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this purpose, an embodiment of the present application provides a storage medium, which stores multiple instructions that can be loaded by a processor to execute the steps of any one of the embodiments of the workpiece defect image generation method provided by the embodiments of the present application.

[0094] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0095] Since the instructions stored in the storage medium can execute the steps in any one of the embodiments of the workpiece defect image generation method provided by the embodiments of the present application, the beneficial effects achievable by any one of the workpiece defect image generation methods provided by the embodiments of the present application can be achieved. For details, see the previous embodiments and will not be repeated here.

[0096] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for generating a workpiece defect image, characterized in that, The method is applied to a paint by example model, which includes a diffusion model with a LoRA module, a CLIP image encoder, a multi-layer perceptron, and a variational autoencoder, and includes: Cropping the input high-resolution defect image to obtain a slice containing the generated defect area; Cropping the slice containing the generated defect area to obtain an image containing the defect area; After performing data augmentation operations on the image containing the defect area, sequentially inputting it into the CLIP image encoder and the multi-layer perceptron to obtain a feature vector; Performing mask shape enhancement processing on the image containing the defect area to generate an image with a defect mask; Concatenating the image with the defect mask and Gaussian noise, and then inputting them together with the feature vector into the diffusion model for training to obtain a trained diffusion model, and using the trained diffusion model for diffusion to obtain a generated image; Pasting the generated image onto the input high-resolution defect image, and then inputting it into the variational autoencoder for decoding to obtain the original-resolution defect image; The step of performing data augmentation operations on the image containing the defect area and then sequentially inputting it into the CLIP image encoder and the multi-layer perceptron to obtain a feature vector includes: Inputting the image containing the defect area into the CLIP image encoder for encoding to obtain an image feature representation; Inputting the image feature vector into the multi-layer perceptron for non-linear transformation to obtain a feature vector; The step of performing mask shape enhancement processing on the image containing the defect area to generate an image with a defect mask includes: Multiplying the image containing the defect area with the mask to generate an image with a defect mask; Using a randomly selected defect-free area or a defect area image as the ground truth for training; The method further includes: Scaling if the slice containing the generated defect area does not meet the processing requirements.

2. An apparatus for generating a workpiece defect image, characterized in that, The device includes: A first cropping module for cropping the input high-resolution defect image to obtain a slice containing the generated defect area; A second cropping module for cropping the slice containing the generated defect area to obtain an image containing the defect area; A data augmentation module for performing data augmentation operations on the image containing the defect area and then sequentially inputting it into the CLIP image encoder and the multi-layer perceptron to obtain a feature vector; A mask shape enhancement module for performing mask shape enhancement processing on the image containing the defect area to generate an image with a defect mask; A training and diffusion module for concatenating the image with the defect mask and Gaussian noise, and then inputting them together with the feature vector into the diffusion model for training to obtain a trained diffusion model, and using the trained diffusion model for diffusion to obtain a generated image; A decoding module for pasting the generated image onto the input high-resolution defect image and then inputting it into the variational autoencoder for decoding to obtain the original-resolution defect image; The data augmentation module is further configured to: input the image containing the defect area into the CLIP image encoder for encoding to obtain an image feature representation; input the image feature vector into the multi-layer perceptron for non-linear transformation to obtain a feature vector; The mask shape enhancement module is further configured to: generate an image with a defect mask after multiplying the image containing the defect area by the mask; The training and diffusion module is further configured to: use randomly selected defect-free area or defect area images as the ground truth for training; scale the slices containing the generated defect areas if they do not meet the processing requirements.

3. A computer device, characterized in that, The computer device includes a processor and a memory, and a computer program is stored in the memory. The computer program is loaded and executed by the processor to implement the workpiece defect image generation method according to claim 1.

4. A computer-readable storage medium, characterized in that, A computer program is stored in the storage medium. The computer program is loaded and executed by a processor to implement the workpiece defect image generation method according to claim 1.

Citation Information

Patent Citations

  • Defect image generation method and device, computer equipment and storage medium

    CN117953321A