Abnormal image generation method and device based on few-sample diffusion model learning

Through the abnormal image generation method based on the few-sample diffusion model learning, the abnormal information feature encoding model and feature adaptive weighting model are used, combined with the potential diffusion model, the problems of insufficient feature extraction of abnormal samples and insufficient accuracy in generating abnormal samples in the prior art are solved, and efficient and real abnormal image generation is achieved.

CN120071086APending Publication Date: 2025-05-30PANOVASIC TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510149050.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing anomaly image generation methods are difficult to extract and utilize abnormal sample features in limited data samples, resulting in limited performance of the abnormal detection algorithm and insufficient accuracy and realisticity of generating abnormal samples.

Method used

An abnormal image generation method based on the learning of the few-sample diffusion model is adopted. By collecting and data augmenting the abnormal image and mask map, an abnormal information feature coding model and feature adaptive weighting model are constructed, and combined with the potential diffusion model, end-to-end abnormal image generation is achieved.

Benefits of technology

Effectively learn abnormal features in the case of few samples, generate more realistic and realistic abnormal images, and improve the data support capability of the abnormal detection algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071086A_ABST
    Figure CN120071086A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of image anomaly detection and image generation, discloses an abnormal image generation method and device based on few-sample diffusion model learning, and solves the problems that an existing abnormal image generation method is insufficient in abnormal sample feature extraction and utilization and insufficient in abnormal sample generation accuracy and verisimilitude. According to the scheme of the invention, the method comprises the steps: collecting and making an abnormal image and a mask image, carrying out the data augmentation operation under the condition of few samples, and obtaining a data set; constructing an abnormal information feature coding model used for extracting appearance, position and semantic information of the abnormal region; constructing a feature adaptive weighting model used for optimizing abnormal image generation according to the pixel-level difference between an intermediate image generated in an abnormal image generation process and a normal image sample in an abnormal mask area; and constructing an end-to-end diffusion model based on the potential diffusion model, the abnormal information feature coding model and the feature adaptive weighting model, and finally training the diffusion model by using the data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of image anomaly detection and image generation, and particularly relates to an abnormal image generation method and device based on few-shot diffusion model learning. Background Art

[0002] Currently, image anomaly detection algorithms are widely used in the industrial manufacturing field, and play a key role in ensuring product quality and improving production efficiency. These algorithms usually not only need to be able to identify anomalies, but also need to accurately locate the specific positions where the anomalies occur and classify them, so as to take targeted corrective measures and post-processing. However, due to the scarcity of abnormal events in the production process, the number of abnormal samples that can be collected and sorted out for training is limited, which poses a severe challenge to the development and optimization of deep learning-based anomaly detection algorithms.

[0003] Current anomaly detection methods are mainly divided into unsupervised learning and few-shot supervised learning based on deep learning. Among them, unsupervised learning methods can identify abnormal patterns without abnormal sample labels, but it is difficult to accurately locate and classify anomalies, especially under the conditions of multiple products and unknown anomaly types. Although few-shot supervised learning methods can be trained using limited abnormal samples, they are highly dependent on abnormal samples. When the number of samples is insufficient or the quality is not high, the performance of the algorithm is severely limited.

[0004] To alleviate the problem of insufficient abnormal samples, researchers have proposed abnormal sample generation technologies, aiming to artificially synthesize abnormal samples to expand the dataset by manually adding abnormal states to normal samples. These technologies are mainly divided into two categories:

[0005] 1) Model-free traditional methods: These methods create synthetic abnormal samples by randomly cropping and pasting random image patches from abnormal texture datasets onto normal samples. This method is direct and easy to implement, but its main disadvantage is that the synthetic abnormal samples may be very different from the actual abnormal situations in production, resulting in insufficient authenticity of the synthetic data, which in turn affects the training effect of the anomaly detection algorithm.

[0006] 2) Methods based on generative adversarial networks (GANs): These methods utilize the powerful generation ability of GANs to create and generate abnormal samples. A GAN consists of a generator and a discriminator. The generator is responsible for generating abnormal samples, while the discriminator evaluates the authenticity of these samples. Although GANs have achieved remarkable results in the field of image generation, most GAN models require a large number of abnormal samples for training to simulate real abnormal generation situations, which is not realistic for industrial scenarios with scarce abnormal samples.

[0007] Generally speaking, the existing abnormal image generation methods face the following problems:

[0008] 1. It is difficult to extract and utilize the features of abnormal samples in a limited data sample, which restricts the performance of downstream anomaly detection algorithms;

[0009] 2. The lack of in-depth research and utilization of deep learning technologies such as self-supervised learning and transfer learning leads to insufficient understanding of the simulation and generation of anomalies;

[0010] 3. There is a lack of accuracy and verisimilitude in generating abnormal samples. Summary of the Invention

[0011] The technical problem to be solved by the present invention is: to propose an abnormal image generation method and device based on few-shot diffusion model learning, and solve the problems existing in the existing abnormal image generation methods, such as insufficient extraction and utilization of abnormal sample features, and insufficient accuracy and verisimilitude in generating abnormal samples.

[0012] The technical solution adopted by the present invention to solve the above technical problems is:

[0013] On the one hand, the present invention provides an abnormal image generation method based on few-shot diffusion model learning, including:

[0014] A. Collect and produce abnormal images for different types of anomalies of the product to be detected and the abnormal region mask images corresponding to the abnormal images respectively;

[0015] B. Perform data augmentation operations on the abnormal images of different types of anomalies and the abnormal region mask images corresponding to the abnormal images respectively to obtain a data set;

[0016] C. Construct an abnormal information feature encoding model for extracting the appearance, position, and semantic information of the abnormal region;

[0017] D. Construct a feature adaptive weighting model for optimizing the generation of abnormal images according to the pixel-level difference between the intermediate images generated during the generation of abnormal images and the normal image samples in the abnormal mask region;

[0018] E. Based on the latent diffusion model, the abnormal information feature encoding model, and the feature adaptive weighting model, construct an end-to-end diffusion model;

[0019] F. Based on the data set obtained in step B, train the diffusion model constructed in step E;

[0020] G. When performing the abnormal image generation task, use the normal image and the randomly generated abnormal region mask image as inputs, and use the trained diffusion model to obtain the abnormal image and the abnormal region mask image corresponding to the abnormal image.

[0021] Further, in step A, the method for separately collecting and creating abnormal images for different types of abnormalities of the product to be detected and the mask images of the abnormal regions corresponding to the abnormal images includes:

[0022] For the product to be abnormally detected, image samples are collected through an industrial imaging device. After annotating the abnormal regions in the abnormal images, they are then converted into mask images to obtain the abnormal images and the corresponding mask images of the abnormal regions.

[0023] Further, in step B, the method for performing data augmentation operations includes:

[0024] Before performing augmentation, first calculate the coordinate positions of the augmented abnormal mask regions according to a randomly combined augmentation strategy, and determine whether they exceed the boundaries of the original image. If they exceed the boundaries, re-randomly combine the image augmentation methods until the coordinate conditions for image augmentation are met. If they do not exceed the boundaries, use the randomly combined augmentation strategy to augment the abnormal images and the corresponding mask images of the abnormal regions.

[0025] Further, the randomly combined augmentation strategy includes, but is not limited to, a random combination of one or more image augmentation methods such as random inversion, random central rotation, contrast change, brightness change, central cropping, random translation, distortion transformation, random noise, and convolutional filtering.

[0026] Further, in step C, the constructed abnormal information feature encoding model includes:

[0027] An abnormal region appearance depth feature extraction network, which takes the abnormal image and the corresponding mask image of the abnormal region as inputs and outputs an abnormal region appearance feature vector;

[0028] An abnormal region position depth feature extraction network, which takes the mask image of the abnormal region corresponding to the abnormal image as an input and outputs an abnormal region position feature vector;

[0029] An abnormal region semantic depth feature extraction network, which takes the product abnormal text as an input and outputs an abnormal region semantic feature vector. The product abnormal text is created and saved independently when creating the dataset.

[0030] Further, both the abnormal region appearance depth feature extraction network and the abnormal region position depth feature extraction network include multiple feature extraction units in the downsampling part and the upsampling part. Among them, the feature extraction units in the downsampling part are composed of a max-pooling layer, a convolutional layer, a batch normalization layer, and a Relu activation layer stacked in sequence; the feature extraction units in the upsampling part are composed of a transposed convolutional layer, a convolutional layer, a batch normalization layer, and a Relu activation layer stacked in sequence.

[0031] Further, the abnormal region semantic depth feature extraction network uses word embedding technology to map words in the product abnormal text into low-dimensional dense vectors, and then performs feature extraction on the text through context embedding, sentence-level embedding, or using a deep learning model, and converts the mapped vectors into abnormal region semantic feature vectors.

[0032] Further, in step D, the feature adaptive weighting model uses an image similarity algorithm to calculate the pixel-level difference between the intermediate image generated during the abnormal image generation process and the normal image sample in the abnormal mask region, and calculates a weight map based on this difference according to the adaptive scaling operation and the self-attention mechanism, and obtains an adaptive weight map to adjust and optimize the generated intermediate image.

[0033] Further, in step E, when the constructed diffusion model generates an abnormal image, it takes the normal image sample and the random mask image as inputs, extracts the abnormal region appearance feature vector, position feature vector, and semantic feature vector of the random mask image through the abnormal information feature encoding model, and combines them with the pixel-level multiplication into the attention mapping module of the latent diffusion model to achieve the embedding of abnormal information, guides the latent diffusion model to generate an intermediate image containing abnormal features, and calculates a weight map based on the pixel-level difference between this intermediate image and the normal image sample in the abnormal mask region through the feature adaptive weighting model to adjust and optimize the generated intermediate image, and finally generates an abnormal image that only has abnormalities in the mask region and the other parts are consistent with the normal sample image.

[0034] On the other hand, the present invention also provides an abnormal image generation device based on few-shot diffusion model learning, which includes:

[0035] A data processing module, configured to collect and produce abnormal images for different types of abnormalities of the product to be detected and the abnormal region mask maps corresponding to the abnormal images respectively, and perform data augmentation operations on the abnormal images of different types of abnormalities and the abnormal region mask maps corresponding to the abnormal images;

[0036] A diffusion model learning module, configured to extract features from the abnormal images and the corresponding masks and serialize them into abnormal information feature vectors, calculate a feature adaptive weight map through the abnormal mask, and realize the learning of the diffusion model for abnormal generation based on the abnormal information characteristic vectors and the feature adaptive weight map;

[0037] A diffusion model inference module, configured to generate a random mask image according to the input normal image sample, and input the generated mask image and the normal image sample into the trained diffusion model to generate an abnormal image and the corresponding mask image.

[0038] The beneficial effects of the present invention are:

[0039] (1) By leveraging the powerful capabilities of the latent diffusion model, the generation of abnormal images and their corresponding masks is achieved.

[0040] (2) By collecting a small number of abnormal image and mask pairs and combining local augmentation operations on the dataset, the amount of data is expanded, enabling the model to learn sufficient abnormal features even in the few-shot scenario.

[0041] (3) Through the abnormal information feature encoding model, the location, appearance, and semantic features of the abnormal regions are extracted, helping the diffusion model to better understand and grasp the abnormal information, and providing guarantee for generating more realistic and practical abnormal images.

[0042] (4) Through the feature adaptive weighting model based on the abnormal mask information, the attention of the model to the regions of the generated abnormal images that are not obvious is enhanced, improving the quality of the generated abnormal images.

[0043] (5) Based on the end-to-end diffusion model structure design, the model can be efficiently trained and produce a large number of effective abnormal images and their corresponding masks, thus providing sufficient data support for downstream abnormal detection algorithms. Brief Description of the Drawings

[0044] Figure 1 It is a flowchart of the abnormal image generation method based on few-shot diffusion model learning in the present invention;

[0045] Figure 2 It is a structural diagram of the deep feature extraction network for the location of abnormal regions in the present invention;

[0046] Figure 3 It is a structural diagram of the deep feature extraction network for the appearance of abnormal regions in the present invention;

[0047] Figure 4 It is a schematic diagram of the principle of the deep feature extraction network for the semantics of abnormal regions in the present invention;

[0048] Figure 5 It is an example diagram of the network structure of the feature extraction unit in the downsampling part of the present invention;

[0049] Figure 6 It is an example diagram of the network structure of the feature extraction unit in the upsampling part of the present invention;

[0050] Figure 7 It is a schematic diagram of the structure of the feature adaptive weighting model based on the abnormal mask in the present invention;

[0051] Figure 8 It is a schematic diagram of the network structure of the abnormal image generation method based on few-shot diffusion model learning in the present invention;

[0052] Figure 9Schematic diagram of the structure of the abnormal image generation device based on few-shot diffusion model learning in the present invention. Detailed implementation manners

[0053] The present invention aims to provide a method and device for generating abnormal images based on few-shot diffusion model learning, and solve the problems existing in the existing abnormal image generation methods, such as insufficient extraction and utilization of abnormal sample features, and insufficient accuracy and vividness of generated abnormal samples. The core idea is as follows: Few-shot data processing: For typical abnormalities of the product to be inspected, a small number of abnormal image samples are collected through an industrial imaging device, and are labeled and converted into mask images using open-source tools. At the same time, local augmentation of the image and mask datasets is performed under few-shot conditions to expand the data volume and alleviate the problem of insufficient samples. Construct an abnormal information feature encoding model: By constructing an abnormal information feature encoding model that includes position, appearance, and semantics, neural network components are used to perform deep feature extraction on the appearance, position, and semantic information of the abnormal region respectively, enhancing the discrimination and utilization rate of limited image features. Construct a feature adaptive weighting model based on the abnormal mask: By calculating the pixel-level difference between the intermediate image generated by the diffusion model and the normal sample in the abnormal mask region, combined with adaptive scaling operations and self-attention mechanisms to calculate the weight mapping, the attention of the model to the regions where less obvious abnormal images are generated is improved, enhancing the overall abnormal generation ability. End-to-end model learning and generation: Based on the latent diffusion model (LDM), an end-to-end diffusion model learning and generation architecture is developed, integrating the abnormal information feature encoding model and the feature adaptive weighting model. After training the model using the augmented dataset, when an anomaly-free sample and a random mask are input, a large number of image samples with anomalies only at the mask can be efficiently and accurately generated, providing sufficient data support for downstream anomaly detection algorithms.

[0054] The solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0055] Embodiment 1

[0056] The process of the abnormal image generation method based on few-shot diffusion model learning provided in this embodiment is shown in Figure 1 and includes the following steps:

[0057] A. Collect and produce pairs of abnormal images and masks

[0058] In this step, for various typical abnormalities of the product to be inspected for anomalies, a small number of abnormal image samples are respectively collected through an industrial imaging device. After labeling the abnormal images, they are then converted into mask images, thereby obtaining pairs of abnormal images and masks (that is, abnormal images and the corresponding mask images of the abnormal images).

[0059] In an exemplary implementation manner, the specific implementation of this step is as follows:

[0060] Collect abnormal image samples: For the products to be inspected, use professional industrial imaging equipment (including cameras, lenses, and light sources) to obtain image data containing standard abnormalities. Taking metal surface defect detection as an example, use a high-resolution industrial camera to take pictures of the metal surface, ensuring that these pictures contain predefined abnormal types, such as cracks, dents, and scratches, etc. The number of images collected for each abnormal category is less than 10.

[0061] Abnormal area annotation: Use open-source image annotation tools, such as LabelImg, VGG ImageAnnotator, or Labelme, etc., to perform polygon manual annotation on the collected abnormal images, and accurately mark the pixel-level positions of the abnormal areas.

[0062] Mask image generation: Based on the coordinates of the annotated abnormal areas, batch generate binary mask images using the OpenCV library. In the mask image, the pixel values of the abnormal areas are set to 255 (white), and the pixel values of the normal areas are set to 0 (black) for subsequent processing.

[0063] B. Data augmentation for abnormal images and mask pairs

[0064] In this step, perform data augmentation operations on the abnormal images of different types of abnormalities and the mask images of the abnormal areas corresponding to the abnormal images respectively. Before applying specific image enhancement techniques to the image mask, first determine the new coordinate positions of the enhanced abnormal mask areas according to the randomly generated enhancement strategy, and check whether these coordinates exceed the boundaries of the original image. If it is found that the coordinates exceed the boundaries, re-randomly select the image enhancement technique until the coordinate conditions are met to ensure that the abnormal areas always remain within the boundaries of the image during the image enhancement process.

[0065] In an exemplary implementation, before performing the augmentation operation, record the coordinates of the abnormal areas of each image to be augmented based on the mask image, so as to track and check the position changes of the abnormal areas during the augmentation process. For the paired abnormal images and masks, calculate and judge whether the coordinates of the augmented abnormal areas exceed the original image according to the image augmentation strategy. If they exceed, re-randomly select the augmentation method until the coordinates of the augmented abnormal areas are within the original image.

[0066] In terms of the augmentation strategy, a combination of one or more image augmentation methods can be adopted, such as random inversion, translation, rotation, contrast change, convolution filtering, etc. Augment the image and the corresponding mask in the same augmentation manner, and save the augmented image pairs. Finally, divide all the augmented image mask pairs into a training set and a validation set according to different products to be inspected, models, and abnormal types for subsequent model training and evaluation.

[0067] C. Construct an Abnormal Information Feature Encoding Model

[0068] In this step, network structure models for extracting and processing the appearance of abnormal regions, the location of abnormal regions, and abnormal semantic features are respectively constructed through components of a neural network to achieve relatively comprehensive extraction of abnormal information features, providing sample feature conditions (constraints) for abnormal generation in the diffusion model.

[0069] In an exemplary implementation, the abnormal information feature encoding model consists of three parts, namely: the deep feature extraction network for the appearance of abnormal regions, the deep feature extraction network for the location of abnormal regions, and the deep feature extraction network for the semantics of abnormal regions, which are specifically described as follows:

[0070] See Figure 2 , the deep feature extraction network for the location of abnormal regions takes the mask map of the abnormal region corresponding to the abnormal image as input and outputs the feature vector of the abnormal region location. It is built by feature extraction components in the downsampling part and the upsampling part. Among them, the structure of the feature extraction component in the downsampling part is shown in Figure 5 , which is composed of a max pooling layer, a convolutional layer with a kernel size of 3, a batch normalization layer, and an activation layer with a ReLU activation function stacked in sequence to achieve downsampling and feature extraction of the input features; the structure of the feature extraction component in the upsampling part is shown in Figure 6 , which is composed of a transposed convolutional layer with a kernel size of 3, a convolutional layer with a kernel size of 3, a batch normalization layer, and an activation layer with a ReLU activation function stacked in sequence to achieve upsampling and feature extraction of the input features.

[0071] In Figure 2 In the shown deep feature extraction network structure for the location of abnormal regions, the number of channels, height, and width of the input abnormal mask are 1, 640, and 640 respectively. This network structure is designed as a typical multi-level encoding and decoding structure, and skip connections (i.e., referring to the Figure 2 operation process in ) are added to combine feature maps of the same size to enhance the stability of the extracted features. Specifically, first, the input abnormal mask (1, 640, 640) undergoes continuous 5 - time downsampling and high - dimensional feature mapping to obtain a feature map of size (512, 16, 16). Then, the high - dimensional features with compressed size are restored to a feature size and dimension of (32, 256, 256) through continuous upsampling and feature extraction components. Finally, the decoded features are input into a multi - layer perceptron MLP to serialize them into abnormal location tokens for subsequent processing by the diffusion backbone network model.

[0072] See Figure 3, the appearance depth feature extraction network for the abnormal area takes the abnormal image and the corresponding abnormal area mask image as inputs and outputs the appearance feature vector of the abnormal area; its network structure is also based on Figure 5 and Figure 6 The shown feature extraction components are built, and it is also designed as a typical multi-level encoding and decoding structure, and skip connection feature operations are added to enhance the stability of the extracted features. Specifically, first, the abnormal image and the abnormal mask are subjected to a pixel-level multiplication operation to obtain an image that only contains the pixels of the abnormal area, and its dimension size is (3, 256, 256). Then, the abnormal area image undergoes 5 consecutive downsamplings and high-dimensional feature mappings to obtain a feature map of size (1024, 16, 16). Then, the obtained high-dimensional features are used to restore the feature size and dimension to (64, 256, 256) through consecutive upsampling feature extraction components. Finally, the features from the previous step are input into a multi-layer perceptron MLP to serialize them into abnormal appearance tokens.

[0073] Figure 4 Schematically shows the process of extracting abnormal semantic information features. The Word2Vec technology is used to map the words in the product abnormal text into low-dimensional dense vectors, and feature vectors are constructed through a context embedding algorithm. Finally, through feature serialization, the process from abnormal text input to abnormal semantic token output is realized. Among them, the product abnormal text is created and saved independently when the dataset is initially constructed. Figure 4 In “Product_i_Anomaly_j”, it represents the product model and its abnormal type. For example, “screw_p1_concave” represents the p1 model screw with a concave abnormality, and “PCB_023_scratch” represents the 023 model PCB board with a scratch abnormality.

[0074] D. Construct a feature adaptive weighting model

[0075] In this step, a feature adaptive weighting model is constructed to optimize the generation of abnormal images based on the pixel-level differences between the intermediate images generated during the generation of abnormal images and the normal image samples within the abnormal mask area.

[0076] In an exemplary implementation, the feature adaptive weighting model uses an image similarity algorithm to evaluate the pixel differences between the intermediate images generated by the diffusion model and the standard samples within the abnormal mask area, and based on these differences, combined with an adaptive scaling operation and a self-attention mechanism, calculates a weight map. Figure 7 Schematically shows the calculation process of the adaptive weight map, and its inputs include a normal image, an abnormal mask, and a generated intermediate abnormal image.

[0077] During the process of calculating the pixel-level differences within the abnormal region by the similarity algorithm, it involves the generation of intermediate abnormal images based on the diffusion model. Specifically: Based on the pre-trained latent diffusion model, with the abnormal semantic token as the condition c, a new image i similar to the input content is generated in the following way 0 :

[0078]

[0079] where b θ is the approximate backward process that combines the mean and variance of the iterative predicted Gaussian distribution in the denoising diffusion model

[0080] The calculation process of the adaptive weight is as follows

[0081] For the intermediate image generated by the diffusion model at the x-th step calculate the pixel-level difference between it and the normal image i within the mask m region OK Based on this difference, combined with the adaptive scaling operation P AS (·) and the self-attention mechanism SA(·), an adaptive weight map m SW is obtained, and the calculation process is as follows

[0082]

[0083] The adaptive scaling operation P AS (·) is defined as

[0084]

[0085] Based on the above feature adaptive weighting model, it can give different degrees of attention to different regions according to the information of the abnormal mask. For the less obvious abnormal regions generated, the model will strengthen the attention, making the generated abnormal images more realistic and comprehensive as a whole, enhancing the model's ability in overall image abnormal generation, improving the quality of the generated abnormal images, and helping to enhance the detection ability of the abnormal detection algorithm for various abnormal situations

[0086] E. Construct an end-to-end diffusion model

[0087] In this step, an end-to-end diffusion model is constructed based on the latent diffusion model, the anomaly information feature encoding model, and the feature adaptive weighting model. That is, this diffusion model is based on the latent diffusion model (LDM) framework and further integrates the feature encoding mechanism of anomaly information and a model for feature adaptive weighting based on the anomaly mask. Among them, the latent diffusion model serves as the basic framework and undertakes the main image generation and denoising tasks. The anomaly information feature encoding model is responsible for extracting the location, appearance, and semantic features of the anomaly region, forming a feature vector and combining it with the attention mapping module of the latent diffusion model through pixel-level multiplication to achieve the embedding of anomaly information and guide the latent diffusion model to generate images containing anomaly features. The adaptive weight mapping calculated by the feature adaptive weighting model based on the anomaly mask participates in the generation process of the latent diffusion model, adjusts and optimizes the generated intermediate images, so that the finally generated images have anomalies only at the mask, and the other parts are consistent with the original images.

[0088] In an exemplary embodiment, Figure 8 The learning network framework of the entire anomaly diffusion model is shown. Specifically, in the diffusion step from t to t - 1, first, the input image i t is directly fed into the latent diffusion model for feature extraction and denoising. At the same time, the input image i t is input into the feature adaptive weighting model based on the anomaly mask to generate an adaptive weight mapping. On the other hand, the anomaly image, anomaly mask, and anomaly text are used as the input of the anomaly information encoding model, and serialized anomaly tokens are output after passing through the relevant feature processing model. Then, the obtained anomaly tokens are combined with the attention mapping module of the LDM through pixel-level multiplication to achieve the embedding of anomaly information.

[0089] The loss functions adopted by the entire diffusion model are mean squared error (MSE), structural similarity loss (SSIM), and variational lower bound (VLB). The total loss of the network is expressed as follows:

[0090] L = λ 1 L MAE + λ 2 L SSIM + λ 3 L VLB

[0091] Among them, the variational lower bound ensures the consistency between the distribution of generated anomaly samples and the distribution of real samples.

[0092] F. Training the diffusion model

[0093] In this step, based on the dataset obtained in step B, the diffusion model constructed in step E is trained.

[0094] In an exemplary embodiment, the training data set is randomly combined according to a preset augmentation strategy. The input sizes of the images and masks are set to 640*640, the batch size is set to 8, the number of diffusion steps is set to 100,000, the learning rate is set to 0.001, and the optimizer such as Adam is selected for network parameter setting and data processing. For each type of anomaly, 500 anomaly image-mask pairs are generated for subsequent anomaly inspection tasks.

[0095] G. Perform the anomaly image generation task

[0096] In this step, when performing the anomaly image generation task, normal images and randomly generated anomaly region mask images are used as inputs, and the trained diffusion model is utilized to obtain anomaly images and the anomaly region mask images corresponding to the anomaly images.

[0097] In an exemplary embodiment, based on anomaly-free images, mask images are randomly generated, and then the anomaly-free images and the random mask images are input into the trained diffusion model. The anomaly information feature encoding model forms a feature vector by extracting the location, appearance, and semantic features of the anomaly regions in the random mask images and combines it with the attention mapping module of the latent diffusion model through pixel-level multiplication to achieve the embedding of anomaly information and guide the latent diffusion model to generate images containing anomaly features. The adaptive weight mapping calculated by the feature adaptive weighting model based on the anomaly mask participates in the generation process of the latent diffusion model to adjust and optimize the generated intermediate images, such that the finally generated images have anomalies only at the mask, and other parts are consistent with the original images, and finally the generated anomaly image-mask pairs are obtained.

[0098] Embodiment 2

[0099] This embodiment provides an anomaly image generation device based on few-shot diffusion model learning. Refer to Figure 9 , this device is composed of a data processing module, a diffusion model learning module, and a diffusion model inference module.

[0100] The data processing module: is used to collect images of various products and anomaly types by using industrial imaging equipment; batch convert the annotation data of the anomaly regions into binary images through code and locally save them; pre-calculate the coordinates of the augmented anomaly regions, and perform the augmentation operation on the image anomaly mask pairs when the conditions are met.

[0101] The diffusion model learning module: extracts the feature information of the anomaly images and the corresponding masks and serializes them into tokens; calculates the feature adaptive weighting mapping through the anomaly masks to improve the anomaly generation of the network model that only focuses on the mask regions; realizes the learning of anomaly generation by the diffusion model based on the anomaly information and the feature adaptive weighting mapping.

[0102] Diffusion model inference module: capable of generating a random mask image based on the input normal image; inputting the generated mask image and the normal image into the well-trained anomaly generation network to generate an anomaly image and its corresponding mask pair.

[0103] Finally, it should be noted that the above embodiments are only preferred embodiments and are not intended to limit the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the spirit of the present invention and the scope protected by the claims, several modifications, equivalent substitutions, improvements, etc. should be included within the protection scope of the present invention.

Claims

1. A method for generating abnormal images based on few-sample diffusion model learning, characterized in that: include: A. Collect and prepare abnormal images of different types of abnormalities of the product to be detected and abnormal area mask images corresponding to the abnormal images; B. Perform data augmentation operations on abnormal images of different types of abnormalities and abnormal region masks of corresponding abnormal images to obtain data sets; C. Construct an abnormal information feature encoding model for extracting the appearance, location and semantic information of abnormal areas; D. constructing a feature adaptive weighting model for optimizing abnormal image generation based on the pixel-level difference between the intermediate image generated during the abnormal image generation process and the normal image sample in the abnormal mask area; E. Construct an end-to-end diffusion model based on the potential diffusion model, the abnormal information feature encoding model and the feature adaptive weighting model; F. Based on the data set obtained in step B, the diffusion model constructed in step E is trained; G. When performing the abnormal image generation task, the normal image and the randomly generated abnormal region mask map are used as input, and the trained diffusion model is used to obtain the abnormal image and the abnormal region mask map corresponding to the abnormal image.

2. The method for generating abnormal images based on few-sample diffusion model learning according to claim 1, characterized in that: In step A, the method of collecting and preparing abnormal images for different types of abnormalities of the product to be detected and abnormal area mask images corresponding to the abnormal images includes: For products to be detected for abnormalities, image samples are collected through industrial imaging equipment, and the abnormal areas in the abnormal images are marked and then converted into mask images to obtain the abnormal images and the corresponding abnormal area mask images.

3. The method for generating abnormal images based on few-sample diffusion model learning according to claim 1, characterized in that: In step B, the method for performing the data augmentation operation includes: Before augmentation, the coordinate position of the augmented abnormal mask area is first calculated according to the randomly combined augmentation strategy, and it is determined whether it exceeds the boundary of the original image. If it exceeds the boundary, the image augmentation method is randomly combined again until the coordinate conditions of image augmentation are met. If it does not exceed the boundary, the randomly combined augmentation strategy is used to augment the abnormal image and the corresponding abnormal area mask map.

4. The method for generating abnormal images based on few-sample diffusion model learning according to claim 3, characterized in that: The randomly combined augmentation strategies include: random inversion, random center rotation, contrast change, brightness change, center cropping, random translation, distortion transformation, random noise, convolution filtering, or a random combination of one or more image augmentation methods.

5. The method for generating abnormal images based on few-sample diffusion model learning according to claim 1, characterized in that: In step C, the abnormal information feature coding model constructed includes: The abnormal region appearance deep feature extraction network takes the abnormal image and the corresponding abnormal region mask as input and outputs the abnormal region appearance feature vector; The abnormal region position deep feature extraction network takes the abnormal region mask image corresponding to the abnormal image as input and outputs the abnormal region position feature vector; The abnormal region semantic deep feature extraction network takes the product abnormal text as input and outputs the abnormal region semantic feature vector. The product abnormal text is created and saved independently when the dataset is created.

6. The method for generating abnormal images based on few-sample diffusion model learning as claimed in claim 5, characterized in that: The abnormal area appearance deep feature extraction network and the abnormal area position deep feature extraction network both include multiple feature extraction units in a downsampling part and an upsampling part, wherein the feature extraction unit in the downsampling part is composed of a maximum pooling layer, a convolution layer, a batch normalization layer and a Relu activation layer stacked in sequence; the feature extraction unit in the upsampling part is composed of a deconvolution layer, a convolution layer, a batch normalization layer and a Relu activation layer stacked in sequence.

7. The method for generating abnormal images based on few-sample diffusion model learning according to claim 5, characterized in that: The abnormal area semantic deep feature extraction network uses word embedding technology to map words in the product abnormality text into low-dimensional dense vectors, and then extracts features from the text through context embedding, sentence level embedding or using a deep learning model to convert the mapped vector into an abnormal area semantic feature vector.

8. The method for generating abnormal images based on few-sample diffusion model learning according to claim 1, characterized in that: In step D, the feature adaptive weighted model uses an image similarity algorithm to calculate the pixel-level difference between the intermediate image generated in the abnormal image generation process and the normal image sample in the abnormal mask area, and calculates the weight mapping based on the difference according to the adaptive scaling operation and self-attention mechanism to obtain an adaptive weight mapping to adjust and optimize the generated intermediate image.

9. The method for generating abnormal images based on few-sample diffusion model learning according to claim 1, characterized in that: In step E, when generating an abnormal image, the constructed diffusion model takes a normal image sample and a random mask image as input, extracts the appearance feature vector, position feature vector and semantic feature vector of the abnormal area of ​​the random mask image through the abnormal information feature encoding model, and combines them into the attention mapping module of the potential diffusion model by pixel-level multiplication to embed the abnormal information, guide the potential diffusion model to generate an intermediate image containing abnormal features, and calculates the weight mapping based on the pixel-level difference between the intermediate image and the normal image sample in the abnormal mask area through the feature adaptive weighting model, adjusts and optimizes the generated intermediate image, and finally generates an abnormal image with an abnormality only in the mask area and other parts consistent with the normal sample image.

10. An abnormal image generation device based on few-sample diffusion model learning, characterized in that: include: A data processing module is used to collect and generate abnormal images of different types of abnormalities of the product to be detected and abnormal area mask images corresponding to the abnormal images, and perform data augmentation operations on the abnormal images of different types of abnormalities and the abnormal area mask images corresponding to the abnormal images; The diffusion model learning module is used to extract features from abnormal images and corresponding masks and serialize them into abnormal information feature vectors, calculate feature adaptive weight mapping through abnormal masks, and realize the learning of abnormal generation by the diffusion model based on the abnormal information feature vectors and feature adaptive weight mapping; The diffusion model inference module is used to generate a random mask image based on the input normal image samples, and input the generated mask image and normal image samples into the trained diffusion model to generate an abnormal image and a corresponding mask image.

Citation Information

Cited By

  • Power transmission line small sample abnormal data generation method and system based on diffusion model and feature migration

    CN120673192A