Self-attention guidance anomaly generation method based on latent variable diffusion model
Through the self-attention guidance method based on the latent variable diffusion model, high-quality normal-exception sample pairs are generated, which solves the problems of insufficient authenticity of abnormal image generation and inability to generate normal-exception sample pairs in the prior art, and achieves the improvement of data generation in the abnormal detection technology.
Patent Information
- Application Number
- CN202510112323.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-23
AI Technical Summary
In the existing abnormality detection technology, the abnormal image generation method is difficult to ensure the authenticity of the generated results, and it is impossible to generate normal-analysis sample pairs at the same time, which cannot meet the requirements of the reconstruction-based abnormality detection method.
The self-attention guidance method based on the latent variable diffusion model is adopted to generate the latent variable diffusion model through the text-guided anomaly-normal sample image pair. This method uses a combination of diffusion process and inverse process to control image generation using a self-attention mechanism to ensure the authenticity and diversity of abnormal areas.
The generated abnormal images perform excellently in terms of authenticity and diversity of abnormal areas, and the abnormal areas and normal areas are naturally fusion, and the normal areas are highly consistent with the corresponding normal images, which solves the problem of scarcity of data and realizes high-quality abnormal data generation.
Smart Images

Figure CN120031730A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of anomaly detection, and in particular to a self-attention guided anomaly generation method based on a latent variable diffusion model. Background Art
[0002] Industrial anomaly detection, that is, the detection, location and classification of anomalies, is of great significance in industrial manufacturing. However, in actual industrial production, abnormal samples are extremely scarce. Therefore, the current mainstream anomaly detection methods are mainly unsupervised methods, which usually only use normal samples to train the model, or use semi-supervised methods to combine normal samples with a small number of abnormal samples for training. Although these methods have shown certain advantages in anomaly detection, their performance in anomaly location tasks is still limited. To this end, researchers have proposed an anomaly generation method to synthesize more abnormal data, so as to improve the overall performance through supervised learning anomaly detection methods.
[0003] There are two main problems with existing abnormal image synthesis methods. First, when generating abnormal images, current methods usually have difficulty in ensuring the authenticity of the generated results. For example, non-model methods usually generate abnormal images by randomly pasting fragments of existing abnormal samples or abnormal textures onto normal samples. The abnormal data generated in this way often lacks authenticity. Although the model based on generative adversarial networks (GANs) can generate more realistic abnormal data, its performance is highly dependent on a large amount of training data. However, due to the scarcity of abnormal samples in the field of anomaly detection, the GAN model cannot fully realize its potential, resulting in limited authenticity of the generated data. In addition, although the existing abnormal image generation methods based on diffusion models focus on the realism of abnormal areas, they often ignore the natural fusion of abnormal areas with normal areas, resulting in poor overall quality of the generated images.
[0004] Secondly, existing methods mainly focus on the generation of abnormal images and their corresponding abnormal masks, ignoring the requirement of some anomaly detection methods for the normal images corresponding to the abnormal images. Especially in the reconstruction-based anomaly detection methods, the abnormal-normal sample pairs are crucial for model training. Therefore, focusing only on the generation of abnormal images and abnormal masks cannot meet the needs of all anomaly detection technologies and has certain limitations. Summary of the invention
[0005] The purpose of the present invention is to propose a self-attention guided normal-abnormal sample pair generation method based on a latent variable diffusion model to overcome the problems of limited authenticity of abnormal images generated in the prior art and inability to simultaneously generate normal-abnormal sample pairs.
[0006] To achieve the above object, the present invention proposes the following technical solutions:
[0007] A self-attention guided anomaly generation method based on a latent variable diffusion model comprises the following steps:
[0008] S1. Construct a text-guided latent variable diffusion model for abnormal-normal sample images and design the diffusion process of the latent variable diffusion model.
[0009] S2, taking random noise and corresponding control text as input, and outputting the predicted noise through the latent variable diffusion model of abnormal-normal sample images guided by the text;
[0010] S3, training the latent variable diffusion model of abnormal-normal sample images based on the diffusion process constructed in S1;
[0011] S4, based on the inverse process of the diffusion model and the influence of the attention mechanism on the image generation results, using the predicted noise, iteratively reconstructs the abnormal-normal sample images at the same time;
[0012] S5. Calculate the pixel-level difference between the normal sample image and the abnormal sample image generated by the model, and convert the difference into a binary mask as the abnormal mask of the generated image, and finally generate a normal image-abnormal image-abnormal mask image pair.
[0013] Preferably, the S1 specifically includes the following contents:
[0014] An initial sample x of normal and abnormal images is passed through a pre-trained encoder 0 Compression is performed to facilitate diffusion operations in the latent representation space:
[0015] Z 0 =Encode(x 0 )
[0016] Among them, x 0 represents the initial sample (normal sample image or abnormal sample image), Encode(·) represents the pre-trained encoder, and Z 0 Represents the latent space variable after the initial sample is compressed by the encoder.
[0017] Through the diffusion process, the variable Z in the latent space 0 Gradually add Gaussian noise, the standard deviation of which is a fixed value β t , the mean is determined by the fixed value β t and the data z at the current time t t Joint decision. This process is a Markov chain process, and its transition probability is:
[0018]
[0019] Among them, z t and z t-1denote the latent variables at step t and step t-1 respectively; q(z t |z t-1 ) indicates that from z t-1 to z t The forward diffusion process; β t is a fixed value, and 0<β t <1;
[0020] According to parameter renormalization, the latent variable z at any time t Can be based on Z 0 and β t Calculate it without iteration:
[0021]
[0022] in, δ is random noise sampled from a standard normal distribution with mean 0 and variance 1;
[0023] In the whole diffusion process, the variable Z in the latent space 0 As input, after T consecutive steps of forward noise addition, the final output is the latent variable z that follows the standard Gaussian distribution t .
[0024] Preferably, the text-guided abnormal-normal sample image pair latent variable diffusion model includes the following three parts:
[0025] ① Pre-trained variational autoencoder (VAE): A generative model based on the Encoder-Decoder architecture, where the encoder compresses the input image x into a low-dimensional latent space feature z, and the decoder can restore the low-dimensional latent space feature z to a pixel-level image x.
[0026] ② Pre-trained CLIP Text Encoder: As a text encoder τ θ (·), using prompts that can reflect the difference between normal sample images and abnormal sample images to guide the generated content of the diffusion model. And the LORA weight is introduced in the text encoder. For each attention layer, consider the query matrix W q , the bond matrix W k , value matrix W v And the output projection matrix W o ; where the linear projection of each matrix is replaced by the following formula:
[0027] h 1,* =W * h l-1 +B * A * h l-1
[0028] Where h represents the projected input / output activations, resulting in the trainable LoRA weights of each attention layer l being
[0029] Among them, the tips that can reflect the difference between normal sample images and abnormal sample images are specifically:
[0030] P src :vfx with nothing
[0031] P dst :vfx with sks
[0032] Among them, the prompt P src and P dst Corresponding to normal sample image I and abnormal sample image I respectively a , both prompts are represented by the text encoder τ θ (·) encoding and then injected into the Unet of the diffusion model. vfx and sks are used as identifiers for class names and anomaly names. These words have weak priors in both the language model and the diffusion model, are easier to fit than other words, and can achieve better generation results.
[0033] ③ Pre-trained U-Net network for noise prediction: First, the input features are downsampled, then the original feature size is restored by upsampling, and the features extracted by the text encoder are fused at each layer. Similarly, we introduce LORA weights in the text encoder. For each attention layer, consider the query matrix W q , the bond matrix W k , value matrix W v And the output projection matrix W o . The linear projection of each matrix is replaced by the following formula:
[0034] h 1,* =W * h l-1 +B * A * h l-1
[0035] Where h represents the projected input / output activations, resulting in the trainable LoRA weights of each attention layer l being
[0036] Preferably, S3 specifically includes the following contents:
[0037] Keep the original parameters of the pre-trained model unchanged and only use the objective function to train the LoRA weights. Weight matrix B * Initialized to 0, A *Initialize randomly. When testing, you can update the weights by Integrate the LoRA weights into the model:
[0038] ||ε-ε θ (z t ,τ θ (P),t)||ε~N(0,1)
[0039] Where ||·|| represents the L1 norm, t is any random value in the set {1,2,…,T}, and N(0,1) is a normal distribution with a mean of 0 and a variance of 1. By back-propagating the gradient, we jointly optimize the text encoder τ θ (·) and the noise prediction network ε θ (·)’s LoRA weight;
[0040] The parameters of the latent variable diffusion model for abnormal-normal sample images guided by the text are optimized, and the deep learning network is updated in an iterative manner until the loss function converges, thereby completing the training process of the latent variable diffusion model for abnormal-normal sample images guided by the text.
[0041] Preferably, the S4 specifically includes the following contents:
[0042] ① For normal sample images, from the Gaussian noise image z in the latent space T The normal sample image is restored through the inverse diffusion process, which is also a Markov chain process:
[0043]
[0044] Among them, p θ (z t-1 |z t ,P src ) indicates a learnable conditional Gaussian distribution; μ θ (P src ,z t ,γ t ) represents the mean of the learnable distribution; represents the distribution variance, which is set to the default value β;
[0045] Assume that z is given t and z 0 , z t-1 The posterior distribution of is a conditional Gaussian distribution:
[0046]
[0047] z can be derived from the Bayesian formula t-1 The mean of the posterior distribution:
[0048]
[0049] Fitting the posterior distribution p based on a deep learning network θ , the noise ε is estimated through the text-guided abnormal-normal sample image latent variable diffusion model, and the approximate latent space representation of the normal sample image is obtained after transformation
[0050]
[0051] Then we can get the learnable distribution p θ The mean of is:
[0052]
[0053] The text-guided abnormal-normal sample image pair latent variable diffusion model finally passes step-by-step reasoning:
[0054]
[0055] Among them, ε~N(0,1), the whole reverse diffusion process is continuous after T steps, and the final potential space representation z of the normal sample image is obtained 0 ;
[0056] Finally, the pre-trained decoder network decodes the latent space representation of the normal sample into a normal sample image x 0 :
[0057] x 0 =decode(z 0 )
[0058] Where decode(·) represents the pre-trained decoder network;
[0059] ②For abnormal samples corresponding to normal samples, we repeat the reasoning process of normal samples, and the text guidance condition is P dst , and replace the self-attention map of the abnormal sample image with the self-attention map of the normal sample image generation process:
[0060]
[0061] Among them, the self-attention map refers to the self-attention layer in the noise prediction network through linear projection l k and l q , from the noise image φ self (z t ) receives the key matrix K self and the query matrix Q self ;
[0062] The formula of the self-attention map is defined as follows:
[0063] Qself = l q (φ self (z t ))
[0064] K self = l k (φ self (z t ))
[0065]
[0066] Among them, d self K self and Q self Dimension; M self Represents the self-attention map, which determines the correlation weight between the i-th and j-th spatial features in the image, thereby affecting the spatial layout and shape details of the generated image. In the noise prediction of abnormal sample images, we replace the self-attention map in the abnormal sample image prediction process with the attention map in the normal sample prediction process at the corresponding number of steps, that is:
[0067]
[0068] Preferably, the S5 specifically includes the following contents:
[0069] Using the same ImageNet pre-trained feature extractor F, from the input normal image x 0 and abnormal images Extract features and calculate anomalies on feature maps of different scales by cosine similarity
[0070]
[0071] Where n represents the nth feature layer f n , the total score of abnormal location of each layer of the input pair is calculated as:
[0072]
[0073] Among them, σ n′ represents the upsampling factor used to maintain the same dimension of the image in pixel space; n' represents the number of feature layers used during inference.
[0074] Compared with the prior art, the present invention provides a method for generating normal-abnormal sample pairs guided by self-attention based on a latent variable diffusion model, which has the following beneficial effects:
[0075] The present invention proposes a method for simultaneously generating normal images and abnormal images by controlling a latent variable diffusion model using a self-attention map, which can effectively ensure the quality of the generated abnormal images in the following aspects: the authenticity and diversity of abnormal regions, the naturalness of the fusion of abnormal regions with normal regions, and the high consistency between the normal regions of the abnormal images and the normal regions in the corresponding normal images, thereby generating realistic and diverse normal image-abnormal image data pairs.
[0076] In addition, the present invention combines LoRA technology with the small sample training framework of the diffusion model, which can generate high-quality real images when data is limited, while still maintaining excellent generation performance, fully solving the problem of data scarcity. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings involved in the embodiments are briefly introduced. Obviously, the drawings in the following description are only schematic illustrations of some embodiments of the present invention. For those skilled in the art, other forms of drawings can also be constructed based on these drawings without creative work.
[0078] Figure 1 It is a schematic diagram of a latent variable diffusion model of abnormal-normal sample images guided by text mentioned in Example 1 of the present invention;
[0079] Figure 2 Schematic diagram of the self-attention map generation process mentioned in Example 1 of the present invention;
[0080] Figure 3 This is a schematic diagram of generating the exception mask mentioned in Embodiment 1 of the present invention;
[0081] Figure 4 The test results of different methods mentioned in Example 2 of the present invention are compared;
[0082] Figure 5 This is the effect of the normal image-abnormal image and abnormal mask generated by the present invention mentioned in Embodiment 2 of the present invention. DETAILED DESCRIPTION
[0083] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.
[0084] Embodiment 1:
[0085] See also Figure 1The present invention proposes a self-attention-guided anomaly generation method based on a latent variable diffusion model. The diffusion model is trained using a small number of abnormal sample images and normal sample images in an existing data set, and the influence of the self-attention mechanism on the generated content of the diffusion model is used to achieve the simultaneous generation of normal sample image-abnormal sample image pairs. Finally, the anomaly mask is obtained based on the difference between the generated normal sample image and the abnormal sample image.
[0086] In this embodiment, the public dataset MVTEC is selected as a reference for image synthesis in industrial scenarios. This dataset covers 15 common industrial objects, each of which contains a large number of normal images without defects, a small number of abnormal images with defects and their corresponding precise abnormal masks, and provides label information of the corresponding object category and defect type for each image. The following steps are included:
[0087] S1. Constructing the diffusion process of the latent variable diffusion model of abnormal sample images and normal sample images guided by text;
[0088] An initial sample x of normal and abnormal images is passed through a pre-trained encoder 0 Compression is performed to facilitate diffusion operations in the latent representation space:
[0089] Z 0 =Encode(x 0 )
[0090] Among them, x 0 represents the initial sample, Encode(·) represents the pre-trained encoder, and Z 0 It represents the latent space variable after the initial sample is compressed by the encoder. The initial sample includes the same number of normal sample images and a certain type of abnormal sample images. In this embodiment, six normal sample images and six abnormal sample images are taken.
[0091] Through the diffusion process, the variable Z in the latent space 0 Gradually add Gaussian noise, the standard deviation of which is a fixed value β t , the mean is determined by the fixed value β t and the data z at the current time t t Joint decision. This process is a Markov chain process, and its transition probability is:
[0092]
[0093] Among them, z t and z t-1 denote the latent variables at step t and step t-1 respectively; q(z t |z t-1 ) indicates that from z t-1 to z tThe forward diffusion process; β t is a fixed value, and 0<β t <1;
[0094] According to parameter renormalization, the latent variable z at any time t Can be based on Z 0 and β t Calculate it without iteration:
[0095]
[0096] in, δ is random noise sampled from a standard normal distribution with mean 0 and variance 1;
[0097] In the whole diffusion process, the variable Z in the latent space 0 As input, after T consecutive steps of forward noise addition, the final output is the latent variable z that follows the standard Gaussian distribution t ;
[0098] S2. Establish a text-guided abnormal-normal sample image pair latent variable diffusion model, take random noise and corresponding control text as input, and output the predicted noise through the text-guided abnormal-normal sample image pair latent variable diffusion model;
[0099] The text-guided latent variable diffusion model for abnormal-normal sample image pairs consists of the following three parts:
[0100] ① Pre-trained variational autoencoder (VAE): A generative model based on the Encoder-Decoder architecture. This embodiment uses the pre-trained variational autoencoder model of Stabel Diffusion-2.1. The encoder compresses the input image x to a low-dimensional latent space feature z, and the decoder can restore the low-dimensional latent space feature z to a pixel-level image x.
[0101] ② Pre-trained CLIP Text Encoder: As a text encoder τ θ (·), using prompts that can reflect the difference between normal sample images and abnormal sample images to guide the generation of content by the diffusion model. This embodiment uses the pre-trained text encoder model of stabeldiffusion-2.1. We introduce LORA weights in the text encoder. For each attention layer, we consider the query matrix W q , the bond matrix W k , value matrix W v And the output projection matrix W o . The linear projection of each matrix is replaced by the following formula:
[0102] h 1,* =W * h l-1 +B * A * h l-1
[0103] In the formula, h represents the input / output activation of the projection, and the trainable LoRA weight of each attention layer l is
[0104] Among them, the tips that can reflect the difference between normal sample images and abnormal sample images are specifically:
[0105] P src :vfx with nothing
[0106] P dst :vfx with sks
[0107] Among them, the prompt P src and P dst Corresponding to normal sample image I and abnormal sample image I respectively a In this embodiment, LoRA weights are trained separately for each type of abnormality. During the training process, normal sample image I and abnormal sample I a Six images each. Both prompts are generated by the text encoder τ θ (·) encoding and then injected into the Unet of the diffusion model. vfx and sks are used as identifiers for class names and anomaly names. These words have weak priors in both the language model and the diffusion model, are easier to fit than other words, and can achieve better generation results.
[0108] ③ Pre-trained U-Net network for noise prediction: First, downsample the input features, then restore the original feature size by upsampling, and fuse the features extracted by the text encoder at each layer. This embodiment uses the pre-trained U-Net model of stabeldiffusion-2.1. Similarly, we introduce LORA weights in the text encoder. For each attention layer, we consider the query matrix W q , the bond matrix W k , value matrix W v And the output projection matrix W o . The linear projection of each matrix is replaced by the following formula:
[0109] h 1,* =W * h l-1 +B * A * h l-1
[0110] In the formula, h represents the input / output activation of the projection, and the trainable LoRA weight of each attention layer l is
[0111] S3, training the latent variable diffusion model of abnormal-normal sample image pairs based on the diffusion process constructed in S1;
[0112] Keep the original parameters of the pre-trained model unchanged and only use the objective function to train the LoRA weights. Weight matrix B * Initialized to 0, A * Initialize randomly. When testing, you can update the weights by Integrate the LoRA weights into the model:
[0113] ||ε-ε θ (z t ,τ θ (P),t)||ε~N(0,1)
[0114] Where ||·|| represents the L1 norm, t is any random value in the set {1,2,…,T}, and N(0,1) is a normal distribution with a mean of 0 and a variance of 1. By back-propagating the gradient, we jointly optimize the text encoder τ θ (·) and the noise prediction network ε θ (·)’s LoRA weight;
[0115] The parameters of the latent variable diffusion model for abnormal-normal sample images guided by the text are optimized, and the deep learning network is updated in an iterative manner until the loss function converges, thereby completing the training process of the latent variable diffusion model for abnormal-normal sample images guided by the text.
[0116] S4. Based on the inverse process of the diffusion model and the influence of the attention mechanism on the image generation results, the predicted noise is used to iteratively reconstruct the abnormal-normal sample images.
[0117] Referring to the inverse process based on the diffusion model and the influence of the attention mechanism on the image generation results as shown below, the pseudo code of the process of iteratively reconstructing abnormal-normal sample images by using predicted noise is shown.
[0118]
[0119] The inference process of the latent variable diffusion model for abnormal-normal sample images guided by text can be defined as a reverse Markov process, which is the inverse process of the forward diffusion process. T and z* T Start, z T and z* T Obey Gaussian distribution:
[0120] zT =z* T ~N(0,1)
[0121] ① For normal sample images, from the Gaussian noise image z in the latent space T The normal sample image is restored through the inverse diffusion process, which is also a Markov chain process:
[0122]
[0123] Among them, p θ (z t-1 |z t ,P src ) indicates a learnable conditional Gaussian distribution; μ θ (P src ,z t ,γ t ) represents the mean of the learnable distribution; represents the distribution variance, which is set to the default value β;
[0124] Assume that z is given t and z 0 , z t-1 The posterior distribution of is a conditional Gaussian distribution:
[0125]
[0126] z can be derived from the Bayesian formula t-1 The mean of the posterior distribution:
[0127]
[0128] Fitting the posterior distribution p based on a deep learning network θ , the noise ε is estimated through the text-guided abnormal-normal sample image latent variable diffusion model, and the approximate latent space representation of the normal sample image is obtained after transformation
[0129]
[0130] Then we can get the learnable distribution p θ The mean of is:
[0131]
[0132] The text-guided abnormal-normal sample image pair latent variable diffusion model finally passes step-by-step reasoning:
[0133]
[0134] Among them, ε~N(0,1), the whole reverse diffusion process is continuous after T steps, and the final potential space representation z of the normal sample image is obtained 0 ;
[0135] Finally, the pre-trained decoder network decodes the latent space representation of the normal sample into a normal sample image x 0 :
[0136] x 0 =decode(z 0 )
[0137] Here, decode(·) represents the pre-trained decoder network.
[0138] ②For abnormal samples corresponding to normal samples, we repeat the reasoning process of normal samples, and the text guidance condition is P dst , and replace the self-attention map of the abnormal sample image with the self-attention map of the normal sample image generation process:
[0139]
[0140] Among them, refer to Figure 2 The calculation process of the self-attention map is shown in Figure 2. The self-attention map refers to the self-attention layer in the noise prediction network through linear projection l k and l q , from the noise image φ self (z t ) receives the key matrix K self and the query matrix Q self .
[0141] The formula of the self-attention map is defined as follows:
[0142] Q self = l q (φ self (z t ))
[0143] K self = l k (φ self (z t ))
[0144]
[0145] Among them, d self K self and Q self Dimensions. M selfRepresents the self-attention map, which determines the correlation weight between the i-th and j-th spatial features in the image, thereby affecting the spatial layout and shape details of the generated image. In the noise prediction of abnormal sample images, we replace the self-attention map in the abnormal sample image prediction process with the attention map in the normal sample prediction process at the corresponding number of steps, that is:
[0146]
[0147] S5. Calculate the pixel-level difference between the normal sample image and the abnormal sample image generated by the model, and convert the difference into a binary mask as the abnormal mask of the generated image, and finally generate a normal image-abnormal image-abnormal mask image pair.
[0148] See attached Figure 3 The abnormal mask generation process shown in Figure 2 is as follows. Using the same ImageNet pre-trained feature extractor F, we start with the input normal image x 0 and abnormal images Extract features and calculate anomalies on feature maps of different scales by cosine similarity
[0149]
[0150] Where n represents the nth feature layer f n , the total score of abnormal location of each layer of the input pair is calculated as:
[0151]
[0152] Among them, σ n′ represents the upsampling factor used to maintain the same dimension of the image in pixel space; n' represents the number of feature layers used during inference.
[0153] Embodiment 2:
[0154] Based on Example 1, but different in that the advanced contrast methods used include Crop-paste, SD-GAN, Defect-GAN, DFMGAN and AnomalyDiffusion. Among them, Crop-paste is a non-model method that generates abnormal data by randomly pasting fragments of existing abnormal samples or abnormal texture datasets onto normal samples; SD-GAN, Defect-GAN and DFMGAN use generative adversarial networks (GAN) to synthesize abnormal images; AnomalyDiffusion is a texture inversion technology based on a diffusion model, which learns the appearance and location information of the anomaly, and generates abnormal images on masked normal samples. We tested the above methods on the same dataset and plotted the results as follows: Figure 4 .from Figure 4 It can be seen that the self-attention guided normal-abnormal sample pair generation method based on the latent variable diffusion model proposed in the present invention can generate higher quality abnormal images and has higher authenticity in the fusion of abnormal areas and normal areas.
[0155] Table 1 shows the experimental results of the generation quality of each comparison method. We use Inception Score (IS) and Intra-cluster pairwise LPIPS distance (IC-LPIPS) to measure the generation quality and generation diversity. The experiment shows that the proposed method achieves the highest level in both the generation quality and diversity of abnormal data.
[0156] Table 1
[0157]
[0158] Table 2. Comparative experiments on pixel-level and image-level anomaly localization on the MVTec dataset were performed by training U-Net on DRAEM, PRN, DFMGAN, AnomalyDiffusion, and data generated by the present invention. We calculated the AUROC, AP, and F1-max indicators at the pixel and image levels, and the results showed that the U-Net trained on the data generated by the present invention achieved a high level of comprehensive performance.
[0159] Table 2
[0160]
[0161] Table 3 shows the application effect of the normal-abnormal data pairs generated by the present invention in abnormality detection and location. Since many existing synthetic abnormality methods are difficult to generate complete normal-abnormal data pairs, we trained the condition-diffusion model on the data generated by Crop-paste and the present invention respectively, and conducted pixel-level and image-level abnormality location experiments on the MVTec dataset. The results show that the data generated by the present invention has a significant advantage in training the diffusion model (generating corresponding normal samples with abnormal samples as conditions).
[0162] Table 3
[0163]
[0164]
[0165] Figure 5The effects of the normal image-abnormal image and abnormal mask generated by the present invention are demonstrated. The results show that the generation method of the present invention can not only ensure the authenticity of the fusion of abnormal areas and normal areas, but also ensure the high consistency between the normal areas in the abnormal images and the corresponding normal images.
[0166] It should be noted that the above description is only a preferred embodiment of the present invention and does not constitute a limitation on the protection scope of the present invention. Any equivalent replacement or change made by any technician familiar with the technical field within the technical scope disclosed in the present invention according to the technical solution and inventive concept of the present invention shall be deemed to be included in the protection scope of the present invention.
Claims
1. A self-attention guided anomaly generation method based on latent variable diffusion model, characterized in that: The following steps are involved: S1. Construct a text-guided latent variable diffusion model for abnormal-normal sample images and design the diffusion process of the latent variable diffusion model. S2, taking random noise and corresponding control text as input, and outputting the predicted noise through the latent variable diffusion model of abnormal-normal sample images guided by the text; S3, training the latent variable diffusion model of abnormal-normal sample images based on the diffusion process constructed in S1; S4, based on the inverse process of the diffusion model and the influence of the attention mechanism on the image generation results, using the predicted noise, iteratively reconstructs the abnormal-normal sample images at the same time; S5. Calculate the pixel-level difference between the normal sample image and the abnormal sample image generated by the model, and convert the difference into a binary mask as the abnormal mask of the generated image, and finally generate a normal image-abnormal image-abnormal mask image pair.
2. According to claim 1, a method for generating normal-abnormal sample pairs guided by self-attention based on latent variable diffusion model is characterized in that: The S1 specifically includes the following contents: The initial samples x0 of normal and abnormal images are compressed through a pre-trained encoder for diffusion operation in the latent representation space: Z0=Encode(x0) Where x0 represents the initial sample, including the same number of normal sample images and a certain type of abnormal sample images; Encode(·) represents the pre-trained encoder, and Z0 represents the latent space variable after the initial sample is compressed by the encoder; In the forward diffusion process, given the initial data distribution Z0, Gaussian noise is continuously added to the distribution, and the standard deviation of the noise is a fixed value β t , the mean is determined by the fixed value β t and the data z at the current time t t Jointly determined; the above process is a Markov chain process, and its transition probability is: Among them, z t and z t-1 denote the latent variables at step t and step t-1 respectively; q(z t |z t-1 ) indicates that from z t-1 to z t The forward diffusion process; β t is a fixed value, and 0<β t <1; According to parameter renormalization, the latent variable z at any time t Both based on Z0 and β t Calculate it without iteration: in, δ is random noise sampled from a standard normal distribution with mean 0 and variance 1; In the whole diffusion process, the variable Z0 in the latent space is used as input, and after T consecutive steps of forward noise addition, the final output is the latent variable z that obeys the standard Gaussian distribution. t .
3. According to claim 2, a method for generating normal-abnormal sample pairs guided by self-attention based on latent variable diffusion model is characterized in that: The text-guided abnormal-normal sample image pair latent variable diffusion model includes the following three parts: ① Pre-trained variational autoencoder: a generative model based on the Encoder-Decoder architecture; The encoder compresses the input image x into a low-dimensional latent space feature z, and the decoder can restore the low-dimensional latent space feature z to a pixel-level image x; ② Pre-trained CLIP Text Encoder: As a text encoder τ θ (·) Use prompts that can reflect the difference between normal sample images and abnormal sample images to guide the generated content of the diffusion model; introduce LORA weights in the text encoder. For each attention layer, consider the query matrix W q , the bond matrix W k , value matrix W v And the output projection matrix W o ; Replace the linear projection of each matrix with the following formula: h1=W * h l-1 +B * AND * h l-1 Where h represents the input / output activation of the projection, and the trainable LoRA weights of each attention layer l are The tips reflecting the difference between normal sample images and abnormal sample images are as follows: P src :vfx with nothing P dst :vfx with sks Among them, the prompt P src and P dst Corresponding to normal sample image I and abnormal sample image I respectively a , both prompts are generated by the text encoder τ θ (·) Encoded and then injected into the Unet of the diffusion model; vfx and sks are used as identifiers of class name and anomaly name; ③ Pre-trained U-Net network: used for noise prediction, first downsample the input features, then restore the original feature size through upsampling, and fuse the features extracted by the text encoder at each layer; introduce LORA weights in the text encoder, the specific content is the same as in ②.
4. The method for generating normal-abnormal sample pairs based on a latent variable diffusion model guided by self-attention according to claim 3, characterized in that: The S3 specifically includes the following contents: Keep the original parameters of the pre-trained model unchanged and only use the objective function to train the LoRA weights; * Initialized to 0, A * Initialize randomly and update the weights at test time Integrate the LoRA weights into the model: ||e-e θ (z t ,t θ (P),t)||ε~N(0,1) Where ||·|| represents the L1 norm, t is any random value in the set {1,2,…,T}, and N(0,1) is a normal distribution with a mean of 0 and a variance of 1. By back-propagating the gradient, we jointly optimize the text encoder τ θ (·) and the noise prediction network ε θ (·)’s LoRA weight; The parameters of the latent variable diffusion model for abnormal-normal sample images guided by text are optimized, and the deep learning network is updated in an iterative manner until the loss function converges, thereby completing the training process of the latent variable diffusion model for abnormal-normal sample images guided by text.
5. According to claim 4, a method for generating normal-abnormal sample pairs guided by self-attention based on latent variable diffusion model is characterized in that: The S4 specifically includes the following contents: ① For normal sample images, from the Gaussian noise image z in the latent space T The normal sample image is restored through the inverse diffusion process, which is also a Markov chain process: Among them, p θ (z t-1 |z t ,P src ) indicates a learnable conditional Gaussian distribution; μ θ (P src ,z t ,γ t ) represents the mean of the learnable distribution; represents the distribution variance, which is set to the default value β; Assume that z is given t and z0,z t-1 The posterior distribution of is a conditional Gaussian distribution: Derived from the Bayesian formula, z t-1 The mean of the posterior distribution: Fitting the posterior distribution p based on a deep learning network θ , the noise ε is estimated through the text-guided abnormal-normal sample image latent variable diffusion model, and the approximate latent space representation of the normal sample image is obtained after transformation Then we can get the learnable distribution p θ The mean of is: The text-guided abnormal-normal sample image pair latent variable diffusion model finally passes step-by-step reasoning: Among them, ε~N(0,1), the whole reverse diffusion process goes through T consecutive steps to obtain the final latent space representation z0 of the normal sample image; Finally, the latent space representation of the normal sample is decoded into a normal sample image x0 by the pre-trained decoder network: x0=decode(z0) Where decode(·) represents the pre-trained decoder network; ② For abnormal samples corresponding to normal samples, repeat the reasoning process of normal samples, and the text guidance condition is P dst , and replace the self-attention map of the abnormal sample image with the self-attention map of the normal sample image generation process: Among them, the self-attention map refers to the self-attention layer in the noise prediction network through linear projection l k and l q , from the noise image φ self (z t ) receives the key matrix K self and the query matrix Q self ; The formula of the self-attention map is defined as follows: Q self =l q (φ self (z t )) K self =l k (φ self (With t )) Among them, d self K self and Q self Dimension; M self Represents the self-attention map, which determines the correlation weight between the i-th and j-th spatial features in the image, thereby affecting the spatial layout and shape details of the generated image; in the noise prediction of abnormal sample images, the self-attention map in the abnormal sample image prediction process is replaced by the attention map in the normal sample prediction process at the corresponding number of steps, that is:
6. The method for generating normal-abnormal sample pairs guided by self-attention based on latent variable diffusion model according to claim 5, characterized in that: The S5 specifically includes the following contents: Using the same ImageNet pre-trained feature extractor F, we input the normal image x0 and the abnormal image Extract features and calculate anomalies on feature maps of different scales by cosine similarity Where n represents the nth feature layer f n , the total score of abnormal location of each layer of the input pair is calculated as: Among them, σ n′ represents the upsampling factor used to maintain the same dimension of the image in pixel space; n' represents the number of feature layers used during inference.
Citation Information
Cited By
Image reconstruction method and device based on potential diffusion model, and medium
CN120510039A
Condition-controllable image sample expansion method and system for surface defect detection
CN120525000A