Abnormity detection method based on continuous diffusion model
By projecting the update gradient of the new task in anomaly detection to a subspace orthogonal to the important representation of the previous task, and using iterative singular value decomposition and anomaly mask network, the problem of forgetting and overfitting the diffusion model during the continuous learning process is solved, improving the reconstruction ability of the anomaly region and the robustness of the model.
Patent Information
- Application Number
- CN202510388996.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-05-30
AI Technical Summary
Existing anomaly detection methods based on diffusion models are susceptible to ‘catastrophic forgetting’ and ‘faithful hallucinations’ during continuous learning, and are prone to ‘overfitting’ normal samples, resulting in a decrease in the ability to reconstruct anomaly regions.
An anomaly detection method based on a continuous diffusion model is proposed. By projecting the update gradient of a new task into a subspace orthogonal to the important representation of the previous task, the impact of the update gradient on the previous feature space is eliminated, and the iterative singular value decomposition (iSVD) method is used to reduce memory consumption. At the same time, an exception mask network is designed to use the non-local encoder and exception mask loss of the Transformer structure to mask exception features and enhance the conditional mechanism of image-to-image generation.
It effectively alleviates the forgetting problem of the diffusion model in continuous anomaly detection, enhances the model's ability to reconstruct the abnormal areas, reduces memory consumption, and improves the robustness of the model in dynamically changing industrial scenarios.
Smart Images

Figure CN120071018A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of continuous industrial anomaly detection, and in particular to an anomaly detection method based on a continuous diffusion model. Background Art
[0002] Anomaly Detection (AD) has been widely applied in both medical and industrial fields. Traditional methods usually adopt a one-to-one paradigm, training a customized model for a specific category or an industrial product. The one-to-one mode severely limits the generalization ability of the model. On the contrary, the one-to-many paradigm proposes that training a single model can perform anomaly detection across all categories. However, in real-world scenarios, the pattern increment is unpredictable, and the model needs to have the ability of continuous learning, which is called the continuous anomaly detection paradigm.
[0003] Diffusion models (DMs) are widely used in anomaly detection. Current DM-based methods generate normal images conditioned on anomaly images through an image-to-image mechanism. The difference between the input and output of the diffusion model is used as the anomaly detection score, which requires the diffusion model to reconstruct the abnormal regions in the input image. However, the image-to-image diffusion model suffers from "catastrophic forgetting" when continuously generating diverse samples. Taking ControlNet as an example, its performance on the base tasks (10 categories in MVTec) is excellent, but as the model continues to be trained on new tasks (new categories), the performance drops severely. In addition, ControlNet trained on the latest categories suffers from "faithfulness hallucination". Specifically, the output of the model becomes inconsistent with the previously trained samples, resulting in the failure of anomaly detection. Moreover, classical anomaly detection methods based on diffusion models (DMs) tend to "overfit" normal samples rather than focusing on "reconstructing" abnormal regions, which weakens the role of the image-to-image conditioning mechanism and exacerbates the hallucination problem in continuous learning. Summary of the Invention
[0004] The object of the present invention is to propose an anomaly detection method based on a continuous diffusion model. This method projects the updated gradients of a new task into a subspace orthogonal to the important representations of previous tasks, thereby eliminating the influence of the updated gradients on the previous feature space. Since the column basis matrix is calculated by performing singular value decomposition (SVD) on the matrix expanded by intermediate features, the gradient projection operation requires a large amount of memory. However, the Markov-based diffusion process significantly increases the memory overhead. Taking the U-Net of Stable Diffusion as an example, calculating the column basis matrix for 10 images consumes almost an additional 157 GB of memory, and more images require more memory. To solve this problem, the present method proposes an iterative singular value decomposition (iSVD) method based on the linear transfer property. Through online iterative operations, iSVD only requires a memory consumption of approximately 17 GB to calculate the column basis matrix for any number of images.
[0005] In addition, the present method proposes an anomaly mask network. After using a CNN to encode the input image and matching its size with the input of the U-Net, the present invention also designs a non-local encoder using a Transformer structure to perceive global information. Neighbor mask self-attention and anomaly mask loss are used to mask abnormal features, thereby retaining the normal features output by the local encoder, enabling the diffusion model to pay more attention to the reconstruction of abnormal regions rather than overfitting to generate normal images.
[0006] The present invention alleviates the problems of the prior art from two aspects: 1) using a stable parameter update strategy, and 2) strengthening the conditional mechanism for image-to-image generation.
[0007] The specific technical solution for achieving the object of the present invention is as follows:
[0008] An anomaly detection method based on a continuous diffusion model. This method uses an anomaly mask network to control the diffusion model to reconstruct abnormal regions, and performs pixel-level distance measurement between the reconstructed image and the original image to perceive abnormal regions. At the same time, a gradient projection strategy is used to alleviate the catastrophic forgetting of the diffusion model in continuous anomaly detection. Finally, an iterative singular value decomposition technique is proposed to save more than 90% of the memory consumption of gradient projection, enabling the network model to better adapt to continuous anomaly detection tasks. The method includes the following specific steps:
[0009] Step 1: Preparation and preprocessing of the dataset
[0010] 1-1: Collect pictures of the object to be detected, including different types of images, and use them as the training set and test set in the dataset. The pictures used in the training set are normal images, that is, they do not contain any abnormalities; the pictures used in the test set include both normal images and abnormal images with abnormalities, as well as the annotation information of the abnormal areas; there are a total of 15 categories; the continuous training of the model includes 14 - 1, 3 5, 10 - 5 and 10 – 1 Four settings of 5, which respectively represent initially training 14 classes and then incrementing 1 class; initially training 3 classes and then incrementing 4 times, 3 classes each time; initially training 10 classes and then incrementing 5 classes; initially training 10 classes and then incrementing 5 times, 1 class each time;
[0011] 1-2: Image and label preprocessing. Adjust the images in the training set and test set to a size of 256×256, and also adjust the label masks of the test set to a size of 256×256, and perform normalization processing;
[0012] Step 2: Visual feature encoding processing
[0013] 2-1: During the training process, for each batch of training data, input the training images into the VAE for encoding to extract the latent space features of the images ;
[0014] 2-2: Add random steps of Gaussian noise to the shallow space features to obtain
[0015] 2-3: Randomly shuffle the pixels at the patch level of the original image to obtain the perturbed image;
[0016] 2-4 Input the original image and the perturbed image into the anomaly mask network to respectively extract the original features and perturbed features ;
[0017] Step 3: Training of the diffusion model and the anomaly mask network
[0018] 3-1: Input the randomly step-added noise features into the U-Net, use as the conditional control of the U-Net and at the same time use as the supervision signal to predict the noise, so that the U-Net ( ) learns the injected noise at each moment. The specific loss function is as follows:
[0019]
[0020] wherein represents the input image, represents the perturbed image, represents the latent space features encoded by VAE, represents the latent space features after adding noise in T steps, represents the standard normal distribution, represents U-Net, represents the anomaly mask network.
[0021] 3-2: Align the original features and the perturbed features by mean squared error loss to train the anomaly mask network, so that the model can learn to mask the abnormal areas with perturbations. The specific loss function is as follows:
[0022]
[0023] wherein the features of the perturbed image, are the original image features.
[0024] Step 4: Anomaly detection and localization prediction
[0025] 4-1: Input the image of the object to be detected into the VAE encoder to obtain the latent space features ;
[0026] 4-2: Then obtain the features after adding noise through T steps of noise addition ;
[0027] 4-3: Input the image of the object to be detected into the trained anomaly mask network to extract the features ;
[0028] 4-4: Use the features as the condition, use the trained U-Net to predict the noise, and finally obtain the reconstructed latent space features through T steps of prediction;
[0029] 4-5: Finally, reconstruct into an anomaly-free image through the VAE decoder;
[0030] 4-6: Input the original image and the reconstructed image into the pre-trained ResNet50 network at the same time to extract the feature maps, denoted as and ;
[0031] 4 - 7: Calculate the feature map and Calculate the Euclidean distance between the feature maps at each position as the pixel - level anomaly score at that position. That is, the larger the distance, the greater the degree of anomaly. In addition, take the maximum anomaly score among all positions in each image as the image - level anomaly score;
[0032]
[0033] where represents the anomaly score at position , respectively represent the feature maps of the original image and the reconstructed image by the diffusion model.
[0034] Step 5: Continual learning
[0035] 5 - 1: Divide the initial training set and the incremental training set for continual learning according to the settings in Step 1, and complete the pre - processing operations.
[0036] 5 - 2: After training and testing the U - Net and the anomaly mask network according to Step 3 and Step 4, extract and retain the overall intermediate - layer features of the current training set;
[0037] 5 - 3: Expand all the intermediate - layer features into matrix form and perform block division in the column direction;
[0038] 5 - 4: Through the proposed iterative singular value decomposition technique, first perform singular value decomposition on the first sub - matrix to obtain the unitary matrix, retain the top 97% eigenvectors with the largest eigenvalues and fuse them with the next sub - matrix, and then perform singular value decomposition iteratively in turn to finally obtain the column - basis matrix of the intermediate - layer features ;
[0039] 5 - 5: When training on the new incremental training set, project the training gradient onto the orthogonal direction of the matrix so that the training of the new task will not affect the old feature space.
[0040]
[0041] where represents the original training gradient of the th layer, represents the column - basis matrix of the intermediate features of the th layer, represents the training gradient after projection of the th layer.
[0042] 5 - 6: Go back to Step 5 - 2 until all incremental training sets are trained.
[0043] Compared with the prior art, the present invention has better continuous learning ability, can better adapt to the change of detected object categories, and prevent the occurrence of forgetting problems, greatly enhancing the robustness of the model in a dynamically changing industrial scenario, and further empowering industrial intelligent quality inspection. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is the overall flowchart of the present invention;
[0045] Figure 2 is the network training flowchart of the present invention;
[0046] Figure 3 is the network testing flowchart of the present invention;
[0047] Figure 4 is the continuous learning flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] For ease of understanding of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0049] Refer to Figure 1 and Figure 2 , the present invention realizes the reconstruction of the abnormal region of the image by fine-tuning the diffusion model under the control of the abnormal mask network, and at the same time alleviates the catastrophic forgetting problem of the diffusion model in continuous learning through gradient projection technology. Finally, the proposed iterative singular value decomposition technology can significantly reduce the memory consumption of gradient projection. Thus, it provides a more effective continuous update ability and improves the abnormal detection and localization effect. It is divided into five stages: S1 dataset preparation and preprocessing; S2 visual feature encoding processing; S3 diffusion model and abnormal mask network training; S4 abnormal detection and localization prediction; S5 continuous learning.
[0050] Refer to Figure 1 , the overall flowchart of the present invention
[0051] S1: Dataset preparation and preprocessing;
[0052] S2: Visual feature encoding processing;
[0053] S3: Diffusion model and abnormal mask network training;
[0054] S4: Abnormal detection and localization prediction;
[0055] S5: Continuous learning.
[0056] Refer to Figure 2 , the model training is carried out according to the following steps:
[0057] S110: Collect the dataset, perform pixel-level annotation on the abnormal regions, and divide the dataset according to different continual learning settings in accordance with Step 1-1.
[0058] S120: Preprocess the images by normalizing them;
[0059] S210: Input the preprocessed batch training image data into the VAE encoder to extract latent space features.
[0060] S220: Perform forward noise addition on the latent space features encoded by VAE.
[0061] S230: Perturb the preprocessed batch training data with patches.
[0062] S240: Input the preprocessed batch training image data and the training data perturbed with patches into the anomaly mask network to extract feature maps and
[0063] S310: Use the feature maps extracted by the anomaly mask network and the latent space features with forward noise addition in S220 to train the U-Net.
[0064] S320: Use the L2 loss to train the anomaly mask network to and align the features.
[0065] Refer to Figure 3 , and conduct model testing according to the following steps
[0066] S110: Collect the dataset, perform pixel-level annotation on the abnormal regions, and divide the dataset according to different continual learning settings in accordance with Step 1-1.
[0067] S120: Preprocess the images by normalizing them;
[0068] S410: In the testing phase, first input the preprocessed images into the VAE encoder to extract latent space features.
[0069] S420: Perform noise addition on the latent space features.
[0070] S430: Input the preprocessed images into the anomaly mask network to extract feature maps .
[0071] S440: Use the feature maps in S420 as control conditions to denoise the latent space features with added noise.
[0072] S450: Use VAE to reconstruct the denoised latent space features into images.
[0073] S510: Input the original preprocessed image in S410 and the reconstructed image in S450 into ResNet50 respectively to extract feature maps and 。
[0074] S520: Calculate the distance between each position of the feature maps and to obtain the final anomaly score map.
[0075] Refer to Figure 4 and perform the continuous learning process according to the following steps
[0076] S510: In the continuous learning stage, first divide and preprocess the dataset.
[0077] S520: Train and test the U-Net and the anomaly mask network.
[0078] S530: Expand all intermediate layer features into matrix form and perform block division in the column vector direction.
[0079] S540: Calculate the column basis matrix through the iterative singular value decomposition technique.
[0080] S550: In the new incremental training, project the training gradient onto the orthogonal direction of the column basis matrix calculated in S440.
[0081] S560: Return to S520 until all incremental training sets are trained.
[0082] The comparison between the present invention and the prior art in terms of anomaly detection (image-level / pixel-level AUROC) is shown in Table 1 below:
[0083]
[0084] Taking 10 images as an example, the memory saving rate of the iterative singular value decomposition of the present invention is shown in Table 2 below:
[0085]
[0086] It can be seen from Table 1 above that the present invention has significantly improved in the comparison of continuous anomaly detection and localization metrics with advanced anomaly detection methods on the MVTec dataset. It can be seen from Table 2 that the iterative singular value decomposition proposed by the present invention can effectively save the memory occupancy rate.
[0087] The above is only a further description of the present invention and is not intended to limit the present invention. All equivalent implementations of the present invention should be included within the scope of the claims of the present invention.
Claims
1. An anomaly detection method based on a continuous diffusion model, characterized in that: The method comprises the following specific steps: Step 1: Dataset preparation and preprocessing 1-1: Collect images of objects to be detected, including images of different types, and use them as training sets and test sets in the data set. The images used in the training set are normal images, that is, they do not contain any abnormalities; the images used in the test set include both normal images and abnormal images with abnormalities, as well as annotation information of abnormal areas, a total of 15 categories; continuous training of the model includes 14 - 1,3 5, 10-5 and 10-1 There are four settings, which represent initial training of 14 classes and then incrementing by 1 class; initial training of 3 classes and then incrementing by 4 times, each time by 3 classes; initial training of 10 classes and then incrementing by 5 classes; initial training of 10 classes and then incrementing by 5 times, each time by 1 class; 1-2: Image and label preprocessing: resize the images in the training set and test set to 256×256, and resize the label masks of the test set to 256×256, and perform normalization. Step 2: Visual feature encoding processing 2-1: During the training process, for each batch of training data, the training image is input into VAE for encoding to extract the latent space features of the image ; 2-2: Latent space features Add Random The Gaussian noise of the step is obtained ; 2-3: Randomly shuffle the original image at the patch level to obtain the perturbed image; 2-4 Input the original image and the perturbed image into the abnormal mask network to extract the original features respectively and disturbance characteristics ; Step 3: Diffusion model and anomaly mask network training 3-1: Random Characteristics of step noise Input to U-Net, using perturbation features Used as a conditional control for U-Net As a supervisory signal, the noise is predicted, so that U-Net ( ) to learn each The noise injected at the moment, the specific loss function is as follows: ;in represents the input image, represents the perturbed image, represents the latent space features after VAE encoding, express The latent space features after step noise addition, represents the standard normal distribution, represents U-Net, represents anomaly mask network; 3-2: Aligning original features through mean square error loss and disturbance characteristics To train the anomaly mask network, the model can learn to mask out abnormal areas with disturbances. The specific loss function is as follows: ;in The image features after perturbation, is the original image feature; Step 4: Anomaly detection and location prediction 4-1: The image of the object to be detected Input to VAE encoder to obtain latent space features ; 4-2: Then, the noise-added features are obtained through T-step noise addition processing. ; 4-3: Image of the object to be detected Input into the trained anomaly mask network to extract features ; 4-4: Use features As a condition, use the trained U-Net to predict The noise is removed, and the reconstructed latent space features are finally obtained through T-step prediction ; 4-5: Finally Reconstructed into an abnormal-free image through the VAE decoder ; 4-6: The original image and the reconstructed image At the same time, it is input into the pre-trained ResNet50 network to extract the feature maps, which are recorded as and ; 4-7: Calculate feature map and The Euclidean distance between the features at each position is taken as the pixel-level anomaly score of that position. That is, the larger the distance, the greater the degree of anomaly. In addition, the maximum anomaly score among all positions of each image is taken as the image-level anomaly score. ;in Indicates at location The anomaly score, Represent the feature maps of the original image and the image reconstructed by the diffusion model respectively; Step 5: Keep learning 5-1: Divide the initial training set and incremental training set for continuous learning according to the settings in step 1, and complete the preprocessing operation; 5-2: After training and testing the U-Net and the anomaly mask network according to steps 3 and 4, extract the overall intermediate layer features of the current training set and retain them; 5-3: Expand all the intermediate layer features into a matrix form and divide them into blocks in the column direction; 5-4: Iterative singular value decomposition: First, perform singular value decomposition on the first sub-matrix to obtain a unitary matrix, retain the first 97% eigenvectors with the largest eigenvalues and merge them with the next sub-matrix, then iterate the singular value decomposition in sequence, and finally obtain the column basis matrix of the intermediate layer features ; 5-5: When training with a new incremental training set, the training gradient Projection to Matrix Orthogonal direction, so that the training of new tasks will not affect the old feature space; ;in, Indicates The original training gradient of the layer, Indicates The column basis matrix of the intermediate features of the layer, Indicates Training gradients after layer projection; 5-6: Return to step 5-2 until all incremental training sets are trained.