Unsupervised anomaly detection method, system and readable storage medium based on diffusion model
By extracting image style features and processing them using the S2A and S2T modules, and combining them with the diffusion model for image reconstruction, the problems of diffusion model reconstruction distortion and high-frequency feature loss are solved, and the accuracy and adaptability of anomaly detection are improved.
Patent Information
- Application Number
- CN202510644737.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The existing diffusion model has problems such as excessive initial input image noise and misjudgment of high-frequency features during the reconstruction process, resulting in poor reconstruction effect and decreased anomaly detection accuracy.
The pre-trained ResNet50 network is used to extract style features, the S2A module and S2T module are used to process the style features, the diffusion model is combined to reconstruct the image, and the abnormal image is output through the abnormality discrimination module.
It improves the accuracy of anomaly detection, enhances the sensitivity and retention of high-frequency features, adapts to the thickness changes and positional relationship deviations of printed patterns, and reduces false anomaly detection and weakening of abnormal areas.
Smart Images

Figure CN120162727B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and more specifically, to an unsupervised anomaly detection method, system and readable storage medium based on a diffusion model. Background Art
[0002] In the current field of unsupervised anomaly detection, reconstruction-based unsupervised methods are the most popular and widely used technical approach. The core idea of these methods is to reconstruct the input image using a model and detect anomalous regions by analyzing the differences between the reconstructed image and the original image. While these methods can detect anomalies in images to a certain extent, their limited reconstruction performance directly affects the accuracy of anomaly detection. The reconstructed image may also differ significantly from the original image in non-anomalous areas.
[0003] Existing diffusion models face two main problems during reconstruction:
[0004] 1. The initial input image is too noisy: The initial input image of the diffusion model consists of "abnormal image + a lot of noise". The model needs to reconstruct the original image from the strong noise, which is extremely difficult and leads to poor reconstruction effect.
[0005] 2. Misjudgment of high-frequency features: The diffusion model is essentially a denoising model. During the denoising process, the model often misjudges high-frequency features in the image (such as edges and textures) as noise, thereby losing this key information and affecting the accuracy of anomaly detection. Summary of the Invention
[0006] The purpose of the present invention is to provide an unsupervised anomaly detection method, system and readable storage medium based on a diffusion model, which can effectively solve the problems of reconstruction distortion and high-frequency feature loss in existing reconstruction methods.
[0007] A first aspect of the present invention provides an unsupervised anomaly detection method based on a diffusion model, comprising the following steps:
[0008] Use the pre-trained ResNet50 network to extract style features and the pre-trained encoder to obtain the latent space representation of the input image;
[0009] The style features are processed using a preset S2A module and a preset S2T module to obtain processed style features and input them into a diffusion model; at the same time, the latent space representation is input into the diffusion model, and a final denoised latent space representation is output;
[0010] Decode the final denoised latent space representation to obtain the reconstructed image;
[0011] The style features and reconstructed images are respectively input into the anomaly discrimination module, and the anomaly map is output.
[0012] In this solution, the style feature extraction using the pre-trained ResNet50 network includes the following steps:
[0013] Given an input image , feature maps are extracted using the first four layers of the pre-trained ResNet50 network;
[0014] ,
[0015] ,
[0016] in, , , , Represent the feature maps extracted from the first four convolutional layers of the ResNet50 network; Represents the output of the nth layer, where n=1,2,3,4.
[0017] Calculate the Gram matrix of each layer feature map:
[0018] For each extracted feature map Perform Gram matrix calculation to describe the style characteristics of the image, where =1,2,3,4:
[0019] ,
[0020] in, Indicates the The Gram matrix of the layer, , , Respectively represent The number of channels, height and width of the layer feature map, Indicates the The feature map of the layer has a dimension of .
[0021] In this solution, the S2A module is used to process the style features, including the following steps:
[0022] Use the attention mechanism to calculate the attention weights through the shallow feature map:
[0023] ,
[0024] in, Represents shallow feature maps After linear transformation The query matrix Q obtained; Represents shallow feature maps After linear transformation The resulting bond matrix K; represents a trainable weight matrix;
[0025] Attention weight Represents the correlation between different positions of the feature map, normalized by the softmax function;
[0026] Apply the generated attention weights to the value matrix V of the feature map to generate an enhanced feature map:
[0027] ,
[0028] in, Represents shallow feature maps After linear transformation The obtained value matrix V; Represents a trainable weight matrix.
[0029] In this solution, the S2T module processes the style features, including the following steps:
[0030] Input the style feature G extracted from the Gram matrix into the conversion network to generate a fixed-length token:
[0031] ,
[0032] in, Represents the token generated by transforming style features; Represents the conversion network function that converts the Gram matrix feature G into a fixed-length token;
[0033] The style token Conditional characteristics of the diffusion model Combined to form the final diffusion model input:
[0034] ,
[0035] in, represents the conditional features used as input to the diffusion model, Indicates that the style token and conditional features Stitching in a specific dimension.
[0036] In this solution, the conversion network generates a fixed-length token. The steps include:
[0037] First the input needs to be flattened into a vector:
[0038] ,
[0039] in Represents a flattening function that flattens a multidimensional data structure into one dimension;
[0040] set up ,but , represents the field of real numbers;
[0041] Then, the dimension is gradually reduced through two fully connected layers, and a multi-head self-attention mechanism (MHA) is inserted between the two fully connected layers; finally, a fixed-length Token vector is output. , where d is the dimension of token;
[0042] ,
[0043] in, represents the fully connected weight matrix, represents the bias term, Represents the nonlinear activation function ReLU.
[0044] In this solution, the style features and the reconstructed image are input into the anomaly discrimination module respectively, and the anomaly map is output. The steps include:
[0045] For the input image and the reconstructed image , respectively use the pre-trained ResNet50 to extract the intermediate layer features of the image, and get ,in ;
[0046] Calculate the cosine similarity of the original image and the reconstructed image for the features extracted from the 2nd, 3rd, and 4th layers respectively:
[0047] ,
[0048] in Indicates the difference between the input image and the reconstructed image in Cosine similarity graph of layers;
[0049] The similarity maps calculated at each layer are resized to the same size as the input image, and all similarity maps are weighted fused:
[0050] ,
[0051] in, represents the final generated anomaly graph; Resizes the similarity map to the resolution of the input image.
[0052] This solution also includes setting up category-aware learning rate scheduling and pre-training models.
[0053] The category-aware learning rate scheduling is designed for multi-classification scenarios. It sets an independent learning rate scheduler for each category to address the problem of training progress differences caused by uneven category distribution:
[0054] ,
[0055] in, :category The current learning rate of : initial learning rate; : Scheduling factor, used to control the learning rate decay speed; : Current training round.
[0056] In this solution, the pre-trained models, namely the diffusion model and the feature extraction module, are loaded with pre-trained weights respectively. The diffusion model is loaded with diffusion model weights pre-trained on a large-scale image generation dataset; the feature extraction module is loaded with ResNet50 weights pre-trained on the ImageNet dataset.
[0057] A second aspect of the present invention provides an unsupervised anomaly detection system based on a diffusion model, comprising a memory and a processor, wherein the memory includes an unsupervised anomaly detection method program based on a diffusion model, and when the unsupervised anomaly detection method program based on a diffusion model is executed by the processor, the following steps are implemented:
[0058] Use the pre-trained ResNet50 network to extract style features and the pre-trained encoder to obtain the latent space representation of the input image;
[0059] The style features are processed using a preset S2A module and a preset S2T module to obtain processed style features and input them into a diffusion model; at the same time, the latent space representation is input into the diffusion model, and a final denoised latent space representation is output;
[0060] Decode the final denoised latent space representation to obtain the reconstructed image;
[0061] The style features and reconstructed images are respectively input into the anomaly discrimination module, and the anomaly map is output.
[0062] The third aspect of the present invention provides a computer-readable storage medium, which includes a machine program for an unsupervised anomaly detection method based on a diffusion model. When the program for an unsupervised anomaly detection method based on a diffusion model is executed by a processor, the steps of an unsupervised anomaly detection method based on a diffusion model as described in any one of the above items are implemented.
[0063] The present invention discloses an unsupervised anomaly detection method, system, and readable storage medium based on a diffusion model. The present invention extracts style features of an image, processes the style features using an S2A module and an S2T module, injects the processed features into a decoding layer of the diffusion model, and simultaneously injects latent space representations into different layers of the diffusion model decoding layer. This enhances the model's sensitivity and retention of high-frequency features, minimizes the difference between the images before and after reconstruction in non-abnormal areas, and improves the accuracy of anomaly detection.
[0064] It can effectively detect unsupervised anomalies based on the diffusion model and can adapt to changes in the thickness of printed pattern strokes and positional relationship deviations.
[0065] In response to positional relationship deviations, the overall template is broken up and only local useful templates are retained for flexible combination, which improves the algorithm's adaptability to positional deviations and detection accuracy.
[0066] In response to changes in the thickness of pattern strokes, by learning the thickness characteristics of the strokes, defects can be corrected using the thickness characteristics during defect detection, effectively avoiding over-inspection caused by stroke thickness.
[0067] For non-defect interference items, a pattern mask is obtained, and a light and dark defect shielding mask is created in the defect detection link. Shielding the detection interference caused by non-defects can effectively avoid over-inspection caused by non-defective parts. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 The flowchart of the present invention is a method for unsupervised anomaly detection based on a diffusion model;
[0069] Figure 2 A block diagram of an unsupervised anomaly detection system based on a diffusion model of the present invention is shown. DETAILED DESCRIPTION
[0070] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein can be combined with each other.
[0071] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0072] Figure 1 The flowchart of the present application is a method for unsupervised anomaly detection based on a diffusion model.
[0073] like Figure 1 As shown, the present application discloses an unsupervised anomaly detection method based on a diffusion model, comprising the following steps:
[0074] S102, extracting style features using a pre-trained ResNet50 network and obtaining a latent space representation of the input image using a pre-trained encoder;
[0075] S104, respectively processing the style features using a preset S2A module and a preset S2T module to obtain processed style features and inputting them into a diffusion model; simultaneously inputting the latent space representation into the diffusion model, and outputting a final denoised latent space representation;
[0076] S106, decoding the final denoised latent space representation to obtain a reconstructed image;
[0077] S108: Input the style features and the reconstructed image into the anomaly discrimination module respectively, and output an anomaly map.
[0078] It should be noted that in anomaly detection tasks, the model needs to accurately distinguish between normal and abnormal regions. To this end, the characteristics of normal regions must be preserved as much as possible during reconstruction. Otherwise, problems such as false anomaly detection and weakening of abnormal regions may occur.
[0079] False anomaly detection: If a normal area undergoes significant changes after reconstruction, it will be mistaken for an abnormal area, resulting in an increased false detection rate.
[0080] Weakening of abnormal regions: When the model fails to accurately capture the characteristics of abnormal regions during the reconstruction process, the abnormal regions may be over-smoothed, making them difficult to detect effectively.
[0081] Therefore, a mechanism needs to be designed to lock the features of normal areas during the reconstruction process to prevent them from changing significantly. At the same time, the model needs to ensure that abnormal areas are appropriately highlighted during the reconstruction process so that they can be clearly shown in the difference map.
[0082] This paper proposes an unsupervised anomaly detection method based on a diffusion model. By improving the diffusion model's reconstruction performance, it overcomes the problems of reconstruction distortion and high-frequency feature loss in existing reconstruction methods. The core of this method is to extract image style features, convert deep-level style features into tokens, and inject them into all decoding layers. Simultaneously, shallow-level style features are injected into different layers of the decoding layer, enhancing the model's sensitivity to and ability to retain high-frequency features.
[0083] It should be noted that the present invention adopts the Latent Diffusion Model (LDM) to significantly reduce computational complexity while maintaining generation quality by performing diffusion and denoising operations in a low-dimensional latent space. The optimization objectives of the diffusion model are:
[0084] ,
[0085] in, represents the latent space representation of the diffusion model at time step t; represents random noise from a standard normal distribution; Represents the noise value predicted by the model.
[0086] By mapping high-dimensional images into latent space representations , the diffusion model completes the step-by-step denoising process in the latent space. The latent space represents the encoder trained by get:
[0087] ,
[0088] in, is the input image.
[0089] The denoised latent space representation is passed through the decoder Restore to a high-dimensional image:
[0090] ,
[0091] in, is the final latent space representation after denoising.
[0092] It should be noted that style features are important representations of images, reflecting texture, structure and global information.
[0093] According to an embodiment of the present invention, extracting style features using a pre-trained ResNet50 network includes the following steps:
[0094] Given an input image , feature maps are extracted using the first four layers of the pre-trained ResNet50 network;
[0095] , , , ,
[0096] in, , , , Represent the feature maps extracted from the first four convolutional layers of the ResNet50 network; Represents the output of the nth layer, where n=1,2,3,4.
[0097] Calculate the Gram matrix of each layer feature map:
[0098] For each extracted feature map (in =1,2,3,4) to calculate the Gram matrix to describe the style characteristics of the image:
[0099] ,
[0100] in, Indicates the The Gram matrix of the layer, , , Respectively represent The number of channels, height and width of the layer feature map, Indicates the The feature map matrix of the layer has the dimension .
[0101] It should be noted that the preset S2A module (i.e. Style-to-Attention) module is designed to utilize shallow style features to enhance the ability to retain high-frequency details.
[0102] According to an embodiment of the present invention, processing the style features using a preset S2A module includes the following steps:
[0103] Use the attention mechanism to calculate the attention weights through the shallow feature map:
[0104] ,
[0105] in, Represents shallow feature maps After linear transformation The query matrix Q obtained; Represents shallow feature maps After linear transformation The resulting bond matrix K; represents a trainable weight matrix;
[0106] Attention weight Represents the correlation between different positions of the feature map, normalized by the softmax function;
[0107] Apply the generated attention weights to the value matrix V of the feature map to generate an enhanced feature map:
[0108] ,
[0109] in, Represents shallow feature maps After linear transformation The obtained value matrix V; Represents a trainable weight matrix.
[0110] According to an embodiment of the present invention, the preset S2T module processes the style features, including the following steps:
[0111] Input the style feature G extracted from the Gram matrix into the conversion network to generate a fixed-length token:
[0112] ,
[0113] in, Represents the token generated by transforming style features; Represents the conversion network function that converts the Gram matrix feature G into a fixed-length token;
[0114] The style token Conditional characteristics of the diffusion model Combined to form the final diffusion model input:
[0115] ,
[0116] in, represents the conditional features used as input to the diffusion model, Indicates that the style token and conditional features Stitching in a specific dimension.
[0117] In this solution, the conversion network generates a fixed-length token. The steps include:
[0118] First the input needs to be flattened into a vector:
[0119] ,
[0120] in Represents a flattening function that flattens a multidimensional data structure into one dimension;
[0121] set up ,but , represents the field of real numbers;
[0122] Then, the dimension is gradually reduced through two fully connected layers. At the same time, in order to enhance the semantic expression of style features, a multi-head self-attention mechanism (MHA) is inserted between the two fully connected layers; finally, a fixed-length Token vector is output. , where d is the dimension of token;
[0123] ,
[0124] in, represents the fully connected weight matrix, represents the bias term, Represents the nonlinear activation function ReLU.
[0125] According to an embodiment of the present invention, the style features and the reconstructed image are respectively input into an abnormality discrimination module, and an abnormality map is output. The steps include:
[0126] For the input image and the reconstructed image , respectively use the pre-trained ResNet50 to extract the intermediate layer features of the image, and get ,in ;
[0127] Calculate the cosine similarity of the original image and the reconstructed image for the features extracted from the 2nd, 3rd, and 4th layers respectively:
[0128] ,
[0129] in Indicates the difference between the input image and the reconstructed image in Cosine similarity graph of layers;
[0130] The similarity graph calculated at each layer Resize to the same size as the input image and perform a weighted fusion of all similarity maps:
[0131] ,
[0132] in, represents the final generated anomaly graph; Resizes the similarity map to the resolution of the input image.
[0133] According to an embodiment of the present invention, the present invention also includes the setting of a training strategy, wherein the training strategy includes category-aware learning rate scheduling and pre-training models,
[0134] The category-aware learning rate scheduling is designed for multi-classification scenarios. It sets an independent learning rate scheduler for each category to address the problem of training progress differences caused by uneven category distribution:
[0135] ,
[0136] in, :category The current learning rate of : initial learning rate; : Scheduling factor, used to control the learning rate decay speed; : Current training round.
[0137] According to an embodiment of the present invention, the pre-trained models, namely the diffusion model and the feature extraction module, respectively load pre-trained weights, wherein the diffusion model: loads the diffusion model weights SD1.5 pre-trained on a large-scale image generation dataset; the feature extraction module loads the ResNet50 weights pre-trained on the ImageNet dataset.
[0138] It should be noted that, in order to illustrate and verify the method of the present invention, the following experimental data are further described:
[0139] Dataset: Digital Domain MVTec contains 15 categories of real-world data for anomaly detection, including 5 categories of texture data and 10 categories of object data. The training set contains 3,629 non-anomaly samples, and the test set contains 1,725 samples containing various types of anomalies or non-anomalies.
[0140] Experimental setup: All images were scaled to 256 × 256. Only positive images from the training set were sampled for each training run. Negative images were not used in training and were only used for model evaluation.
[0141] Experimental Results: Table 1 shows the pixel-level AUROC mean for anomaly detection on different datasets in MVTec. The experimental results show that this method can effectively improve the anomaly detection accuracy compared to the initial diffusion model (SD1.5);
[0142] Table 1. Mean pixel-level AUROC of anomaly detection in different categories of datasets
[0143] .
[0144] Figure 2 A block diagram of an unsupervised anomaly detection system based on a diffusion model of the present invention is shown.
[0145] like Figure 2 As shown, the second aspect of the present invention provides an unsupervised anomaly detection system based on a diffusion model, including a memory and a processor. The memory includes an unsupervised anomaly detection method program based on a diffusion model. When the unsupervised anomaly detection method program based on a diffusion model is executed by the processor, the following steps are implemented:
[0146] S102, extracting style features using a pre-trained ResNet50 network and obtaining a latent space representation of the input image using a pre-trained encoder;
[0147] S104, respectively processing the style features using a preset S2A module and a preset S2T module to obtain processed style features and inputting them into a diffusion model; simultaneously inputting the latent space representation into the diffusion model, and outputting a final denoised latent space representation;
[0148] S106, decoding the final denoised latent space representation to obtain a reconstructed image;
[0149] S108: Input the style features and the reconstructed image into the anomaly discrimination module respectively, and output an anomaly map.
[0150] It should be noted that in anomaly detection tasks, the model needs to accurately distinguish between normal and abnormal regions. To this end, the characteristics of normal regions must be preserved as much as possible during reconstruction. Otherwise, problems such as false anomaly detection and weakening of abnormal regions may occur.
[0151] False anomaly detection: If a normal area undergoes significant changes after reconstruction, it will be mistaken for an abnormal area, resulting in an increased false detection rate.
[0152] Weakening of abnormal regions: When the model fails to accurately capture the characteristics of abnormal regions during the reconstruction process, the abnormal regions may be over-smoothed, making them difficult to detect effectively.
[0153] Therefore, a mechanism needs to be designed to lock the features of normal areas during the reconstruction process to prevent them from changing significantly. At the same time, the model needs to ensure that abnormal areas are appropriately highlighted during the reconstruction process so that they can be clearly shown in the difference map.
[0154] This paper proposes an unsupervised anomaly detection method based on a diffusion model. By improving the diffusion model's reconstruction performance, it overcomes the problems of reconstruction distortion and high-frequency feature loss in existing reconstruction methods. The core of this method is to extract image style features, convert deep-level style features into tokens, and inject them into all decoding layers. Simultaneously, shallow-level style features are injected into different layers of the decoding layer, enhancing the model's sensitivity to and ability to retain high-frequency features.
[0155] It should be noted that the present invention adopts the Latent Diffusion Model (LDM) to significantly reduce computational complexity while maintaining generation quality by performing diffusion and denoising operations in a low-dimensional latent space. The optimization objectives of the diffusion model are:
[0156] ,
[0157] in, represents the latent space representation of the diffusion model at time step t; represents random noise from a standard normal distribution; Represents the noise value predicted by the model.
[0158] By mapping high-dimensional images into latent space representations , the diffusion model completes the step-by-step denoising process in the latent space. The latent space represents the encoder trained by get:
[0159] ,
[0160] in, is the input image.
[0161] The denoised latent space representation is passed through the decoder Restore to a high-dimensional image:
[0162] ,
[0163] in, is the final latent space representation after denoising.
[0164] It should be noted that style features are important representations of images, reflecting texture, structure and global information.
[0165] According to an embodiment of the present invention, extracting style features using a pre-trained ResNet50 network includes the following steps:
[0166] Given an input image , feature maps are extracted using the first four layers of the pre-trained ResNet50 network;
[0167] , , , ,
[0168] in, , , , Represent the feature maps extracted from the first four convolutional layers of the ResNet50 network; Represents the output of the nth layer, where n=1,2,3,4.
[0169] Calculate the Gram matrix of each layer feature map:
[0170] For each extracted feature map (in =1,2,3,4) to calculate the Gram matrix to describe the style characteristics of the image:
[0171] ,
[0172] in, Indicates the The Gram matrix of the layer, , , Respectively represent The number of channels, height and width of the layer feature map, Indicates the The feature map matrix of the layer has the dimension .
[0173] It should be noted that the preset S2A module (i.e. Style-to-Attention) module is designed to utilize shallow style features to enhance the ability to retain high-frequency details.
[0174] According to an embodiment of the present invention, processing the style features using a preset S2A module includes the following steps:
[0175] Use the attention mechanism to calculate the attention weights through the shallow feature map:
[0176] ,
[0177] in, Represents shallow feature maps After linear transformation The query matrix Q obtained; Represents shallow feature maps After linear transformation The resulting bond matrix K; represents a trainable weight matrix;
[0178] Attention weight Represents the correlation between different positions of the feature map, normalized by the softmax function;
[0179] Apply the generated attention weights to the value matrix V of the feature map to generate an enhanced feature map:
[0180] ,
[0181] in, Represents shallow feature maps After linear transformation The obtained value matrix V; Represents a trainable weight matrix.
[0182] According to an embodiment of the present invention, the preset S2T module processes the style features, including the following steps:
[0183] Input the style feature G extracted from the Gram matrix into the conversion network to generate a fixed-length token:
[0184] ,
[0185] in, Represents the token generated by transforming style features; Represents the conversion network function that converts the Gram matrix feature G into a fixed-length token;
[0186] The style token Conditional characteristics of the diffusion model Combined to form the final diffusion model input:
[0187] ,
[0188] in, represents the conditional features used as input to the diffusion model, Indicates that the style token and conditional features Stitching in a specific dimension.
[0189] In this solution, the conversion network generates a fixed-length token. The steps include:
[0190] First the input needs to be flattened into a vector:
[0191] ,
[0192] in Represents a flattening function that flattens a multidimensional data structure into one dimension;
[0193] set up ,but , represents the field of real numbers;
[0194] Then, the dimension is gradually reduced through two fully connected layers. At the same time, in order to enhance the semantic expression of style features, a multi-head self-attention mechanism (MHA) is inserted between the two fully connected layers; finally, a fixed-length Token vector is output. , where d is the dimension of token;
[0195] ,
[0196] in, represents the fully connected weight matrix, represents the bias term, Represents the nonlinear activation function ReLU.
[0197] According to an embodiment of the present invention, the style features and the reconstructed image are respectively input into an abnormality discrimination module, and an abnormality map is output. The steps include:
[0198] For the input image and the reconstructed image , respectively use the pre-trained ResNet50 to extract the intermediate layer features of the image, and get ,in ;
[0199] Calculate the cosine similarity of the original image and the reconstructed image for the features extracted from the 2nd, 3rd, and 4th layers respectively:
[0200] ,
[0201] in Indicates the difference between the input image and the reconstructed image in Cosine similarity graph of layers;
[0202] The similarity maps calculated at each layer are resized to the same size as the input image, and all similarity maps are weighted fused:
[0203] ,
[0204] in, represents the final generated anomaly graph; Resizes the similarity map to the resolution of the input image.
[0205] According to an embodiment of the present invention, the present invention also includes the setting of a training strategy, wherein the training strategy includes category-aware learning rate scheduling and pre-training models,
[0206] The category-aware learning rate scheduling is designed for multi-classification scenarios. It sets an independent learning rate scheduler for each category to address the problem of training progress differences caused by uneven category distribution:
[0207] ,
[0208] in, :category The current learning rate of : initial learning rate; : Scheduling factor, used to control the learning rate decay speed; : Current training round.
[0209] According to an embodiment of the present invention, the pre-trained models, namely the diffusion model and the feature extraction module, respectively load pre-trained weights, wherein the diffusion model: loads the diffusion model weights SD1.5 pre-trained on a large-scale image generation dataset; the feature extraction module loads the ResNet50 weights pre-trained on the ImageNet dataset.
[0210] The third aspect of the present invention provides a computer-readable storage medium, which includes a machine program for an unsupervised anomaly detection method based on a diffusion model. When the program for an unsupervised anomaly detection method based on a diffusion model is executed by a processor, the steps of an unsupervised anomaly detection method based on a diffusion model as described in any one of the above items are implemented.
[0211] The present invention discloses an unsupervised anomaly detection method, system and readable storage medium based on a diffusion model. The present invention discloses an unsupervised anomaly detection method, system and readable storage medium based on a diffusion model. The present invention extracts style features of an image, processes the style features using a preset S2A module and a preset S2T module, injects the processed features into a decoding layer of the diffusion model, and simultaneously injects latent space representations into different layers of the diffusion model decoding layer, thereby enhancing the model's sensitivity and retention ability to high-frequency features, so that the difference between the images before and after reconstruction in non-abnormal areas is as small as possible, thereby improving the accuracy of anomaly detection.
[0212] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0213] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0214] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0215] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned method embodiments may be implemented by hardware associated with program instructions, and the aforementioned program may be stored in a computer-readable storage medium. When the program is executed, the program executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0216] Alternatively, if the integrated units described above are implemented as software modules and sold or used as standalone products, they can also be stored on a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product, stored on a storage medium, includes instructions for enabling a computer device (such as a personal computer, server, or network device) to execute all or part of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as removable storage devices, ROM, RAM, magnetic disks, or optical disks.
Claims
1. An unsupervised anomaly detection method based on a diffusion model, characterized in that: The following steps are involved: Use the pre-trained ResNet50 network to extract the style features of the input image, and use the pre-trained encoder to obtain the latent space representation of the input image; The style features are processed using a preset S2A module and a preset S2T module to obtain processed style features and input them into a diffusion model. The deep style features are converted into tokens and injected into all decoding layers. The shallow style features are injected into different layers of the decoding layer. At the same time, the latent space representation is input into the diffusion model, and the final denoised latent space representation is output. Decode the final denoised latent space representation to obtain the reconstructed image; The style features and reconstructed images are input into the anomaly discrimination module respectively, and an anomaly map is output; The method of extracting style features using the pre-trained ResNet50 network includes the following steps: Given an input image , feature maps are extracted using the first four layers of the pre-trained ResNet50 network; , , , , in, , Represent the feature maps extracted from the first four convolutional layers of the ResNet50 network; Represents the output of the nth layer, where n=1,2,3,4; Calculate the Gram matrix of each layer feature map: For each extracted feature map Perform Gram matrix calculation to describe the style characteristics of the image, where =1,2,3,4: , in, Indicates the The Gram matrix of the layer, , , Respectively represent The number of channels, height and width of the layer feature map, Indicates the The feature map of the layer has a dimension of ; Processing the style features using the S2A module includes the following steps: Use the attention mechanism to calculate the attention weights through the shallow feature map: , in, Represents shallow feature maps After linear transformation The query matrix Q obtained; Represents shallow feature maps After linear transformation The resulting bond matrix K; represents a trainable weight matrix; Attention weight Represents the correlation between different positions of the feature map, normalized by the softmax function; Apply the generated attention weights to the value matrix V of the feature map to generate an enhanced feature map: , in, Represents shallow feature maps After linear transformation The obtained value matrix V; Represents a trainable weight matrix; The S2T module processes the style features, including the following steps: Input the style feature G extracted from the Gram matrix into the conversion network to generate a fixed-length token: , in, Represents the token generated by transforming style features; Represents the conversion network function that converts the Gram matrix feature G into a fixed-length token; The style token Conditional characteristics of the diffusion model Combined to form the final diffusion model input: , in, represents the conditional features used as input to the diffusion model, Indicates that the style token and conditional features Stitching in a specific dimension.
2. The unsupervised anomaly detection method based on a diffusion model according to claim 1, characterized in that: The conversion network generates a fixed-length token. The steps include: First the input needs to be flattened into a vector: , in () represents the flatten function, which flattens a multidimensional data structure into one dimension; set up ,but , represents the field of real numbers; Then, the dimension is gradually reduced through two fully connected layers, and a multi-head self-attention mechanism (MHA) is inserted between the two fully connected layers; finally, a fixed-length Token vector is output. , where d is the dimension of token; , in, , represents the fully connected weight matrix, , represents the bias term, Represents the nonlinear activation function ReLU.
3. The unsupervised anomaly detection method based on a diffusion model according to claim 2, characterized in that: The style features and the reconstructed image are input into the anomaly discrimination module respectively, and the anomaly map is output. The steps include: For the input image and the reconstructed image , respectively use the pre-trained ResNet50 to extract the intermediate layer features of the image, and get ,in ; Calculate the cosine similarity of the original image and the reconstructed image for the features extracted from the 2nd, 3rd, and 4th layers respectively: , in Indicates the difference between the input image and the reconstructed image in Cosine similarity graph of layers; The similarity graph calculated at each layer Resize to the same size as the input image and perform a weighted fusion of all similarity maps: , in, represents the final generated anomaly graph; Resizes the similarity map to the resolution of the input image.
4. The unsupervised anomaly detection method based on a diffusion model according to claim 1, characterized in that: It also includes setting up category-aware learning rate scheduling and pre-training models. The category-aware learning rate scheduling is designed for multi-classification scenarios. It sets an independent learning rate scheduler for each category to address the problem of training progress differences caused by uneven category distribution: , in, :category The current learning rate of : initial learning rate; : Scheduling factor, used to control the learning rate decay speed; : Current training round.
5. The unsupervised anomaly detection method based on a diffusion model according to claim 4, characterized in that: The pre-trained models, namely the diffusion model and the feature extraction module, are loaded with pre-trained weights respectively, wherein the diffusion model is loaded with diffusion model weights pre-trained on a large-scale image generation dataset; the feature extraction module is loaded with ResNet50 weights pre-trained on the ImageNet dataset.
6. An unsupervised anomaly detection system based on a diffusion model, characterized in that: The system includes a memory and a processor, wherein the memory includes an unsupervised anomaly detection method program based on a diffusion model, and when the unsupervised anomaly detection method program based on a diffusion model is executed by the processor, the following steps are implemented: Use the pre-trained ResNet50 network to extract the style features of the input image, and use the pre-trained encoder to obtain the latent space representation of the input image; The style features are processed using a preset S2A module and a preset S2T module to obtain processed style features and input them into a diffusion model. The deep style features are converted into tokens and injected into all decoding layers. The shallow style features are injected into different layers of the decoding layer. At the same time, the latent space representation is input into the diffusion model, and the final denoised latent space representation is output. Decode the final denoised latent space representation to obtain the reconstructed image; The style features and reconstructed images are input into the anomaly discrimination module respectively, and an anomaly map is output; The method of extracting style features using the pre-trained ResNet50 network includes the following steps: Given an input image , feature maps are extracted using the first four layers of the pre-trained ResNet50 network; , , , , in, , Represent the feature maps extracted from the first four convolutional layers of the ResNet50 network; Represents the output of the nth layer, where n=1,2,3,4; Calculate the Gram matrix of each layer feature map: For each extracted feature map Perform Gram matrix calculation to describe the style characteristics of the image, where =1,2,3,4: , in, Indicates the The Gram matrix of the layer, , , Respectively represent The number of channels, height and width of the layer feature map, Indicates the The feature map of the layer has a dimension of ; Processing the style features using the S2A module includes the following steps: Use the attention mechanism to calculate the attention weights through the shallow feature map: , in, Represents shallow feature maps After linear transformation The query matrix Q obtained; Represents shallow feature maps After linear transformation The resulting bond matrix K; represents a trainable weight matrix; Attention weight Represents the correlation between different positions of the feature map, normalized by the softmax function; Apply the generated attention weights to the value matrix V of the feature map to generate an enhanced feature map: , in, Represents shallow feature maps After linear transformation The obtained value matrix V; Represents a trainable weight matrix; The S2T module processes the style features, including the following steps: Input the style feature G extracted from the Gram matrix into the conversion network to generate a fixed-length token: , in, Represents the token generated by transforming style features; Represents the conversion network function that converts the Gram matrix feature G into a fixed-length token; The style token Conditional characteristics of the diffusion model Combined to form the final diffusion model input: , in, represents the conditional features used as input to the diffusion model, Indicates that the style token and conditional features Stitching in a specific dimension.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a program for an unsupervised anomaly detection method based on a diffusion model. When the program for an unsupervised anomaly detection method based on a diffusion model is executed by a processor, the steps of an unsupervised anomaly detection method based on a diffusion model as described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Image generation method and system based on substation equipment defects
CN119516042A
Latent diffusion model autodecoders
US20240395028A1