Image anomaly detection method based on twin coding diffusion model and flow model
By combining the twin coded diffusion model, quantization module and flow model, an efficient and accurate model for image anomaly detection is built, solving the problems of missed detection and real-time requirements in the prior art.
Patent Information
- Application Number
- CN202510503289.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Existing image abnormality detection methods have problems of missed detection when identifying image abnormalities, and it is difficult to meet the real-time requirements of industrial industries.
An image anomaly detection method based on twin coded diffusion model and stream model is adopted. By combining twin coded diffusion model, quantization module and stream model, an image anomaly detection model with universality, high detection efficiency and low error detection rate is constructed.
It realizes efficient and accurate image abnormality detection, reduces the error detection rate, improves detection efficiency, and adapts to industrial real-time needs.
Smart Images

Figure CN120013949A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image detection, and specifically relates to an image anomaly detection method based on a twin coding diffusion model and a flow model. Background Art
[0002] Image anomaly detection is an important research direction in the field of computer vision. It aims to identify abnormal parts of an image that do not conform to the normal pattern. It has a wide range of applications in the industry. In the industrial field, image anomaly detection can be used to detect scratches, cracks, dents and other defects on the surface of products to ensure product quality; it can detect solder joint defects on circuit boards, component misalignment and other problems to ensure the reliability of electronic products. In medical images, abnormalities in brain CT images can be detected to help doctors assist in diagnosis.
[0003] The embedding-based image anomaly detection method uses a pre-trained model to extract image features and constructs interfaces or scoring rules in the feature space to perform anomaly detection. This method does not rely on the construction of additional negative samples and the design of proxy tasks. It mainly considers the differences in the feature space to perform anomaly detection in a high-dimensional feature space. However, this method relies on a pre-trained model, which leads to its lack of interpretability.
[0004] Reconstruction-based image anomaly detection methods can be divided into autoencoder-based, GAN-based, and diffusion model-based methods. The autoencoder-based and GAN-based methods are trained on normal images. When reconstructing abnormal images, the abnormal parts will be reconstructed into normal parts, so as to perform anomaly detection based on the reconstruction error. However, due to the generalization of neural networks, the autoencoder-based and GAN-based methods will also reconstruct the abnormal parts, resulting in false detection and missed detection.
[0005] The method based on diffusion model reconstruction has a training mechanism of adding noise to the image in multiple steps, learning to predict the added noise, and then mastering the process of restoring the image from noise. In anomaly detection reconstruction, this method can effectively reconstruct the anomaly as normal because it adds noise to the abnormal area. However, this mechanism will inevitably introduce noise to the normal area, and the subsequent multiple denoising steps may change the original normal image part. This type of method relies on reconstruction errors to identify anomalies. Any changes in the normal part may confuse the model's ability to distinguish between normal and abnormal, thereby causing false positives or negatives. In addition, the method based on diffusion model reconstruction is difficult to meet the real-time requirements of industrial anomaly detection because it needs to perform multiple denoising steps. Summary of the invention
[0006] In response to the above problems, the present invention proposes an image anomaly detection method based on the twin coding diffusion model and the flow model. Image anomaly detection is performed by combining the twin coding diffusion model, the quantization module and the flow model. The goal is to build an image anomaly detection model that is universal, has high detection efficiency, low false detection rate and is explainable.
[0007] The present invention is implemented by the following technical solution: an image anomaly detection method based on a twin coding diffusion model and a flow model, the steps are as follows: S1: Obtain image samples and perform preprocessing to obtain normal training samples and validation set samples; S2: After adding fixed-scale noise to the normal training sample, a noisy image and a known noise are obtained. The original normal image and the noisy image are sent to the twin coding diffusion model to obtain the predicted noise, and the noise prediction loss and the flow discrimination loss are defined; S3: Construct a first-stream model, extract features of predicted noise and known noise to obtain latent features and known features, use the first-stream model to estimate the probability distribution of latent features, and use the known features to regularize the first-stream model, and perform joint alternating training with the twin coding diffusion model; construct a quantization module, the quantization module obtains quantization features that match the input image, obtains quantization residual features based on the quantization features and constructs a second-stream model; use the second-stream model to estimate the probability distribution of quantization residual features, and train the first-stream model, the second-stream model, and the twin coding diffusion model; S4: After training is completed, the probability distribution of the latent features and the probability distribution of the quantized residual features are dynamically adjusted using the validation set to obtain the optimal dynamic weight parameters and integrated into the final anomaly score; S5: Use the trained twin encoding diffusion model, the first-stream model, and the second-stream model to perform anomaly detection on the image to be detected.
[0008] Further preferably, in step S5, after adding noise of a fixed scale to the image to be detected during the diffusion process, a noisy image and a known noise are obtained, the image to be detected and the noisy image are sent to the twin coding diffusion model to obtain predicted noise, and feature extraction is performed on the predicted noise and the known noise respectively to obtain potential features and known features respectively, and the probability distribution of the potential features is obtained through the first stream model; the quantized residual features of the image to be detected are obtained based on the quantization module, and the probability distribution of the quantized residual features is obtained through the second stream model; the probability distribution of the potential features and the probability distribution of the quantized residual features are integrated through the optimal dynamic weight parameters to obtain the anomaly score, the anomaly score is compared with the artificially set anomaly detection threshold to obtain a binary mask image, and whether the image to be detected is abnormal is judged according to the binary mask image.
[0009] Further preferably, the twin coding diffusion model includes three parts, the first part is the twin encoder, the second part is the spatial feature fusion module, and the third part is the intermediate layer and decoder; the twin encoder is composed of two encoders with the same structure, and the original normal image and the noisy image are input in parallel. The twin encoder performs feature interaction inside, and the first layer of features of the noisy image is added to the first layer of features of the original normal image. The last six layers of features of the original normal image are passed through the spatial feature fusion module and added to the features of the corresponding noisy image, and then jump-connected to the corresponding decoder. The intermediate layer serves as a bridge connecting the twin encoder and the decoder to further extract feature information; finally, the predicted noise is obtained through the decoder.
[0010] Further preferably, the spatial feature fusion module is composed of a 3×3 convolutional layer, an instance normalization layer and a SiLU activation layer.
[0011] More preferably, the noise prediction loss adopts L2 loss.
[0012] Further preferably, the total loss of the twin coding diffusion model is the sum of the flow discrimination loss and the noise prediction loss, and the flow discrimination loss is defined as the arithmetic mean sum of the discrimination losses of all coupling layers in the first flow model.
[0013] Further preferably, the process of constructing the quantization module is as follows: First, use the feature extractor pre-trained on the ImageNet dataset to extract local features of the original normal image using the following formula to obtain the local features of each original normal image and form a local feature library: ; in is the feature of the d-th original normal image, is the dth original normal image, is a pre-trained feature extractor; The feature vector is flattened as follows to construct the original patch feature of the image: ; in It represents the e-th feature vector in the features of the d-th original normal image, where e is the feature vector number of the original image feature, C represents the channel, H is the height, and W is the width. Represents a C-dimensional vector, and flatten is a vector flattening operation; The original patch features of all the original normal images are gathered together to form a feature set, and the K-means clustering algorithm is used to perform dictionary learning on the set to obtain m cluster centers: ; in, is the set of cluster centers, is the vth cluster center feature, is the serial number of the cluster center, and m is the number of cluster centers; The cluster center is used as the codebook, i.e., the visual dictionary; the input image is extracted using a pre-trained feature extractor, and the feature is divided into S×S block-level sub-vectors: ; Among them, S 2 is the number of blocks into which the feature is divided, is the divided feature block, u is the sequence number of the feature block; The block-level subvectors are flattened to obtain patch features, and then their block-level features are encoded in a quantized manner, that is, the distance from each point in the patch feature to the cluster center is calculated, and each point in the feature is assigned to its nearest cluster to obtain the codebook-based quantized features as the final block features: ; in, is the quantized block feature, is an m-dimensional vector, is the feature vector of the g-th block feature, g is the feature vector sequence number of the block feature, and quantify is the quantization operation; All sub-blocks are processed in the same way to obtain sub-block features, which are used to form image quantization features. The quantization features are residualized with the image features to obtain quantization residual features, and the probability distribution of the quantization residual features is estimated using the second stream model.
[0014] The present invention directly uses the flow model to estimate the probability distribution based on the prediction noise instead of adopting the prediction noise reconstruction method. This can not only avoid the requirement for the fidelity of the normal part of the image during reconstruction, but also improve the efficiency of anomaly detection by using only single-step prediction noise. At the same time, the twin encoder structure and spatial feature fusion module are used in the diffusion model, so that the model can better identify the detailed information of the normal part of the image, thereby improving the model's ability to distinguish between normal and abnormal.
[0015] The present invention uses the twin coding diffusion model to predict noise and trains the twin coding diffusion model in normal images. Different from the mainstream diffusion model training method, the present invention adds single-step Gaussian noise to normal images, and then uses the twin coding diffusion model to predict noise. The twin coding diffusion model is trained using L2 loss and flow discriminant loss. The probability distribution of noise is estimated using the flow model, and the estimation of the noise characteristics of abnormal images will deviate from the predicted distribution. Image anomaly detection is performed based on this difference.
[0016] Since the forward diffusion process introduces noise that destroys the image and will lose the original image detail information to a certain extent, the present invention uses a twin coding structure to input the information of the original normal image and the noisy image in parallel, and uses a spatial feature fusion module at the same time. During this process, the first-layer features of the original normal image will be input into the first-layer features of the noisy image, so that the original normal image information will not excessively affect the noise prediction, and the original normal image features and the noisy image features will be added to the decoder for feature fusion, which can better identify the detail information of the normal part of the image and improve the model's ability to distinguish between normal and abnormal. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flow chart of the present invention.
[0018] Figure 2 Schematic diagram of the image anomaly detection model of the present invention.
[0019] Figure 3 Schematic diagram of the overall architecture of the twin coding diffusion model.
[0020] Figure 4 Graph of the twin encoding diffusion model.
[0021] Figure 5 It is the spatial feature fusion module structure. DETAILED DESCRIPTION
[0022] The present invention will be further described in detail below with reference to the accompanying drawings.
[0023] Reference Figure 1 and Figure 2 , an image anomaly detection method based on twin coding diffusion model and flow model, the steps are as follows: S1: Obtain image samples and perform preprocessing to obtain normal training samples and validation set samples; S2: After adding fixed-scale noise to the normal training sample, a noisy image and a known noise are obtained. The original normal image and the noisy image are sent to the twin coding diffusion model to obtain the predicted noise, and the noise prediction loss and the flow discrimination loss are defined; S3: Construct a first-stream model, extract features of predicted noise and known noise to obtain latent features and known features, use the first-stream model to estimate the probability distribution of latent features, and use the known features to regularize the first-stream model, and perform joint alternating training with the twin coding diffusion model; construct a quantization module, the quantization module obtains quantization features that match the input image, obtains quantization residual features based on the quantization features and constructs a second-stream model, and uses the second-stream model to estimate the probability distribution of quantization residual features; train the first-stream model, the second-stream model, and the twin coding diffusion model; S4: After training is completed, the probability distribution of the latent features and the probability distribution of the quantized residual features are dynamically adjusted using the validation set to obtain the optimal dynamic weight parameters and integrated into the final anomaly score; S5: Use the trained twin coding diffusion model, the first-stream model and the second-stream model to perform anomaly detection on the image to be detected. After adding noise of a fixed scale to the image to be detected during the diffusion process, a noisy image and a known noise are obtained. The image to be detected and the noisy image are sent to the twin coding diffusion model to obtain the predicted noise. The predicted noise and the known noise are respectively subjected to feature extraction to obtain potential features and known features. The probability distribution of the potential features is obtained through the first-stream model. The quantized residual features of the image to be detected are obtained based on the quantization module, and the probability distribution of the quantized residual features is obtained through the second-stream model. The probability distribution of the potential features and the probability distribution of the quantized residual features are integrated through the optimal dynamic weight parameter to obtain the anomaly score. The anomaly score is compared with the artificially set anomaly detection threshold to obtain a binary mask map. The image to be detected is judged to be abnormal based on the binary mask map. Among them, the pixels above the anomaly detection threshold are marked as abnormal, and the pixels below the anomaly detection threshold are marked as normal. If all pixels of the image are marked as normal, the image to be detected is judged to be a normal image, otherwise it is judged to be an abnormal image.
[0024] In step S1 of this embodiment, an image sample is obtained, the image is transformed into a 256×256 image size, the image pixel value is divided by 255 and limited to the range of [0,1], the input data of the model is unified so that the model can better process the input data, and the image is divided into normal training samples and validation set samples in a ratio of 7:3, where the normal training samples only contain normal image samples, and the validation set contains normal image samples and abnormal image samples.
[0025] like Figure 3 and Figure 4As shown, the twin coding diffusion model of this embodiment includes three parts, the first part is the twin encoder, the second part is the spatial feature fusion module, and the third part is the intermediate layer and decoder. The twin encoder consists of two encoders with the same structure, and the original normal image and the noisy image are input in parallel to ensure that the detail information of the original image is not lost in the encoding stage. At the same time, the spatial feature fusion module is used to integrate the high-level semantic information into the underlying semantic information, so as to better allow the network to recognize the detail information of the normal part of the image. The above structure will improve the ability of the model to distinguish between normal and abnormal. The twin encoder performs feature interaction inside, adds the first layer of features of the noisy image to the first layer of features of the original normal image, passes the last six layers of features of the original normal image through the spatial feature fusion module and adds them to the features of the corresponding noisy image, jumps to the corresponding decoder, and the intermediate layer serves as a bridge connecting the twin encoder and the decoder to further extract feature information; finally, the predicted noise is obtained through the decoder. The noise prediction loss is defined, that is, random Gaussian noise is generated to destroy the original normal image, the noisy image and the original normal image are input into the twin coding diffusion model to obtain the predicted noise, and the predicted noise is calculated with the known noise. L2 loss. Define the stream discriminant loss, add the discriminant loss of each coupling layer in the first stream model to form the stream discriminant loss. Add the noise prediction loss and the stream discriminant loss to form the loss of the twin coding diffusion model.
[0026] Spatial feature fusion module such as Figure 5 As shown in Figure 1, it consists of a 3×3 convolutional layer, an instance normalization layer, and a SiLU activation layer. The instance normalization layer can retain the information of individual images, and the SiLU activation layer can save more input information. The spatial feature fusion module integrates the features of the first three layers of the original normal image encoding features into the features of the last three layers, as shown in Figure 1. Figure 5 The first three layers of features are integrated through convolution blocks and added to the fourth layer of features, and the final features are obtained through the SiLU activation layer. The fusion process of the fifth and sixth layers of features is similar. The fused features are added to the corresponding encoding features of the noisy image and then jump-connected to the corresponding decoder.
[0027] Use Gaussian noise of fixed scale to add noise to the original normal image to obtain the noisy image. The formula is as follows: ; in is the multiplication of coefficients at each time step from 0 to t, where t is a preset fixed time step. is a randomly sampled Gaussian noise, represents a Gaussian distribution, is the original normal image, Represents the noisy image after adding noise to the original normal image.
[0028] The noise prediction loss uses L2 loss: ; in, is the noise prediction loss, is the average of the squared differences between the predicted noise and the known noise, is a known noise, represents the prediction noise.
[0029] When estimating an image, the diffusion model has difficulty balancing the fidelity of reconstructing the normal part of the image and the reconstruction requirements of the abnormal part, making it unable to distinguish between normal and abnormal, resulting in false detection and missed detection. In addition, the estimation process requires multiple denoising steps, which takes a long time and is not conducive to industrial deployment. The present invention only uses the twin coding diffusion model to transfer the image to the noise feature for fixed-scale noise, and uses the flow model to estimate the noise distribution for anomaly detection, avoiding the tediousness of multi-step denoising, improving the efficiency of anomaly detection, and being more adaptable to industrial deployment and application.
[0030] The diffusion model-based method will change the details of the original normal image. Since the original normal image is corrupted by noise during input, even if the prediction is only through single-step diffusion, it may be difficult for the diffusion model to identify the normal part of the image when predicting noise. By using the twin encoding diffusion model to input the original normal image and the noisy image in parallel, the network will not lose the feature information of the original image. At the same time, the spatial feature fusion module is added to retain the low-level semantic information of the original normal image, thereby improving the model's ability to distinguish between normal and abnormal.
[0031] The coupling layer discriminant loss for constructing the first-rate model consists of the mean square error and is defined as: ; in, is the mean square error of the i-th coupling layer of the first-stream model, n is the number of samples, is the output of the i-th coupling layer of the first-stream model for the j-th sample, are the corresponding known features.
[0032] Flow discrimination loss It is defined as the arithmetic mean sum of the discriminative losses of all coupled layers in the first-stream model: ; Where k is the number of coupling layers in the first flow model.
[0033] The total loss of the final twin encoding diffusion model It can be expressed as: .
[0034] The process of building a quantization module is as follows: First, use the feature extractor pre-trained on the ImageNet dataset to extract local features of the original normal image using the following formula to obtain the local features of each original normal image and form a local feature library: ; in is the feature of the d-th original normal image, is the dth original normal image, is a pre-trained feature extractor.
[0035] The feature vector is flattened as follows to construct the original patch feature of the image: ; in It represents the e-th feature vector in the features of the d-th original normal image, where e is the feature vector number of the original image feature, C represents the channel, H is the height, and W is the width. Represents a C-dimensional vector, and flatten is a vector flattening operation.
[0036] The original patch features of all the original normal images are gathered together to form a feature set, and the K-means clustering algorithm is used to perform dictionary learning on the set to obtain m cluster centers: ; in, is the set of cluster centers, is the vth cluster center feature, is the serial number of the cluster center, and m is the number of cluster centers.
[0037] The cluster centers are used as codebooks, i.e., visual dictionaries.
[0038] The input image is extracted using a pre-trained feature extractor and the feature is divided into S×S block-level sub-vectors: ; Among them, S 2 is the number of blocks into which the feature is divided, is the divided feature block, and u is the sequence number of the feature block.
[0039] The block-level subvectors are flattened to obtain patch features, and then their block-level features are encoded in a quantized manner, that is, the distance from each point in the patch feature to the cluster center is calculated, and each point in the feature is assigned to its nearest cluster to obtain the codebook-based quantized features as the final block features: ; in, is the quantized block feature, is an m-dimensional vector, is the feature vector of the g-th block feature, g is the feature vector sequence number of the block feature, and quantify is the quantization operation.
[0040] All sub-blocks are processed in the same way to obtain sub-block features, which are used to form image quantization features. The quantization features are residualized with the image features to obtain quantization residual features, and the probability distribution of the quantization residual features is estimated using the second stream model.
[0041] A pre-trained network is used to process the predicted noise and known noise output by the diffusion process to obtain latent features and known features, a first-stream model is used to estimate the latent features, and the known features are used to regularize the first-stream model.
[0042] Build the flow model: The flow model can obtain a bidirectional mapping relationship from target distribution to noise distribution and from noise distribution to target distribution. It is composed of multiple layers of reversible coupling layers, and the reversible coupling layer structure uses the coupling layer structure of RealNVP.
[0043] The reversible coupling layer learns the bidirectional mapping relationship between the latent features and the intermediate features. The bidirectional mapping relationship between the latent features and the Gaussian features is learned by stacking multiple coupling layers. The stacking of coupling layers can be expressed as: ; Among them, H i is the intermediate feature output by the i-th coupling layer, f i is the i-th coupling layer, i={1,2,3,…,k}, X is the latent feature, and Z is the Gaussian feature.
[0044] The final flow model can be expressed as , where F is the mapping of estimated latent features to Gaussian features, is a stack of coupled layers. The probability estimation loss L of the flow model F It can be expressed as: ; in, is the standard Gaussian distribution probability function, is the mapping operation of the input potential feature element x to the Gaussian feature, yes The Jacobian matrix of the reversible coupling layer, log is the logarithm, It is to find the absolute value of the determinant of the matrix.
[0045] Use known features to calculate the flow discrimination loss and regularize the first-flow model. Finally, the total loss of the first-flow model is It can be expressed as: ; The loss of the second-stream model It is composed of the probability estimation loss of the flow model: .
[0046] The twin coding diffusion model and the first-stream model are trained alternately to optimize the two models together. The alternating training period q is set. When the number of training rounds is 2pq to (2p+1)q, p is a non-negative integer, the parameters of the first-stream model are fixed, and training is performed using the loss function of the twin coding diffusion model. When the number of training rounds is (2p+1)q to (2p+2)q, p is a non-negative integer, the parameters of the twin coding diffusion model are fixed, and training is performed using the loss function of the first-stream model. The alternating process is independent of the second-stream model, and the second-stream model is trained simultaneously during the alternating training until the training is completed.
[0047] After training is completed, the validation set is used to dynamically adjust the abnormal proportion of potential features and quantized residual features and integrate them into the final abnormal score. The steps are as follows: (1) Normalization: and First, the maximum and minimum values are normalized to the range of 0 to 1, and then bilinear interpolation is used to adjust to the image size. ; ; in is the probability distribution of the latent features, calculated from the loss of the first-rate model, is the probability distribution of the quantized residual features, calculated by the loss of the second-stream model. Q is a bilinear interpolation operation that resizes the probability distribution of the feature dimension to the image size. is the probability distribution of the normalized latent features, is the probability distribution of the normalized quantized residual feature, min is the minimum value operation, and max is the maximum value operation. (2) Verification set search for optimal weights: Traverse the candidate values of r on the validation set (such as r∈{0.1,0.2,...,0.9}) and select the weight coefficient r that maximizes AUC-ROC ∗ : ; Among them, r is the dynamic weight coefficient, ranging from 0 to 1. AUC-ROC is an indicator used to evaluate the performance of anomaly detection models. It measures the ability of the model to distinguish between normal and abnormal under different thresholds. argmax refers to the value of the variable when the formula reaches the maximum value.
[0048] (3) Final anomaly score S final The calculation formula is: ; The twin encoding diffusion model and the flow model are jointly trained alternately. The image features are first transferred to the noise space, and then the noise space is transferred to the latent space, so as to accurately estimate the probability distribution of normal images. The twin encoding diffusion model is trained by combining L2 loss and flow discriminant loss to better fit the prediction noise. In the process of flow model training, the coupling layer is regularized using known Gaussian noise data, which is conducive to the flow model learning more meaningful feature representation.
[0049] The block feature clustering method quantizes the image to be detected and estimates the quantized residual features of the input image and the quantized image. The quantized residual features are dynamically combined with the abnormality discrimination of the latent features to improve the robustness of anomaly detection.
[0050] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0051] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. An image anomaly detection method based on twin coding diffusion model and flow model, characterized in that: Here are the steps: S1: Obtain image samples and perform preprocessing to obtain normal training samples and validation set samples; S2: After adding fixed-scale noise to the normal training sample, a noisy image and a known noise are obtained. The original normal image and the noisy image are sent to the twin coding diffusion model to obtain the predicted noise, and the noise prediction loss and the flow discrimination loss are defined; S3: Build a first-stream model, extract features from predicted noise and known noise to obtain latent features and known features, use the first-stream model to estimate the probability distribution of latent features, and use known features to regularize the first-stream model, and perform joint alternating training with the twin coding diffusion model; build a quantization module, the quantization module obtains quantization features that match the input image, obtains quantization residual features based on the quantization features, and builds a second-stream model; The probability distribution of quantized residual features is estimated using the second-stream model, and the first-stream model, the second-stream model and the twin coding diffusion model are trained; S4: After training is completed, the probability distribution of the latent features and the probability distribution of the quantized residual features are dynamically adjusted using the validation set to obtain the optimal dynamic weight parameters and integrated into the final anomaly score; S5: Use the trained twin encoding diffusion model, the first-stream model, and the second-stream model to perform anomaly detection on the image to be detected.
2. The image anomaly detection method according to claim 1, characterized in that: In step S5, after adding noise of a fixed scale to the image to be detected during the diffusion process, a noisy image and a known noise are obtained, and the image to be detected and the noisy image are sent to the twin coding diffusion model to obtain predicted noise. Feature extraction is performed on the predicted noise and the known noise respectively to obtain potential features and known features respectively, and the probability distribution of the potential features is obtained through the first-stream model; the quantized residual features of the image to be detected are obtained based on the quantization module, and the probability distribution of the quantized residual features is obtained through the second-stream model; the probability distribution of the potential features and the probability distribution of the quantized residual features are integrated through the optimal dynamic weight parameters to obtain the anomaly score, and the anomaly score is compared with the artificially set anomaly detection threshold to obtain a binary mask image, and whether the image to be detected is abnormal is judged according to the binary mask image.
3. The image anomaly detection method according to claim 1, characterized in that: The twin coding diffusion model includes three parts. The first part is the twin encoder, the second part is the spatial feature fusion module, and the third part is the intermediate layer and decoder. The twin encoder is composed of two encoders with the same structure. The original normal image and the noisy image are input in parallel. The twin encoder performs feature interaction inside, and the first layer of features of the noisy image is added to the first layer of features of the original normal image. The last six layers of features of the original normal image are passed through the spatial feature fusion module and added to the features of the corresponding noisy image, and then jump-connected to the corresponding decoder. The intermediate layer serves as a bridge connecting the twin encoder and the decoder to further extract feature information. Finally, the predicted noise is obtained through the decoder.
4. The image anomaly detection method according to claim 3, characterized in that: The spatial feature fusion module is composed of a 3×3 convolutional layer, an instance normalization layer and a SiLU activation layer.
5. The image anomaly detection method according to claim 3, characterized in that: The noise prediction loss adopts L2 loss.
6. The image anomaly detection method according to claim 3, characterized in that: The total loss of the twin coding diffusion model is the sum of the stream discriminant loss and the noise prediction loss. The stream discriminant loss is defined as the arithmetic mean sum of the discriminant losses of all coupled layers in the first stream model.
7. The image anomaly detection method according to claim 1, characterized in that: The process of building a quantization module is as follows: First, use the feature extractor pre-trained on the ImageNet dataset to extract local features of the original normal image using the following formula to obtain the local features of each original normal image and form a local feature library: ; in is the feature of the d-th original normal image, is the dth original normal image, is a pre-trained feature extractor; The feature vector is flattened as follows to construct the original patch feature of the image: ; in It represents the e-th feature vector in the features of the d-th original normal image, where e is the feature vector number of the original image feature, C represents the channel, H is the height, and W is the width. Represents a C-dimensional vector, and flatten is a vector flattening operation; The original patch features of all the original normal images are gathered together to form a feature set, and the K-means clustering algorithm is used to perform dictionary learning on the set to obtain m cluster centers: ; in, is the set of cluster centers, is the vth cluster center feature, is the serial number of the cluster center, and m is the number of cluster centers; The cluster center is used as the codebook, i.e., the visual dictionary; the input image is extracted using a pre-trained feature extractor, and the feature is divided into S×S block-level sub-vectors: ; Among them, S 2 is the number of blocks into which the feature is divided, is the divided feature block, u is the sequence number of the feature block; The block-level subvectors are flattened to obtain patch features, and then their block-level features are encoded in a quantized manner, that is, the distance from each point in the patch feature to the cluster center is calculated, and each point in the feature is assigned to its nearest cluster to obtain the codebook-based quantized features as the final block features: ; in, is the quantized block feature, is an m-dimensional vector, is the feature vector of the g-th block feature, g is the feature vector sequence number of the block feature, and quantify is the quantization operation; All sub-blocks are processed in the same way to obtain sub-block features, which are used to form image quantization features. The quantization features are residualized with the image features to obtain quantization residual features, and the probability distribution of the quantization residual features is estimated using the second stream model.
Citation Information
Patent Citations
Abnormity detection method based on twin auto-encoders and bidirectional information depth supervision
CN116645369A
Remote sensing image building extraction method and device based on diffusion model
CN117372873A
Industrial internet of things digital twin modeling method and system based on diffusion model
CN118296944A
Multi-class typical foreign matter detection method based on diffusion model
CN118485641A
Denoising diffusion generative adversarial networks
US20230095092A1