An Image Anomaly Detection Method Based on Siamese Encoder Diffusion Model and Flow Model
By combining the twin coded diffusion model, quantization module and flow model, the error detection problem in the existing image anomaly detection methods is solved, and efficient image anomaly detection with low error detection rate is achieved, which is suitable for real-time industrial applications.
Patent Information
- Application Number
- CN202510503289.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The existing image abnormality detection methods have the problem of missed detection when detecting abnormalities, and the method based on diffusion model is difficult to meet the real-time requirements of industrial industries.
An image anomaly detection method based on twin coded diffusion model and stream model is adopted. By combining twin coded diffusion model, quantization module and stream model, an image anomaly detection model with universality, high detection efficiency and low error detection rate is constructed.
It realizes efficient and low error detection rate image abnormality detection, which can better identify the detailed information of the normal part of the image, improves the model's ability to distinguish between normal and abnormal, and is suitable for industrial real-time applications.
Smart Images

Figure CN120013949B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image detection, and particularly relates to an image anomaly detection method based on a twin-coded diffusion model and a flow model. Background Art
[0002] Image anomaly detection is an important research direction in the field of computer vision, aiming to identify abnormal parts in images that do not conform to normal patterns, and has a wide range of applications in the industrial field. In the industrial field, image anomaly detection can be used to detect defects such as scratches, cracks, and depressions on the surface of products to ensure product quality; it can detect problems such as solder joint defects and component misalignment on circuit boards to ensure the reliability of electronic products. In medical images, it can detect abnormal conditions in brain CT images to assist doctors in diagnosis.
[0003] The image anomaly detection method based on embedding uses a pre-trained model to extract image features, and constructs a separation surface or scoring rule in the feature space for anomaly detection. This method does not rely on the construction of additional negative samples and the design of proxy tasks, mainly considering the differences in the feature space to perform anomaly detection in the high-dimensional feature space. However, this method depends on the pre-trained model, resulting in a lack of interpretability.
[0004] The image anomaly detection method based on reconstruction can be divided into methods based on autoencoders, GANs, and diffusion models. The methods based on autoencoders and GANs are trained on normal images. When reconstructing abnormal images, the abnormal parts will be reconstructed into normal parts, and then anomaly detection is performed according to the reconstruction error. However, due to the generalization of neural networks, the methods based on autoencoders and GANs will also reconstruct the abnormal parts, resulting in false detection and missed detection.
[0005] The method based on diffusion model reconstruction has a training mechanism that adds noise to the image in multiple steps and learns to predict these added noises, thereby mastering the process of restoring the image from the noise. In the anomaly detection reconstruction of this method, because noise is added to the abnormal area, it can effectively promote the reconstruction of the abnormal into the normal. However, this mechanism will also inevitably introduce noise to the normal area, and there are multiple subsequent denoising steps, which may change the originally normal image part. Such methods rely on reconstruction error to identify anomalies, and any change in the normal part may confuse the model's ability to distinguish between normal and abnormal, thereby causing false alarms or missed detections. In addition, the method based on diffusion model reconstruction is difficult to meet the real-time requirements of industrial anomaly detection due to the need to perform multiple denoising steps. Summary of the Invention
[0006] In view of the above problems, the present invention proposes an image anomaly detection method based on a twin-coded diffusion model and a flow model. By combining the twin-coded diffusion model, a quantization module, and a flow model for image anomaly detection, the goal is to construct a general, highly efficient, low false-alarm-rate, and interpretable image anomaly detection model.
[0007] The present invention is realized through the following technical solutions: An image anomaly detection method based on a twin-coded diffusion model and a flow model, the steps are as follows:
[0008] S1: Obtain image samples and perform preprocessing to obtain normal training samples and validation set samples;
[0009] S2: After adding noise of a fixed scale to the normal training samples, obtain the noisy image and the known noise. Send the original normal image and the noisy image into the twin-coded diffusion model to obtain the predicted noise, and define the noise prediction loss and the flow discriminant loss;
[0010] S3: Construct a first flow model to extract features from the predicted noise and the known noise to obtain latent features and known features. Use the first flow model to estimate the probability distribution of the latent features, and use the known features to regularize the first flow model, and perform joint alternating training with the twin-coded diffusion model; Construct a quantization module. The quantization module obtains quantization features matching the input image, obtains quantization residual features based on the quantization features, and constructs a second flow model; Use the second flow model to estimate the probability distribution of the quantization residual features, and train the first flow model, the second flow model, and the twin-coded diffusion model;
[0011] S4: After training is completed, dynamically adjust the probability distributions of the latent features and the quantization residual features using the validation set to obtain the optimal dynamic weight parameters, and integrate them into the final anomaly score;
[0012] S5: Use the trained twin-coded diffusion model, the first flow model, and the second flow model to perform anomaly detection on the image to be detected.
[0013] Further preferably, in step S5, after adding noise of a fixed scale to the image to be detected during the diffusion process, obtain the noisy image and the known noise. Send the image to be detected and the noisy image into the twin-coded diffusion model to obtain the predicted noise, perform feature extraction on the predicted noise and the known noise respectively to obtain latent features and known features respectively, and obtain the probability distribution of the latent features through the first flow model; Based on the quantization module, obtain the quantization residual features of the image to be detected, and obtain the probability distribution of the quantization residual features through the second flow model; Integrate the probability distributions of the latent features and the quantization residual features through the optimal dynamic weight parameters to obtain the anomaly score, compare the anomaly score with the artificially set anomaly detection threshold to obtain a binary mask image, and judge whether the image to be detected is abnormal according to the binary mask image.
[0014] Further preferably, the twin-coded diffusion model includes three parts. The first part is the twin encoder, the second part is the spatial feature fusion module, and the third part is the middle layer and the decoder. The twin encoder consists of two encoders with the same structure, which parallelly input the original normal image and the noisy image. Feature interaction occurs inside the twin encoder, adding the first-layer feature of the noisy image to the first-layer feature of the original normal image, and adding the last six-layer features of the original normal image after passing through the spatial feature fusion module to the corresponding features of the noisy image, and then making a skip connection to the corresponding decoder. The middle layer serves as a bridge connecting the twin encoder and the decoder, for further extracting feature information. Finally, the predicted noise is obtained through the decoder.
[0015] Further preferably, the spatial feature fusion module is composed of a 3×3 convolutional layer, an instance normalization layer, and a SiLU activation layer.
[0016] Further preferably, the noise prediction loss uses the L2 loss.
[0017] Further preferably, the total loss of the twin-coded diffusion model is the sum of the flow discriminant loss and the noise prediction loss. The flow discriminant loss is defined as the arithmetic mean of the discriminant losses of all coupling layers in the first flow model.
[0018] Further preferably, the process of constructing the quantization module is as follows:
[0019] First, use the feature extractor pre-trained on the ImageNet dataset to perform local feature extraction on the original normal image using the following formula to obtain the local features of each original normal image, forming a local feature library:
[0020] ;
[0021] where is the feature of the d-th original normal image, is the d-th original normal image, is the pre-trained feature extractor;
[0022] Flatten the feature vector using the following formula to construct the original patch feature of the image:
[0023] ;
[0024] where represents the e-th feature vector in the feature of the d-th original normal image, e is the serial number of the feature vector of the original image feature, C represents the channel, H is the height, W is the width, represents a C-dimensional vector, and flatten is the vector flattening operation;
[0025] Aggregate the original patch features of all original normal images to form a feature set, and use the K-means clustering algorithm on the set for dictionary learning to obtain m clustering centers:
[0026] ;
[0027] Among them, is the set of clustering centers, is the feature of the v-th clustering center, is the serial number of the clustering center, and m is the number of clustering centers;
[0028] Use the clustering centers as the codebook, that is, the visual dictionary; extract features from the input image using a pre-trained feature extractor, and divide the feature into block-level sub-vectors of S×S:
[0029] ;
[0030] Among them, S 2 is the number of blocks for feature division, is the divided feature block, and u is the serial number of the feature block;
[0031] Flatten the block-level sub-vectors to obtain patch features, and then encode their block-level features in a quantization manner, that is, calculate the distance from each point in the patch feature to the clustering center, and assign each point in the feature to its nearest cluster to obtain the quantization feature based on the codebook as the final block feature:
[0032] ;
[0033] Among them, is the quantized block feature, is an m-dimensional vector, is the feature vector of the g-th block feature, g is the serial number of the feature vector of the block feature, and quantify is the quantization operation;
[0034] Perform the same processing on all sub-blocks to obtain sub-block features, form the image quantization feature from the sub-block features, obtain the quantization residual feature by taking the residual between the quantization feature and the image feature, and use the second flow model to estimate the probability distribution of the quantization residual feature.
[0035] The present invention directly uses a flow model for probability distribution estimation based on predicted noise, rather than adopting the method of reconstructing the predicted noise. It can not only avoid the requirement for the fidelity of the normal part of the image during reconstruction, but also improve the anomaly detection efficiency by only using single-step predicted noise. At the same time, a twin encoder structure and a spatial feature fusion module are used in the diffusion model, enabling the model to better identify the detailed information of the normal part of the image, thereby improving the model's ability to distinguish between normal and abnormal.
[0036] The present invention uses a twin-encoded diffusion model to predict noise. The twin-encoded diffusion model is trained on normal images. Different from the mainstream diffusion model training method, the present invention adds single-step Gaussian noise to normal images, then uses the twin-encoded diffusion model to predict the noise, and uses the L2 loss and the flow discriminant loss to train the twin-encoded diffusion model. The flow model is used to estimate the probability distribution of the noise. The estimation of the noise characteristics of abnormal images will deviate from the predicted distribution, and image anomaly detection is performed based on this difference.
[0037] Since the forward diffusion process introduces noise and damages the image, some detail information of the original image will be lost to a certain extent. The present invention uses a twin-encoded structure to parallelly input the information of the original normal image and the noise-added image. At the same time, a spatial feature fusion module is used. In this process, the first-layer features of the original normal image will be input into the first-layer features of the noise-added image, so that the information of the original normal image will not overly affect the prediction of the noise. Both the original normal image features and the noise-added image features will be added to the decoder for feature fusion, which can better identify the detail information of the normal part of the image and improve the ability of the model to distinguish normal from abnormal. Brief Description of the Drawings
[0038] Figure 1 It is a flowchart of the present invention.
[0039] Figure 2 It is a schematic diagram of the image anomaly detection model of the present invention.
[0040] Figure 3 It is a schematic diagram of the overall architecture of the twin-encoded diffusion model.
[0041] Figure 4 It is a diagram of the twin-encoded diffusion model.
[0042] Figure 5 It is the structure of the spatial feature fusion module. Detailed Embodiment
[0043] The present invention will be further described in detail below with reference to the accompanying drawings.
[0044] Refer to Figure 1 and Figure 2 , an image anomaly detection method based on a twin-encoded diffusion model and a flow model, the steps are as follows:
[0045] S1: Obtain image samples and perform preprocessing to obtain normal training samples and validation set samples;
[0046] S2: After adding noise with a fixed scale to the normal training samples, obtain the noise-added images and the known noise, send the original normal images and the noise-added images into the twin-encoded diffusion model to obtain the predicted noise, and define the noise prediction loss function and the flow discriminant loss;
[0047] S3: Construct a first - class model to extract features from the predicted noise and known noise, obtaining latent features and known features. Use the first - class model to estimate the probability distribution of the latent features, and use the known features to regularize the first - class model, and perform joint alternating training with the twin - coded diffusion model; construct a quantization module. The quantization module obtains quantization features matching the input image, obtains quantization residual features based on the quantization features and constructs a second - class model, and uses the second - class model to estimate the probability distribution of the quantization residual features; train the first - class model, the second - class model and the twin - coded diffusion model;
[0048] S4: After training is completed, dynamically adjust the probability distribution of the latent features and the probability distribution of the quantization residual features using the validation set to obtain the optimal dynamic weight parameters, and integrate them into the final anomaly score;
[0049] S5: Use the trained twin - coded diffusion model, the first - class model and the second - class model to perform anomaly detection on the image to be detected. After adding noise with a fixed scale to the image to be detected during the diffusion process, a noisy image and known noise are obtained. Send the image to be detected and the noisy image into the twin - coded diffusion model to obtain the predicted noise. Extract features from the predicted noise and the known noise respectively, obtaining latent features and known features respectively, and obtain the probability distribution of the latent features through the first - class model; obtain the quantization residual features of the image to be detected based on the quantization module, and obtain the probability distribution of the quantization residual features through the second - class model; integrate the probability distribution of the latent features and the probability distribution of the quantization residual features through the optimal dynamic weight parameters to obtain the anomaly score. Compare the anomaly score with the artificially set anomaly detection threshold to obtain a binary mask image, and judge whether the image to be detected is abnormal according to the binary mask image. Among them, the pixel points higher than the anomaly detection threshold are marked as abnormal, and the pixel points lower than the anomaly detection threshold are marked as normal. If all pixel points of the image are marked as normal, then judge that the image to be detected is a normal image, otherwise judge it as an abnormal image.
[0050] In step S1 of this embodiment, image samples are obtained, the shape of the image is transformed into an image size of 256×256, the pixel values of the image are divided by 255, and they are restricted within the range of [0, 1] to unify the input data of the model, enabling the model to better process the input data. The image is divided into normal training samples and validation set samples according to a ratio of 7:3, where the normal training samples only contain normal image samples, and the validation set contains normal image samples and abnormal image samples.
[0051] As Figure 3 and Figure 4As shown in the figure, the twin-coded diffusion model of this embodiment includes three parts. The first part is the twin encoder, the second part is the spatial feature fusion module, and the third part is the intermediate layer and the decoder. The twin encoder consists of two encoders with the same structure, which parallelly input the original normal image and the noisy image to ensure that the details of the original image are not lost during the encoding stage. At the same time, the spatial feature fusion module integrates high-level semantic information into low-level semantic information to better enable the network to recognize the details of the normal part of the image. The above structure will improve the model's ability to distinguish normal from abnormal. Feature interaction occurs inside the twin encoder, adding the first layer feature of the noisy image to the first layer feature of the original normal image, passing the last six layer features of the original normal image through the spatial feature fusion module and adding them to the corresponding features of the noisy image, and then making a skip connection to the corresponding decoder. The intermediate layer serves as a bridge connecting the twin encoder and the decoder for further extracting feature information; finally, the predicted noise is obtained through the decoder. Define the noise prediction loss, that is, generate random Gaussian noise to damage the original normal image, input the noisy image and the original normal image into the twin-coded diffusion model to obtain the predicted noise, and calculate the L2 loss between the predicted noise and the known noise. Define the flow discriminator loss, which is formed by adding the discriminator losses of each coupling layer in the first flow model. Add the noise prediction loss and the flow discriminator loss to form the loss of the twin-coded diffusion model.
[0052] The spatial feature fusion module is as Figure 5 shown, and is composed of a 3×3 convolutional layer, an instance normalization layer, and a SiLU activation layer. The instance normalization layer can retain the information of individual images, and the SiLU activation layer can preserve more input information. The spatial feature fusion module integrates the features of the first three layers among the last six layers of the encoded features of the original normal image into the features of the last three layers. As Figure 5 shown, the features of the first three layers are integrated through a convolutional block and added to the fourth layer feature, and the final feature is obtained through the SiLU activation layer. The fusion process of the fifth and sixth layer features is similar. The fused features are added to the corresponding encoded features of the noisy image and then make a skip connection to the corresponding decoder.
[0053] Add noise to the original normal image using Gaussian noise with a fixed scale to obtain the noisy image. The formula is as follows:
[0054] ;
[0055] where is the product of the coefficients at each time step from 0 to t, t is a preset fixed time step, is randomly sampled Gaussian noise, represents the Gaussian distribution, is the original normal image, represents the noisy image after adding noise to the original normal image.
[0056] The noise prediction loss uses the L2 loss:
[0057] ;
[0058] Wherein, is the noise prediction loss, is the average value of calculating the square of the difference between the predicted noise and the known noise, is the known noise, represents the predicted noise.
[0059] When the diffusion model estimates an image, it is difficult to balance the fidelity of reconstructing the normal part of the image and the reconstruction requirements of the abnormal part, resulting in its inability to distinguish between normal and abnormal, leading to false negatives and false positives. Moreover, the estimation process requires multiple denoising steps, with a long time overhead, which is not conducive to industrial deployment. In the present invention, only for noise of a fixed scale, the twin-encoding diffusion model is used to transfer the image to the noise features, and the flow model is used to estimate the noise distribution for anomaly detection, avoiding the cumbersome multi-step denoising, improving the efficiency of anomaly detection, and being more adaptable to industrial deployment and application.
[0060] The method based on the diffusion model will change the detailed information of the original normal image. Since the original normal image is damaged by noise during input, even the prediction through a single-step diffusion may make it difficult for the diffusion model to recognize the normal part of the image. By using the twin-encoding diffusion model to input the original normal image and the noisy image in parallel, the network will not lose the feature information of the original image. At the same time, a spatial feature fusion module is added to retain the low-level semantic information of the original normal image, thereby improving the model's ability to distinguish between normal and abnormal.
[0061] Construct the coupling layer discrimination loss of the first flow model, which is composed of the mean square error and is defined as:
[0062] ;
[0063] Wherein, is the mean square error of the i-th coupling layer of the first flow model, n is the number of samples, is the output of the i-th coupling layer of the first flow model for the j-th sample, is the corresponding known feature.
[0064] Flow discrimination loss is defined as the arithmetic mean sum of the discrimination losses of all coupling layers in the first flow model:
[0065] ;
[0066] Wherein, k is the number of coupling layers in the first flow model.
[0067] Total loss of the final twin coding diffusion model can be expressed as:
[0068] .
[0069] The process of constructing the quantization module is as follows:
[0070] First, use the feature extractor pre-trained on the ImageNet dataset to perform local feature extraction on the original normal images using the following formula to obtain the local features of each original normal image, forming a local feature library:
[0071] ;
[0072] where is the feature of the d-th original normal image, is the d-th original normal image, is the pre-trained feature extractor.
[0073] Flatten the feature vector using the following formula to construct the original patch features of the image:
[0074] ;
[0075] where represents the e-th feature vector in the feature of the d-th original normal image, e is the serial number of the feature vector of the original image feature, C represents the channel, H is the height, and W is the width, represents a C-dimensional vector, and flatten is the vector flattening operation.
[0076] Aggregate the original patch features of all original normal images to form a feature set, and use the K-means clustering algorithm for dictionary learning on the set to obtain m clustering centers:
[0077] ;
[0078] where, is the set of clustering centers, is the feature of the v-th clustering center, is the serial number of the clustering center, and m is the number of clustering centers.
[0079] Use the clustering centers as the codebook, that is, the visual dictionary.
[0080] Extract features from the input image using the pre-trained feature extractor, and divide the feature into S×S block-level sub-vectors:
[0081] ;
[0082] where S2 The number of blocks divided by features is the divided feature block, and u is the serial number of the feature block.
[0083] Flatten the block-level sub-vectors to obtain patch features, and then encode their block-level features in a quantization manner, that is, calculate the distance from each point in the patch features to the cluster center, assign each point in the features to its nearest cluster, and obtain the quantization features based on the codebook as the final block features:
[0084] ;
[0085] Among them, is the quantized block feature, is an m-dimensional vector, is the feature vector of the g-th block feature, g is the serial number of the feature vector of the block feature, and quantify is the quantization operation.
[0086] Perform the same processing on all sub-blocks to obtain sub-block features. The sub-block features form the image quantization features. The quantization features and the image features are used to calculate the residual to obtain the quantization residual features, and the probability distribution of the quantization residual features is estimated using the second flow model.
[0087] Use the pre-trained network to process the predicted noise and the known noise output by the diffusion process to obtain the latent features and the known features. Use the first flow model to estimate the latent features and use the known features to regularize the first flow model.
[0088] Construct the flow model:
[0089] The flow model can obtain the bidirectional mapping relationship between the target distribution and the noise distribution, and between the noise distribution and the target distribution. It consists of multiple layers of invertible coupling layers, and the structure of the invertible coupling layer uses the coupling layer structure of RealNVP.
[0090] The invertible coupling layer learns the bidirectional mapping relationship between the latent features and the intermediate features. By stacking multiple coupling layers, the bidirectional mapping relationship between the latent features and the Gaussian features is learned. The stacking of the coupling layers can be expressed as:
[0091] ;
[0092] where H i is the intermediate feature output by the i-th coupling layer, and f i is the i-th coupling layer, i = {1, 2, 3,..., k}, X is the latent feature, and Z is the Gaussian feature.
[0093] The final flow model can be expressed as , where F is the mapping estimating the latent features to the Gaussian features, is the stacking of the coupling layers. The probability estimation loss L of the flow modelF Can be expressed as:
[0094] ;
[0095] Wherein, is the standard Gaussian distribution probability function, is the mapping operation from the element x of the input latent feature to the Gaussian feature, is the Jacobian matrix of the invertible coupling layer of, log is to take the logarithm, is to find the absolute value of the matrix determinant.
[0096] Calculate the flow discriminant loss using the known features, and regularize the first flow model. Finally, the total loss of the first flow model Can be expressed as:
[0097] ;
[0098] The loss of the second flow model Consists of the probability estimation loss of the flow model:
[0099] .
[0100] Adopt the way of alternating training of the twin-encoded diffusion model and the first flow model to jointly optimize the two models. Set the period q of alternating training. When the number of training rounds is from 2pq to (2p + 1)q, p is a non-negative integer, fix the parameters of the first flow model, and use the loss function of the twin-encoded diffusion model for training. When the number of training rounds is from (2p + 1)q to (2p + 2)q, p is a non-negative integer, fix the parameters of the twin-encoded diffusion model, and use the loss function of the first flow model for training. This alternating process has nothing to do with the second flow model. The second flow model is trained simultaneously during the alternating training until the training is completed.
[0101] After the training is completed, dynamically adjust the abnormal proportion of the latent feature and the quantization residual feature using the validation set, and integrate them into the final abnormal score. The steps are as follows:
[0102] (1)Normalization processing: For and First, perform min-max normalization processing, adjust it to the range of 0 to 1, and then use bilinear interpolation to the image size.
[0103] ;
[0104] ;
[0105] Wherein is the probability distribution of the latent feature, calculated from the loss of the first flow model, is the probability distribution of the quantized residual features, which is calculated from the loss of the second flow model. Q is a bilinear interpolation operation that adjusts the probability distribution of the feature dimension to the image size. is the probability distribution of the normalized latent features, is the probability distribution of the normalized quantized residual features, min is the minimum operation, and max is the maximum operation.
[0106] (2)Search for the optimal weight on the validation set:
[0107] Traverse the candidate values of r on the validation set (e.g., r ∈ {0.1, 0.2,..., 0.9}), and select the weight coefficient r that maximizes AUC-ROC ∗ :
[0108] ;
[0109] where r is the dynamic weight coefficient, ranging from 0 to 1, AUC-ROC is a metric for evaluating the performance of an anomaly detection model, which measures the ability of the model to distinguish normal and abnormal at different thresholds. argmax is the value of the variable when the expression reaches the maximum.
[0110] (3)Final anomaly score S final The calculation formula of is:
[0111] ;
[0112] The twin-encoded diffusion model and the flow model are jointly and alternately trained. First, the image features are transferred to the noise space, and then the noise space is transferred to the latent space, so as to accurately estimate the probability distribution of normal images. When training the twin-encoded diffusion model, the L2 loss and the flow discriminant loss are combined to better fit the predicted noise. During the training process of the flow model, the known Gaussian noise data is used to regularize the coupling layer, which is beneficial for the flow model to learn more meaningful feature representations.
[0113] The method based on block feature clustering is used to quantize the image to be detected, and the quantized residual features between the input image and the quantized image are estimated. The anomaly discrimination of the quantized residual features and the latent features is dynamically combined to improve the robustness of anomaly detection.
[0114] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.
[0115] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.
Claims
1. An image anomaly detection method based on twin coding diffusion model and flow model, characterized in that: Here are the steps: S1: Obtain image samples and perform preprocessing to obtain normal training samples and validation set samples; S2: After adding fixed-scale noise to the normal training sample, a noisy image and a known noise are obtained. The original normal image and the noisy image are sent to the twin coding diffusion model to obtain the predicted noise, and the noise prediction loss and the flow discrimination loss are defined; S3: Build a first-stream model, extract features from predicted noise and known noise to obtain latent features and known features, use the first-stream model to estimate the probability distribution of latent features, and use known features to regularize the first-stream model, and perform joint alternating training with the twin coding diffusion model; build a quantization module, the quantization module obtains quantization features that match the input image, obtains quantization residual features based on the quantization features, and builds a second-stream model; The probability distribution of quantized residual features is estimated using the second-stream model, and the first-stream model, the second-stream model and the twin coding diffusion model are trained; S4: After training is completed, the probability distribution of the latent features and the probability distribution of the quantized residual features are dynamically adjusted using the validation set to obtain the optimal dynamic weight parameters and integrated into the final anomaly score; S5: Use the trained twin encoding diffusion model, the first-stream model, and the second-stream model to perform anomaly detection on the image to be detected.
2. The image anomaly detection method according to claim 1, characterized in that: In step S5, after adding noise of a fixed scale to the image to be detected during the diffusion process, a noisy image and a known noise are obtained, and the image to be detected and the noisy image are sent to the twin coding diffusion model to obtain predicted noise. Feature extraction is performed on the predicted noise and the known noise respectively to obtain potential features and known features respectively, and the probability distribution of the potential features is obtained through the first-stream model; the quantized residual features of the image to be detected are obtained based on the quantization module, and the probability distribution of the quantized residual features is obtained through the second-stream model; the probability distribution of the potential features and the probability distribution of the quantized residual features are integrated through the optimal dynamic weight parameters to obtain the anomaly score, and the anomaly score is compared with the artificially set anomaly detection threshold to obtain a binary mask image, and whether the image to be detected is abnormal is judged according to the binary mask image.
3. The image anomaly detection method according to claim 1, characterized in that: The twin coding diffusion model includes three parts. The first part is the twin encoder, the second part is the spatial feature fusion module, and the third part is the intermediate layer and decoder. The twin encoder is composed of two encoders with the same structure. The original normal image and the noisy image are input in parallel. The twin encoder performs feature interaction inside, and the first layer of features of the noisy image is added to the first layer of features of the original normal image. The last six layers of features of the original normal image are passed through the spatial feature fusion module and added to the features of the corresponding noisy image, and then jump-connected to the corresponding decoder. The intermediate layer serves as a bridge connecting the twin encoder and the decoder to further extract feature information. Finally, the predicted noise is obtained through the decoder.
4. The image anomaly detection method according to claim 3, characterized in that: The spatial feature fusion module is composed of a 3×3 convolutional layer, an instance normalization layer and a SiLU activation layer.
5. The image anomaly detection method according to claim 3, characterized in that: The noise prediction loss adopts L2 loss.
6. The image anomaly detection method according to claim 3, characterized in that: The total loss of the twin coding diffusion model is the sum of the stream discriminant loss and the noise prediction loss. The stream discriminant loss is defined as the arithmetic mean sum of the discriminant losses of all coupled layers in the first stream model.
7. The image anomaly detection method according to claim 1, characterized in that: The process of building a quantization module is as follows: First, use the feature extractor pre-trained on the ImageNet dataset to extract local features of the original normal image using the following formula to obtain the local features of each original normal image and form a local feature library: ; in is the feature of the d-th original normal image, is the dth original normal image, is a pre-trained feature extractor; The feature vector is flattened as follows to construct the original patch feature of the image: ; in It represents the e-th feature vector in the features of the d-th original normal image, where e is the feature vector number of the original image feature, C represents the channel, H is the height, and W is the width. Represents a C-dimensional vector, and flatten is a vector flattening operation; The original patch features of all the original normal images are gathered together to form a feature set, and the K-means clustering algorithm is used to perform dictionary learning on the set to obtain m cluster centers: ; in, is the set of cluster centers, is the vth cluster center feature, is the serial number of the cluster center, and m is the number of cluster centers; The cluster center is used as the codebook, i.e., the visual dictionary; the input image is extracted using a pre-trained feature extractor, and the feature is divided into S×S block-level sub-vectors: ; Among them, S 2 is the number of blocks into which the feature is divided, is the divided feature block, u is the sequence number of the feature block; The block-level subvectors are flattened to obtain patch features, and then their block-level features are encoded in a quantized manner, that is, the distance from each point in the patch feature to the cluster center is calculated, and each point in the feature is assigned to its nearest cluster to obtain the codebook-based quantized features as the final block features: ; in, is the quantized block feature, is an m-dimensional vector, is the feature vector of the g-th block feature, g is the feature vector sequence number of the block feature, and quantify is the quantization operation; All sub-blocks are processed in the same way to obtain sub-block features, which are used to form image quantization features. The quantization features are residualized with the image features to obtain quantization residual features, and the probability distribution of the quantization residual features is estimated using the second stream model.
Citation Information
Patent Citations
Denoising diffusion generative adversarial networks
US20230095092A1
Synthesizing three-dimensional shapes using latent diffusion models in content generation systems and applications
US20240005604A1
Cited By
Industrial image unsupervised anomaly detection method and system based on normalized stream
CN121190416A