Unsupervised anomaly detection method based on feature alignment and memory bank
By adopting an alignment-based method in unsupervised anomaly detection, aligning image features and establishing a memory database, the problem of detection performance degradation caused by target angle diversity is solved, and more efficient detection and lower storage costs are achieved.
Patent Information
- Application Number
- CN202510263031.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-06
AI Technical Summary
The existing unsupervised anomaly detection methods will significantly reduce the detection performance when the target angles are diverse, resulting in the need of more samples and a larger memory bank, increasing spatial complexity.
The unsupervised anomaly detection method based on alignment is adopted to align image features through image-level alignment networks and feature-level alignment networks, and a memory bank is established using the method of image block feature extraction, reducing the dependence on sample size and memory bank size.
Improves the accuracy and robustness of detection, reduces the storage cost required for memory bank construction, simplifies network structure, and reduces training and computing costs.
Smart Images

Figure CN120107226A_ABST
Abstract
Description
Technical field:
[0001] The invention relates to the field of anomaly detection of industrial images, and is an unsupervised anomaly detection method based on alignment and memory library. Technical background:
[0002] Anomaly detection in industrial images is an important research topic today. With the great progress of artificial intelligence related technologies and the gradual popularization of high-performance computing devices in recent years, the performance of anomaly detection algorithms has also improved. Among them, deep neural networks (DNNs), as a type of machine learning, are widely used in various scenarios. Their performance has been proven to gradually surpass traditional algorithms as the number of samples increases.
[0003] Deep neural networks can be divided into three categories: fully supervised, semi-supervised, and unsupervised, depending on the annotation of the data set. Since abnormal images are generally far less than normal images, semi-supervised and unsupervised learning models are more popular in the field of image anomaly detection. Especially when abnormal samples are missing, the performance of unsupervised anomaly detection algorithms is better than that of fully supervised algorithms. In addition, unsupervised anomaly detection can be divided into feature-embedding-based algorithms and reconstructed-based algorithms according to different usage methods. Both of them have been widely studied recently.
[0004] Among the existing unsupervised anomaly detection methods, the method based on feature embedding has good performance, and the method based on memory library can effectively improve the accuracy of detection. However, the performance of this method will drop significantly when the target angle is diverse. Solving this problem requires more samples, which requires the establishment of a larger memory library, resulting in the problem of linear increase in space complexity. Summary of the invention:
[0005] The purpose of the invention is to provide an unsupervised anomaly detection method based on alignment and memory library, aiming to improve the accuracy and robustness of industrial image anomaly detection, while reducing the storage cost required for memory library construction, so as to solve the technical problems mentioned in the background technology.
[0006] In order to solve the above technical problems, the specific technical solutions of the present invention are as follows:
[0007] An unsupervised anomaly detection method based on feature alignment and memory library, characterized by comprising the following steps:
[0008] Step S1: Collect sample images, divide the collected sample images into positive samples (normal images) and negative samples (abnormal images), annotate the sample images at the pixel level, use the negative samples and their pixel-level annotations as test samples, and preprocess the sample images to form a data set; since it is an unsupervised method, positive samples are used in the training process; and negative samples are mainly used in the testing stage;
[0009] Step S2: Take the normal image as a sample and train the image-level alignment network; the image-level alignment network consists of a spatial transformation network, which extracts image features and converts them into parameters of a transformation matrix, and uses the transformation matrix to perform an affine transformation on the original image to align all normal images; during the training process, it is necessary to take pairs of normal sample images and input them into the image-level alignment network to calculate the l between the paired outputs. 2 loss to optimize the image-level alignment network;
[0010] Step S3: train the feature-level alignment network, which is responsible for extracting the features of the aligned images in step S2 and further aligning the features of the aligned images; the feature-level alignment network uses the residual network as the main feature extraction part; the outputs of the second and third layers are used for alignment, and the alignment task is also performed by the spatial transformation network (slightly different from the spatial transformation network structure in step S2). The output after alignment is the feature required for the subsequent establishment of the memory library.
[0011] Step S4: Use the features output by the feature-level alignment network obtained in step S3 to establish a memory bank; this method uses an image block feature extraction method to establish a memory bank.
[0012] Step S5: Input the test sample into the image-level alignment network and the feature-level alignment network in turn to extract features; compare the extracted features with the features in the memory library, and calculate the distance between the extracted features and the nearest features in the memory library to generate an abnormal segmentation image and an abnormality score.
[0013] Further, the specific process of step S1 is:
[0014] S11: During the image enhancement process, the collected sample images are scaled to H reshape ×W reshape , where H reshape and W reshape They are the height and width of the scaled image respectively;
[0015] S12: After scaling, for each image d i , generate a random affine matrix use Right i and its pixel-level annotation M iPerform an affine transformation. When transforming, the edge is filled in a "reflective" way, and we get
[0016] S13: After scaling, Center align and crop to size H crop ×W crop , H crop and W crop are the height and width after cropping, and the final enhanced sample is obtained And divided into positive samples and negative samples
[0017] Further, the specific process of step S2 is:
[0018] S21: Divide the normal samples obtained in step S1 into multiple batches, each batch contains 2 images d m , d n ;
[0019] S22: Set the image-level alignment network to be represented as D m , d n Input image-level alignment network Get the aligned image By reducing and l 2 Loss to train image-level alignment network
[0020] Further, the specific process of step S3 is:
[0021] S31: In the feature-level alignment network, the second and third layers of the residual network are used 2 、layer 3 Each layer is followed by a spatial transformation network as a feature-level alignment network
[0022] S32: Use the image-level alignment network trained in step S2 The positive samples in step S1 To align; The output of the feature-level alignment network training samples;
[0023] S33: Using the Naive Siamese Network Training Method to Train Feature-Level Alignment Networks l between the twin network pair outputs 2 The norm is used as the loss and back-propagated to train the feature-level alignment network
[0024] Further, the specific process of step S4 is as follows:
[0025] S41: Align the image level network obtained in step S2 and step S3 and feature-level alignment network Connect as a feature extractor
[0026] S42: The positive sample obtained in step S1 enter For feature extraction, use layer 2 and layer 3 The output features establish a memory bank, and the method used to establish the memory bank is: image block feature extraction.
[0027] Further, the specific process of step S5 is:
[0028] S51: Use the feature extractor obtained in step S4 For the negative sample image obtained in step S1 Perform feature extraction and retain the affine transformation matrix generated by the spatial transformation network during the process and in For image-level alignment network The affine transformation matrix output by the spatial transformation network is and Feature-level alignment network Middle layer 2 and layer 3 The affine transformation matrix output by the subsequent spatial transformation network;
[0029] S52: Use the image block feature extraction method to calculate the distance between the extracted feature and its adjacent feature in the memory library to obtain a differential feature map and in and Layer 2 Layers and layers 3 Differential feature maps formed by layers;
[0030] S53: Use the affine transformation matrix retained in S51 to process the differential feature map in S52 to obtain the final abnormal segmentation map S:
[0031]
[0032] Where ⊙ is the symbol of affine transformation, and Inv(·) is the matrix inversion;
[0033] S54: Take the maximum value in the segmentation map S obtained in S53 as the abnormality score of the sample.
[0034] The present invention provides an unsupervised anomaly detection method based on feature alignment and memory library, which has the following advantages:
[0035] 1. The present invention reduces the sample size required for establishing a memory bank, and at the same time reduces the spatial complexity of the memory bank, making storage more space-saving.
[0036] 2. The present invention improves the accuracy and robustness of detection, and is particularly effective when facing targets that are not easily deformed.
[0037] 3. The network structure of the present invention is simple, the cost required for training is low, the calculation complexity is low, and the efficiency is high. Description of the drawings:
[0038] Figure 1 A flow chart of an unsupervised anomaly detection method based on feature alignment and memory library of the present invention;
[0039] Figure 2 This is the structure diagram of the image-level alignment network of the present invention;
[0040] Figure 3 This is a feature-level alignment network structure diagram of the present invention;
[0041] Figure 4 It is a complete structural diagram of the network of the present invention;
[0042] Figure 5 The results of the tests on capsule, pill, and screw on the MVTec dataset are shown in Figure 2.
[0043] Figure 6 A comparison chart of the present invention and other algorithms on the MVTec data set; Specific implementation method:
[0044] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings and specific implementation examples.
[0045] The unsupervised anomaly detection method based on feature alignment and memory library of the present invention specifically comprises the following steps:
[0046] Step S1: Collect samples and divide the collected samples into positive samples and negative samples. Positive samples are normal images and negative samples are abnormal images. Negative samples are annotated at the pixel level to form a data set. Since it is an unsupervised method, positive samples are used in the training process. Negative samples are mainly used in the testing phase.
[0047] Step S2: Train the image-level alignment network, which consists of a simple spatial transformation network. It extracts image features and converts them into parameters of a transformation matrix, and uses the transformation matrix to perform an affine transformation on the original image. Assume that the image-level alignment network is Randomly select a pair of samples d m , d n , through the network Get the aligned image Then, by reducing and l 2 Loss to train the network
[0048] Step S3: Train a feature-level alignment network, which is responsible for extracting the features of the aligned images in step 2 and further aligning their features.
[0049] The network consists of a residual network as the main feature extraction part, which is divided into three layers. 1 、layer 2 、layer 3 The layer needs to be followed by a spatial transformation network. In order to train the network more effectively, the feature extraction network is followed by a coding layer and a prediction layer. The training method of the twin network is used to pass a single positive sample image through the network (including the coding layer and the prediction layer) to calculate the l between the coding layer output and the prediction layer output. 2 The loss is calculated and back propagated on the path of the prediction layer output, while back propagation on the prediction layer output path is prohibited. The feature-level alignment network (excluding the encoding layer and the prediction layer) is thus trained.
[0050] Step S4: Use the features output by the feature-level alignment network obtained in step 3 to build a memory bank. Use the image block feature extraction method to divide the outputs of the last two layers of the feature-level alignment network, and use each divided part as a vector. Collect the vectors extracted from all training samples, sample them, and form a more compact cluster to reduce the size of the memory bank.
[0051] Step S5: Connect the parts obtained in steps S2, S3, and S4 in sequence to perform the final detection. After the negative sample is input into the image-level alignment network, it is input into the feature-level alignment network to extract features. The image block feature extraction method is used to divide the features and compare them with the vectors in the memory library. The difference between them is calculated to obtain the outlier value and outlier segmentation of the negative sample.
[0052] For the establishment of image-level alignment network, feature-level network and memory library, the specific steps are as follows:
[0053] 1. In the image anomaly detection method based on self-supervised learning and feature denoising, the structure and training method of the image-level alignment network are as follows:
[0054] The structure of the image-level alignment network is similar to the spatial transformation network. It extracts features through two layers of convolutional pooling, and predicts the transformation matrix through two layers of fully connected networks. Then, the coordinates of the input samples are transformed. The spatial transformation network is represented as
[0055] Assume that the training sample is For a sample d i , the source coordinate point can be expressed as Then the sample gets the transformation matrix through the network And after coordinate transformation, the target coordinates can be obtained As shown below:
[0056]
[0057] During the training process, a pair of samples d is randomly selected m , d n , through the network Get the aligned image Then, by reducing l 2 Loss to optimize the network
[0058] 2. In the image anomaly detection method based on self-supervised learning and feature denoising, the structure and training method of the feature-level alignment network are as follows:
[0059] In order to improve the accuracy of subsequent abnormal segmentation and reduce the space required to build the memory library, the feature pixels are aligned on this basis. Here, the residual network is used as the backbone of feature extraction, and the first two layers layer2 and layer3 are taken. The spatial transformation network is inserted between each layer and at the end of the last layer to enable the network to obtain the ability to align feature pixels. Here, it is assumed that the pixel-level alignment module is
[0060] To train the pixel-level alignment module, As a twin network, followed by an encoder ε(·) and a prediction period Here, instead of a fully connected network, a two-dimensional convolutional network is used, so that both the encoder and the prediction period are composed of multiple convolutional modules.
[0061] Will pass The module takes as input an image rectified by, e.g. Through the module Get F m , F nThe obtained features are passed through the encoder ε(·) and the prediction period get and z is obtained only through the encoder ε(·) m =ε(F m ), z n =ε(F n ), the final loss function is as follows:
[0062]
[0063] Here m , z n They do not participate in the back propagation of gradients, but participate in the calculation of loss as constants.
[0064] 3. In the image anomaly detection method based on self-supervised learning and feature denoising, the memory library is constructed as follows:
[0065] Based on the backbone network described above, the feature extraction network The output features of layer2 and layer3 construct the memory bank.
[0066] Assume the output feature of layeri is C, H and W represent the number of channels, length and width of the feature respectively, then F i The coordinate position of a channel c The neighborhood of can be expressed as follows, where p represents the size of the neighborhood. .
[0067]
[0068] Flatten the neighborhood of the same coordinate position in each channel and connect them in the channel dimension to obtain the feature vector of that position 0≤x i ≤W,0≤y i ≤H. Then, Adaptive average pooling is used to further compress the feature vector, and we get (mod represents the dimension of the vector after pooling, which is a hyperparameter.) Here, feature vectors of different layers are not merged, but processed separately to reduce loss and improve accuracy.
[0069] Next, multiple batches of feature vectors Sampling is performed to reduce the size of the memory bank. In addition, due to the alignment process of feature pixels in the previous step, more compact clusters are formed between similar features, which can further reduce the capacity of the memory bank.
[0070] In the test phase, the image also extracts feature vectors at different locations The affine transformation matrix generated by the spatial transformation network during the preservation process and (in is the affine transformation matrix output by the spatial transformation network of the image-level alignment network, and They are respectively the layers in the feature-level alignment network 2 and layer 3 The affine transformation matrix output by the subsequent spatial transformation network) is then used to find the n nearest eigenvectors in the memory bank. Will and The average distance between them is taken as the abnormal score at that position, and the differential feature map is obtained. and And perform the following transformation to obtain the abnormal segmentation map S
[0071]
[0072] Where ⊙ is the symbol of affine transformation, and Inv(·) is the matrix inversion. The maximum value of S is taken as the final anomaly score of the sample.
[0073] Embodiment 1:
[0074] The unsupervised anomaly detection method based on feature alignment and memory library has the following specific process: Figure 1 As shown, the details are as follows:
[0075] Step 1: Collect samples and divide them into positive samples and negative samples. Positive samples are normal images and negative samples are abnormal images. Negative samples are annotated at the pixel level to form a data set. Since it is an unsupervised method, positive samples are used in the training process. Negative samples are mainly used in the testing phase.
[0076] Step 2: In the training process of the image-level alignment network, first, the collected image is scaled to a resolution of 224×224. If the sample is not rich enough, it can be enhanced in advance, and a series of affine transformations are performed before inputting it into the image-level alignment network for training. Assume that the image-level alignment network is Randomly select a pair of samples d m , d n , through the network Get the aligned image Then, by reducing and l 2 Loss to train the network The structure of the network is as follows Figure 2As shown in the figure, Conv represents the convolutional layer, Max-Pool represents the maximum pooling layer, and FC represents the fully connected layer. The specific training rounds perform differently on different samples, depending on the results. The network parameters can be randomly initialized. Table 1 introduces the structure of each layer of the image-level alignment network (the output of Linear-2 depends on the selected affine transformation).
[0077] Table 1 Structure of each layer of the image-level alignment network
[0078]
[0079]
[0080] Step 3: Since the image resolution has been normalized to 224×224 in step 1, in this step, we only need to normalize the aligned image obtained in step 1. As a twin network, followed by an encoder ε(·) and a prediction period Here, instead of a fully connected network, a two-dimensional convolutional network is used, so that both the encoder and the prediction period are composed of multiple convolutional modules.
[0081] Will pass The module takes as input an image rectified by, e.g. Through the module Get F m , F n The obtained features are passed through the encoder ε(·) and the prediction period get and z is obtained only through the encoder ε(·) m =ε(F m ),z n =ε(F n ), the final loss function is as follows:
[0082]
[0083] Here m ,z n They do not participate in the back propagation of gradients, but participate in the calculation of loss as constants. Figure 3 As shown in Figure 2, the training rounds are about 50 times. ResNet-layer is a layer in the residual network pre-trained on ImageNet. The corresponding module parameters can be initialized, and other modules can be randomly initialized. In addition, ResNet-layer must be frozen during training. Table 2 introduces the specific structure of each layer of the network.
[0084] Table 2 Structure of each layer of feature-level alignment network
[0085]
[0086]
[0087] Step 4: Through step 2, the alignment features extracted from the positive samples can be obtained, and a memory library can be established. The construction method of the memory library is as follows:
[0088] Based on the backbone network described above, the feature extraction network The output features of layer2 and layer3 construct the memory bank.
[0089] Assume the output feature of layeri is Then F i The coordinate position of a channel c The neighborhood of can be expressed as follows.
[0090]
[0091] Flatten the neighborhood of the same coordinate position in each channel and connect them in the channel dimension to obtain the feature vector of that position 0≤x i ≤W,0≤y i ≤H. Then, Adaptive average pooling is used to further compress the feature vector, and we get (mod is the dimension of the vector after pooling, which is a hyperparameter.) Here, feature vectors of different layers are not merged, but processed separately to reduce loss and improve accuracy.
[0092] Next, multiple batches of feature vectors Sampling is performed to reduce the size of the memory bank. In addition, due to the alignment process of feature pixels in the previous step, more compact clusters are formed between similar features, which can further reduce the capacity of the memory bank.
[0093] Step 5: After the above 4 steps, the complete detection network has been completely established, and its structure is as follows Figure 4 As shown. In the test phase, the image also extracts feature vectors at different positions Then find the n nearest eigenvectors in the memory bank Will and The average distance between the two positions is taken as the anomaly score at that position, which is used as the basis for anomaly segmentation. The maximum value of the anomaly scores at all positions is taken as the anomaly score of the sample.
[0094] Figure 6The performance comparison between the detection method of the present invention and some other detection methods on the MVTec-AD dataset is shown. The comparison indicators include image-level AUROC and pixel-level AUROC. It can be seen that the performance of the method described in this patent exceeds other methods on multiple sample classes and has relatively good performance. In addition, we also visualized the experimental results, such as Figure 5 As shown in the figure, from left to right are the original image, the pixel-level annotation map, and the heat map obtained by anomaly segmentation.
[0095] It is to be understood that the present invention is described by some embodiments, and it is known to those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the scope of protection of the present invention.
Claims
1. An unsupervised anomaly detection method based on feature alignment and memory library, characterized in that: The following steps are involved: Step S1: Collect sample images, divide the collected sample images into positive samples and negative samples, where positive samples are normal images and negative samples are abnormal images; perform pixel-level annotation on sample images, use negative samples and their pixel-level annotations as test samples, and preprocess sample images to form a data set; use positive samples in the training process; and negative samples are mainly used in the test stage; Step S2: Use normal images as samples to train an image-level alignment network; the image-level alignment network consists of a spatial transformation network, which extracts image features and converts them into parameters of a transformation matrix, and uses the transformation matrix to perform an affine transformation on the original image, so as to align all normal images; During the training process, it is necessary to take pairs of normal sample images and input them into the image-level alignment network, and calculate the l2 loss between the paired outputs to optimize the image-level alignment network; Step S3: training a feature-level alignment network, which is responsible for extracting features of the aligned images in step S2 and further aligning the features of the aligned images; The feature-level alignment network uses the residual network as the feature extraction part, uses the output of the second and third layers for alignment, and the alignment task is handed over to the spatial transformation network, which outputs the features after alignment; Step S4: Use the features output by the feature-level alignment network obtained in step S3 to build a memory library; Step S5: input the test sample into the image-level alignment network and the feature-level alignment network in sequence to extract features; The distance between the extracted features and the nearest features in the memory bank is calculated to generate an anomaly segmentation image and anomaly score.
2. The unsupervised anomaly detection method based on feature alignment and memory library according to claim 1, characterized in that: The specific process of step S1 is as follows: S11: During the image enhancement process, the collected sample images are scaled to H reshape ×W reshape , where H reshape and W reshape They are the height and width of the scaled image respectively; S12: After scaling, for each image D i , generate a random affine matrix use Right D i and its pixel-level annotation M i Perform an affine transformation. When transforming, the edge is filled in a reflective manner, and we get S13: After scaling, Center align and crop to size H crop ×W crop , H crop and W crop are the height and width after cropping, and the final enhanced sample is obtained And divided into positive samples and negative samples 3. The unsupervised anomaly detection method based on feature alignment and memory library according to claim 2, characterized in that: The specific process of step S2 is: S21: Divide the normal samples obtained in step S1 into multiple batches, each batch contains 2 images d m , d n ; S22: Set the image-level alignment network to be represented as D m , d n Input image-level alignment network Get the aligned image By reducing and The l2 loss between them is used to train the image-level alignment network 4. The unsupervised anomaly detection method based on feature alignment and memory library according to claim 3, characterized in that: The specific process of step S3 is as follows: S31: In the feature-level alignment network, the second and third layers of the residual network, layer2 and layer3, are used, and each layer is connected to a spatial transformation network as a feature-level alignment network. S32: Use the image-level alignment network trained in step S2 The positive samples in step S1 To align; The output of the feature-level alignment network training samples; S33: Using the Naive Siamese Network Training Method to Train Feature-Level Alignment Networks The l2 norm between the twin network pair outputs is used as the loss for backpropagation to train the feature-level alignment network.
5. The unsupervised anomaly detection method based on feature alignment and memory library according to claim 4, characterized in that: The specific process of step S4 is as follows: S41: Align the image level network obtained in step S2 and step S3 and feature-level alignment network Connect as a feature extractor S42: The positive sample obtained in step S1 enter Feature extraction is performed in , and the features output by layer2 and layer3 are used to establish a memory bank. The method used to establish the memory bank is: image block feature extraction.
6. The unsupervised anomaly detection method based on feature alignment and memory library according to claim 5, characterized in that: The specific process of step S5 is as follows: S51: Use the feature extractor obtained in step S4 For the negative samples obtained in step S1 Perform feature extraction and retain the affine transformation matrix generated by the spatial transformation network during the process and in For image-level alignment network The affine transformation matrix output by the spatial transformation network is and Feature-level alignment network The affine transformation matrix output by the spatial transformation network following layer2 and layer3; S52: Use the image block feature extraction method to calculate the distance between the extracted feature and its adjacent feature in the memory library to obtain a differential feature map and S53: Use the affine transformation matrix retained in S51 to process the differential feature map in S52 to obtain the final abnormal segmentation map S: Where ⊙ is the symbol of affine transformation, and Inv(·) is the matrix inversion; S54: Take the maximum value in the segmentation map S obtained in S53 as the abnormality score of the sample.