A method for image anomaly detection based on memory network
By adopting a memory network-based method in image abnormality detection, using multi-decoder and knowledge distillation technology, the problem of low image reconstruction quality and inability to completely eliminate abnormal areas in the prior art is solved, and higher quality image reconstruction and more accurate abnormality detection are achieved.
Patent Information
- Application Number
- CN202210641017.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-07
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-06-07
AI Technical Summary
The existing image abnormality detection method based on the autoencoder has low quality for reconstruction of the image and cannot completely eliminate the abnormal areas in the input image, resulting in the normal samples being misjudged as abnormal.
Using the image abnormality detection method based on the memory network, a network model including a first encoder, a memory network and multiple decoders is constructed, and matching mapping features are queried through the memory network and multi-decoder reconstruction is performed. Combined with knowledge distillation, lightweight encoder is extracted to improve reconstruction quality and abnormal detection accuracy.
The image reconstruction quality is improved, the detection accuracy of abnormal images is enhanced, and the possibility that normal images are misjudged as abnormal.
Smart Images

Figure CN114882007B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer vision technology, and in particular, relates to an image anomaly detection method based on a memory network. Background Art
[0002] Image anomaly detection is a hot research direction in the field of computer vision. Its research goal is to use normal samples to train a specific model to detect various abnormal images that may appear without using real abnormal samples. It has high research significance and application value in the fields of industrial product defect detection, medical image analysis, video surveillance and security. The difficulty of image anomaly detection is relatively large, mainly reflected in the following points: ① The heterogeneity and unknown nature of abnormal categories in images: Anomalies are irregular, and one type of anomaly may show completely different abnormal characteristics from another type of anomaly. For example, in industrial products, the shape and position of the defects of the outer packaging are unknown. But when the anomaly does not occur, it is unknown what the anomaly is. ② Unbalanced categories and few abnormal samples: Anomalies are usually rare data instances, while normal instances usually account for the vast majority of the data. Therefore, it is difficult or even impossible to collect a large number of labeled abnormal instances. This leads to the inability to have positive and negative samples to learn and train models like conventional supervised learning.
[0003] Due to the above research status, the goal of image anomaly detection can only be to detect abnormal images or local abnormal areas that are different from normal images through unsupervised or semi-supervised learning (a small number of negative samples or artificially created negative samples). At present, this research direction has made some research progress with the joint efforts of many researchers. According to whether neural networks are involved in the model construction stage, the existing image anomaly detection methods can be divided into two categories: traditional methods and deep learning-based methods. The anomaly detection technology based on traditional methods generally includes the following branches: anomaly detection models based on template matching, statistical models, frequency domain analysis, and classification surface construction. The methods based on deep learning are roughly divided into the following categories: anomaly detection methods based on high-dimensional feature vector distance measurement, based on input image reconstruction comparison, and combined with traditional methods.
[0004] In recent years, traditional machine learning methods have been widely used in the field of image anomaly detection. With the development of deep learning technology, combining neural networks to realize image anomaly detection has become a new research technology. Among them, the anomaly detection method based on the reconstruction of the input image of the input neural network has become more and more popular. The core idea of the method based on the reconstruction of the input normal image is to encode the input normal image through the neural network, and use the decoder to decode and reconstruct the extracted high-dimensional features, and train the neural network with the reconstructed input as the goal. Then, in the detection stage, the purpose of anomaly detection is achieved by comparing the difference between the input normal image and the reconstructed image. According to the training mode adopted, the commonly used methods based on the reconstruction of the input normal image roughly include two types: based on autoencoders and based on generative adversarial networks (GAN).
[0005] In the method based on the reconstruction of the input normal image, the most commonly used network structure is the autoencoder (AE). The autoencoder built by training only with normal samples is expected to reconstruct the normal image with higher quality in the test phase. For the abnormal images in the test, there will be differences with the normal images in the image encoding and subsequent decoding and reconstruction process, and the size of the difference can be used as an indicator to measure the degree of abnormality of the sample to be tested. The structure of the autoencoder generally consists of an encoder and a decoder, and the network structures of the two are generally symmetrical. Among them, the encoder continuously reduces the width and height of the feature map while increasing its image channel dimension during the forward propagation of the network to delete redundant information. The decoder is responsible for decoding the features to obtain an image of the same size as the input image, and trains the network model by comparing and calculating the difference between the input normal image and the reconstructed normal image. The most commonly used loss function in this process is the mean square error (MSE). MSE uses the square mean of the difference between the pixel values of all pixels in the image before and after reconstruction to measure the quality of image reconstruction. After training, due to the existence of the bottleneck structure, for some samples with small abnormal areas, the autoencoder can eliminate the influence of the abnormal area during the image encoding and decoding process, reconstruct a normal image as a reference, and then obtain the abnormal area by pixel-by-pixel comparison. This method can not only realize the detection of abnormal images but also locate the abnormal detection area of the image, that is, locate the specific abnormal location area.
[0006] However, the image reconstruction method based on the autoencoder has an obvious shortcoming that the reconstructed image is relatively blurry whether in the training or testing phase, which may cause the network model to reconstruct normal samples into abnormal images. In addition to the problem of low quality of reconstructed images, the method based on the autoencoder also has the problem of not being able to completely eliminate abnormal areas in the input image. When the training samples are more diverse, the autoencoder will show strong learning ability and have too strong adaptability to potential abnormal samples. Summary of the invention
[0007] The purpose of this application is to provide an image anomaly detection method based on a memory network, so as to solve the problem that the image reconstructed by the prior art solution is of low quality and cannot guarantee the complete elimination of abnormal areas in the input image.
[0008] In order to achieve the above purpose, the technical solution of this application is as follows:
[0009] A memory network-based image anomaly detection method, comprising:
[0010] Constructing an image anomaly detection network model, wherein the image anomaly detection network model includes a first encoder, a memory network, and at least two decoders, wherein the first encoder uses a neural network VGG-16;
[0011] The constructed image anomaly detection network model is trained using a normal image training data set, the training sample is input into the first encoder to extract high-dimensional features, the mapping features matching the high-dimensional features are queried in the memory network, and then the mapping features are respectively input into the decoder to reconstruct the image, the reconstructed image with the smallest covariance value with the original training sample is taken as the output reconstructed image, the joint loss is calculated to update the parameters of the image anomaly detection network model, and the training is completed;
[0012] A lightweight second encoder based on the first encoder is extracted through knowledge distillation, the maximum pooling layer in the last four convolutional blocks of the first encoder is passed to the second encoder as a knowledge distillation layer, and the first encoder in the image anomaly detection network model is replaced by the second encoder to generate the final image anomaly detection network model;
[0013] The image to be detected is input into the final image anomaly detection network model, the reconstructed image is output, the anomaly detection scores of the input image to be detected and the reconstructed image are calculated, and it is determined whether the input image to be detected is abnormal.
[0014] Furthermore, based on the neural network VGG-16, the second encoder removes the last layer of convolution of the last three convolution blocks, discards the last fully connected layer of VGG-16, and passes the maximum pooling layer of the last four convolution blocks of the first encoder VGG-16 as the knowledge distillation layer to the last four convolution blocks of the second encoder.
[0015] Further, the querying of the mapping feature matching the high-dimensional feature in the memory network includes:
[0016] The high-dimensional features extracted by the first encoder are used as a query feature vector item set of the memory network, and each feature vector item in the high-dimensional features is used as a query feature vector item;
[0017] The matching probability between each query feature vector item and all prototype feature vector items stored in the memory network is calculated, and then the weighted average of the prototype feature vector items and their corresponding matching probabilities is calculated as the queried feature vector item, and all the queried feature vector items are combined into a mapping feature that matches the input high-dimensional feature.
[0018] Furthermore, the calculation formula for calculating the matching probability between each query feature vector item and all prototype feature vector items stored in the memory network is as follows:
[0019]
[0020] Among them, w t,m is the calculated matching probability, exp is an exponential function with the natural constant e as the base, p m represents the prototype eigenvector term, q t Represents the query feature vector item, and M represents the number of prototype feature vector items stored in the memory network.
[0021] Furthermore, the image anomaly detection method based on memory network also includes:
[0022] The high-dimensional features extracted by the first encoder are used as a query feature vector item set of the memory network, and each feature vector item in the high-dimensional features is used as a query feature vector item;
[0023] Calculate the matching probability v between each prototype feature vector item stored in the memory network and all query feature vector items t,m :
[0024]
[0025] Among them, p m represents the prototype eigenvector term, q t represents the query feature vector item, Q is the number of query feature vector items;
[0026] The matching probability v t,m Standardize to get v′ t,m , the standardized formula is as follows:
[0027]
[0028] Finally, the prototype feature vector item is updated by the following formula:
[0029] p m =f(p m +∑ t∈Q v′ t,m q t );
[0030] where f() is the L2 function.
[0031] Furthermore, the calculation of the abnormality detection scores of the input image to be detected and the reconstructed image includes:
[0032] Calculate the L2 distance between each query feature vector item after the image to be detected passes through the second encoder and the best matching feature vector item in the memory network:
[0033]
[0034] Where Q represents the number of query feature vector items, q t represents the query feature vector item, p s Represents the best matching prototype feature vector item in the memory network;
[0035] Calculate the peak signal-to-noise ratio of the image to be detected and the reconstructed image:
[0036]
[0037] Where N is the number of pixels in the image to be detected, x represents the image to be detected, represents the reconstructed image, It means to find the best reconstructed image;
[0038] The L2 distance and peak signal-to-noise ratio are normalized, and then the weighted sum of the two is calculated as the anomaly detection score.
[0039] Furthermore, the image anomaly detection method based on memory network also includes:
[0040] Calculate the input image x and the output image The weighted reconstruction error between t , the calculation formula is as follows:
[0041]
[0042] Among them, W t (.) is the weight function, and the calculation formula is as follows:
[0043]
[0044] When the fraction ε t When x is higher than a threshold γ, it is regarded as an abnormal image and is not used to update the prototype feature vector item in the memory network. Otherwise, it is used to update the prototype feature vector item in the memory network.
[0045] Furthermore, the weighted sum of the two is calculated as the anomaly detection score, and the calculation formula is as follows:
[0046]
[0047] Among them, g(.) is the normalization operation, λ is the weight coefficient, S t Represents the calculated anomaly detection score.
[0048] The present application proposes a method for detecting anomalies in an image based on a memory network. On the basis of the memory network, multiple decoders are used to improve the reconstruction quality of normal images. Then, when detecting abnormal samples, the abnormal samples will also be reconstructed according to normal samples, thereby highlighting the detection accuracy of abnormal images. With the help of knowledge distillation, the highly sensitive characteristics of the teacher network to normal samples are refined to the student network, so that the student network can maintain sensitivity to normal images during testing, but when encountering abnormal images, the extracted features can be significantly different from the features of normal images, so that the feature query feature vector items obtained are mostly abnormal features. By introducing the knowledge distillation lightweight feature extraction network model to improve the encoder to improve the encoding sensitivity to abnormal images and introducing multiple decoders to improve the reconstruction quality of normal samples, an effective method for detecting anomalies in images is realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 This is a flow chart of the image anomaly detection method based on memory network in this application;
[0050] Figure 2 This is a schematic diagram of the structure of the image anomaly detection network model in the embodiment of the present application;
[0051] Figure 3 Schematic diagram of encoder knowledge distillation in an embodiment of the present application. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0053] In one embodiment, Figure 1 As shown in FIG. 1 , a memory network-based image anomaly detection method is proposed, including:
[0054] Step S1: construct an image anomaly detection network model, wherein the image anomaly detection network model includes a first encoder, a memory network and at least a decoder, and the first encoder adopts VGG-16.
[0055] In this embodiment, the image anomaly detection network model is as follows: Figure 2 As shown, it includes a first encoder, a memory module and at least two decoders with the same structure. Considering the cost of computing performance, the number of decoders is preferably three.
[0056] In a specific embodiment, the first encoder uses VGG-16, which is a neural network commonly used in the machine learning library Pytorch, which often comes with pre-trained network parameters. The networks of each decoder in this embodiment can also use the VGG-16 structure.
[0057] Step S2, using a normal image training data set to train the constructed image anomaly detection network model, inputting the training sample into the first encoder to extract high-dimensional features, querying the mapping features matching the high-dimensional features in the memory network, and then inputting the mapping features into the decoder to reconstruct the image, taking the reconstructed image with the smallest covariance value with the original training sample as the output reconstructed image, calculating the joint loss to update the parameters of the image anomaly detection network model, and completing the training.
[0058] The training data set used in this embodiment takes the Ped2 data set of UCSD as an example. The Ped2 data set of UCSD contains 16 training data image sets and 12 test image sets, including 12 irregular events, including riding a bicycle and driving a vehicle. First, the data is preprocessed to adjust the size of the image to 256×256×3, where the three values are the width of the image, the height of the image, and the number of channels of the image. Four images are used as a batch as the input of the encoder for training.
[0059] During the training process, the training sample image passes through the first encoder to extract high-dimensional features with a feature size of 14×14×512, which is used as the query feature vector item set of the memory network, including 14×14 feature vector items. For any feature vector item q t(t∈Q, 14×14 in this embodiment), the closest prototype feature vector item is searched in the memory network. The memory network saves the feature vector item corresponding to the normal data as the prototype feature vector item, that is, if the input is normal data, the memory network will save its corresponding feature vector item as the prototype feature vector item for query.
[0060] After the memory network queries the closest prototype feature vector item, it outputs the closest prototype feature vector item obtained through the query. After all 14×14 feature vector items of the high-dimensional feature are queried, all prototype feature vector items output by the memory network are combined into a mapping feature that matches the input high-dimensional feature, and the mapping feature has the same size as the high-dimensional feature.
[0061] The obtained mapping features are input into each decoder for decoding and reconstructing the image, and then the obtained multiple reconstructed images are compared with the original input image, and the reconstructed image with the smallest covariance value with the original training sample is taken as the output reconstructed image. After a batch, the joint loss is calculated to update the parameters of the image anomaly detection network model, and training is performed batch by batch until the network converges and the training is completed.
[0062] It should be pointed out that, in order to query the mapping features matching the high-dimensional features in the memory network, the high-dimensional features extracted by the first encoder can be directly used as the query feature vector item set of the memory network, that is, each feature vector item in the high-dimensional features is used as the query feature vector item, and the closest prototype feature vector item is queried in the memory network. After the memory network queries the closest prototype feature vector item, the closest prototype feature vector item obtained by the query is output, and all the prototype feature vector items output by the memory network are combined into a mapping feature matching the input high-dimensional feature.
[0063] In a specific embodiment, the present application queries the memory network for mapping features that match the high-dimensional features, including:
[0064] The high-dimensional features extracted by the first encoder are used as a query feature vector item set of the memory network, and each feature vector item in the high-dimensional features is used as a query feature vector item;
[0065] The matching probability between each query feature vector item and all prototype feature vector items stored in the memory network is calculated, and then the weighted average of the prototype feature vector items and their corresponding matching probabilities is calculated as the queried feature vector item, and all the queried feature vector items are combined into a mapping feature that matches the input high-dimensional feature.
[0066] For example, the memory network stores M 1×1×512 prototype feature vector items, recording the most typical features of various normal data. This application uses p m∈M(m=1,…,M) represents a prototype feature vector item stored in the memory network.
[0067] This embodiment first calculates each query feature vector item q t and the prototype eigenvector term p m The matching probability w between t,m , the calculation formula is as follows:
[0068]
[0069] Among them, exp is an exponential function with the natural constant e as the base.
[0070] For each query feature vector item q t , by calculating the prototype eigenvector term p m With matching probability w t,m The weighted average of the queried feature vector item q t ′, the calculation formula is as follows:
[0071]
[0072] After obtaining the eigenvector item q obtained by the query t ′∈R 14×14×512 Finally, they are aggregated to obtain mapping features that match the input high-dimensional features, and then decoded and reconstructed by the decoder.
[0073] This embodiment uses all feature items instead of the closest feature items, which allows the network model of this application to understand the feature distribution of different normal data and take into account the overall normal features. That is, this application uses the prototype feature vector item p in the memory network m The query feature vector item q is represented by a combination of t This embodiment applies the read operation to each query feature vector item to obtain a converted feature mapping item q t ′∈R 14×14×512 , and then decode and reconstruct them by the decoder. This enables the decoder to reconstruct the input frame using the most typical feature items of the normal samples stored in the memory network, so that the reconstructed image is more inclined to the normal image, reducing the ability of the decoder to reconstruct abnormal images.
[0074] In a specific embodiment, the memory network needs to store feature vector items corresponding to normal data as prototype feature vector items. This embodiment provides a method for updating prototype feature vector items in a memory network, including:
[0075] The high-dimensional features extracted by the first encoder are used as a query feature vector item set of the memory network, and each feature vector item in the high-dimensional features is used as a query feature vector item;
[0076] Calculate the matching probability v between each prototype feature vector item stored in the memory network and all query feature vector items t,m :
[0077]
[0078] Among them, p m represents the prototype eigenvector term, q t represents the query feature vector item, Q is the number of query feature vector items;
[0079] The matching probability v t,m Standardize to get v′ t,m , the standardized formula is as follows:
[0080]
[0081] Finally, the prototype feature vector item is updated by the following formula:
[0082] p m =f(p m +∑ t∈Q v′ t,m q t );
[0083] where f() is the L2 function.
[0084] This embodiment calculates the matching probability between each prototype feature vector item and all query feature vector items, and selects all query feature vector items to update the closest prototype feature vector item. By using the weighted average of the query feature vector items instead of aggregating and summing them, the present application can focus more on the query feature vector items near the prototype feature vector items.
[0085] In this embodiment, the joint loss includes the reconstruction loss l rec , feature compactness loss l compact and feature separation loss l separateness , and add the weight coefficient λ c and λ s To balance the weight of the last two loss functions, the calculation formula is as follows:
[0086] Total loss = l rec +λ c l compact +λ s l separateness
[0087] The image reconstruction loss calculation formula is as follows:
[0088]
[0089]
[0090] where x 1 , x 2 , x 3 are the outputs of the three decoders respectively, and x is the original input image.
[0091] The loss of feature compactness (compression) is calculated as follows:
[0092]
[0093] Where s is the query q t The index number of the most matching item in the corresponding prototype feature vector item is calculated as:
[0094]
[0095] That is, p s Represents the best matching prototype feature vector item in the memory network, that is, the prototype feature vector item with the highest matching probability.
[0096] Feature separation loss function, similar queries should be assigned to the same items to reduce the number of items and memory size. Training the model using feature compression loss will only make all memory feature items similar, so all query feature items are tightly mapped into the embedding space, losing the ability to record different normal patterns. However, the feature items in memory should be far enough apart to consider various feature styles of normal data. In order to prevent this problem when obtaining compact feature representation, a feature separation loss is designed, and the α factor is used to adjust the feature separation loss function. The calculation formula is as follows:
[0097]
[0098] Where n is the query feature item q t The second most recent index number is calculated as follows:
[0099]
[0100] Step S3, extract a lightweight second encoder based on the first encoder through knowledge distillation, pass the maximum pooling layer in the last four convolution blocks of the first encoder as the knowledge distillation layer to the second encoder, replace the first encoder in the image anomaly detection network model with the second encoder, and generate the final image anomaly detection network model.
[0101] In this embodiment, for a trained first encoder, a lightweight second encoder based on the first encoder is extracted through knowledge distillation.
[0102] Specifically, Figure 2 As shown, the first encoder is VGG-16 ( Figure 2 The second encoder is ( Figure 2 Next) Based on the pre-trained VGG-16 provided in Pytorch, the last convolution layer of the last three convolution blocks (Conv2-Conv4) is removed (from the original three convolution layers to two convolution layers), and the last fully connected layer of VGG-16 is discarded, and 14×14×512 is used as the final network output. And the maximum pooling layer of the last four convolution blocks (Conv1-Conv4) of the first encoder VGG-16 is passed to the last four convolution blocks of the second encoder as the knowledge distillation layer.
[0103] The final image anomaly detection network model retains the memory network and decoder in the trained image anomaly detection network model. The network structure layer of each decoder is consistent with the encoder during training, which will not be repeated here.
[0104] Step S4: input the image to be detected into the final image anomaly detection network model, output the reconstructed image, calculate the anomaly detection scores of the input image to be detected and the reconstructed image, and determine whether the input image to be detected is abnormal.
[0105] The final image anomaly detection network model is used to detect the input image to be detected, the anomaly detection scores of the input image to be detected and the reconstructed image are calculated, and it is determined whether the input image to be detected is abnormal.
[0106] Among them, the peak signal-to-noise ratio (PSNR) of the input image to be detected and the reconstructed image can be directly used as the anomaly detection score. When the image to be detected is abnormal, a lower PSNR value is obtained, otherwise it is a normal image.
[0107] In a specific embodiment, calculating the anomaly detection scores of the input image to be detected and the reconstructed image includes:
[0108] Calculate the L2 distance between each query feature vector item after the image to be detected passes through the second encoder and the best matching feature vector item in the memory network:
[0109]
[0110] Where Q represents the number of query feature vector items, q t represents the query feature vector item, p s Represents the best matching feature vector item in the memory network;
[0111] Calculate the peak signal-to-noise ratio of the image to be detected and the reconstructed image:
[0112]
[0113] Where N is the number of pixels in the image to be detected, x represents the image to be detected, represents the reconstructed image, It means to find the best reconstructed image;
[0114] The L2 distance and peak signal-to-noise ratio are normalized, and then the weighted sum of the two is calculated as the anomaly detection score.
[0115] Specifically, the anomaly detection score S t The calculation formula is as follows:
[0116]
[0117] g(.) is the normalization operation, λ is the weight coefficient, and the specific normalization formula is as follows:
[0118]
[0119] After the anomaly detection score is calculated, it is compared with the set threshold, and the image to be detected whose anomaly detection score is greater than the set threshold is determined as an abnormal image, otherwise it is determined as a normal image.
[0120] It should be pointed out that after the network model is trained, when the network model is tested or the network model is used to detect the image to be detected, the input may be a normal image or an abnormal image. In order to expand the prototype feature vector items stored in the memory network, the feature vector items corresponding to the normal image can also be stored in the memory network as the prototype feature vector items.
[0121] To this end, this application also includes:
[0122] Calculate the input image x and the output image The weighted reconstruction error between t , the calculation formula is as follows:
[0123]
[0124] Among them, W(.) is the weight function, and the calculation formula is as follows:
[0125]
[0126] When the fraction ε tWhen x is higher than a threshold γ, x is regarded as an abnormal image, so it is not used to update the prototype feature vector item in the memory network. Otherwise, it is used to update the prototype feature vector item in the memory network. How to update the prototype feature vector item in the memory network has been explained in the previous steps and will not be repeated here.
[0127] The anomaly detection method proposed in the present application improves the reconstruction quality of normal images in the process of reconstructing images based on a decoder, thereby improving the accuracy of anomaly detection.
[0128] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A memory network-based image anomaly detection method, characterized in that: The image anomaly detection method based on memory network includes: Constructing an image anomaly detection network model, wherein the image anomaly detection network model includes a first encoder, a memory network, and at least two decoders, wherein the first encoder uses a neural network VGG-16; The constructed image anomaly detection network model is trained using a normal image training data set, the training sample is input into the first encoder to extract high-dimensional features, the mapping features matching the high-dimensional features are queried in the memory network, and then the mapping features are respectively input into the decoder to reconstruct the image, the reconstructed image with the smallest covariance value with the original training sample is taken as the output reconstructed image, the joint loss is calculated to update the parameters of the image anomaly detection network model, and the training is completed; A lightweight second encoder based on the first encoder is extracted through knowledge distillation, the maximum pooling layer in the last four convolutional blocks of the first encoder is passed to the second encoder as a knowledge distillation layer, and the first encoder in the image anomaly detection network model is replaced by the second encoder to generate the final image anomaly detection network model; Input the image to be detected into the final image anomaly detection network model, output the reconstructed image, calculate the anomaly detection scores of the input image to be detected and the reconstructed image, and determine whether the input image to be detected is abnormal; The step of calculating the abnormality detection scores of the input image to be detected and the reconstructed image includes: Calculate the L2 distance between each query feature vector item after the image to be detected passes through the second encoder and the best matching feature vector item in the memory network: Where Q represents the number of query feature vector items, q t represents the query feature vector item, p s Represents the best matching prototype feature vector item in the memory network; Calculate the peak signal-to-noise ratio of the image to be detected and the reconstructed image: Where N is the number of pixels in the image to be detected, x represents the image to be detected, represents the reconstructed image, It means to find the best reconstructed image; The L2 distance and peak signal-to-noise ratio are normalized, and then the weighted sum of the two is calculated as the anomaly detection score.
2. The image anomaly detection method based on memory network according to claim 1, characterized in that: Based on the neural network VGG-16, the second encoder removes the last layer of convolution of the last three convolution blocks, discards the last fully connected layer of VGG-16, and passes the maximum pooling layer of the last four convolution blocks of the first encoder VGG-16 as the knowledge distillation layer to the last four convolution blocks of the second encoder.
3. The image anomaly detection method based on memory network according to claim 1, characterized in that: The step of searching the memory network for a mapping feature that matches the high-dimensional feature includes: The high-dimensional features extracted by the first encoder are used as a query feature vector item set of the memory network, and each feature vector item in the high-dimensional features is used as a query feature vector item; The matching probability between each query feature vector item and all prototype feature vector items stored in the memory network is calculated, and then the weighted average of the prototype feature vector items and their corresponding matching probabilities is calculated as the queried feature vector item, and all the queried feature vector items are combined into a mapping feature that matches the input high-dimensional feature.
4. The image anomaly detection method based on memory network according to claim 3 is characterized in that: The calculation formula for calculating the matching probability between each query feature vector item and all prototype feature vector items stored in the memory network is as follows: Among them, w t,m is the calculated matching probability, exp is an exponential function with the natural constant e as the base, p m represents the prototype eigenvector term, q t Represents the query feature vector item, and M represents the number of prototype feature vector items stored in the memory network.
5. The image anomaly detection method based on memory network according to claim 1, characterized in that: The memory network-based image anomaly detection method further includes: The high-dimensional features extracted by the first encoder are used as a query feature vector item set of the memory network, and each feature vector item in the high-dimensional features is used as a query feature vector item; Calculate the matching probability v between each prototype feature vector item stored in the memory network and all query feature vector items t,m : Among them, p m represents the prototype eigenvector term, q t represents the query feature vector item, Q is the number of query feature vector items; The matching probability v t,m Standardize to get v′ t,m , the standardized formula is as follows: Finally, the prototype feature vector item is updated by the following formula: p m =f(p m +∑ t∈Q v′ t,m q t ); where f() is the L2 function.
6. The image anomaly detection method based on memory network according to claim 1, characterized in that: The memory network-based image anomaly detection method further includes: Calculate the input image x and the output image The weighted reconstruction error between is taken as the conventional score εt, which is calculated as follows: Among them, W t (.) is the weight function, and the calculation formula is as follows: When the score εt is higher than a threshold γ, x is regarded as an abnormal image and is not used to update the prototype feature vector item in the memory network. Otherwise, it is used to update the prototype feature vector item in the memory network.
7. The image anomaly detection method based on memory network according to claim 1, characterized in that: The weighted sum of the two is calculated as the anomaly detection score, and the calculation formula is as follows: Among them, g(.) is the normalization operation, λ is the weight coefficient, S t Represents the calculated anomaly detection score.
Citation Information
Patent Citations
Multi-mode two-stage unsupervised video anomaly detection method
CN114332053A
Method and means for detection of imperfections in products
SE1930421A1