Adaptive memory matching multi-class anomaly detection method
By constructing a generator network of memory matching mechanism and a dual receptive field feature fusion network, combined with an encoder-decoder-encoder structure discriminator network, the problem of insufficient quality of robust characterization and reconstruction in multi-class anomaly detection is solved, and efficient multi-class anomaly detection effect is achieved.
Patent Information
- Application Number
- CN202510315224.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
AI Technical Summary
When existing multi-category anomaly detection methods deal with complex and high-dimensional data, it is difficult to effectively learn robust representations and distinguish subtle differences between normal and abnormal, and the quality of generative model reconstruction is insufficient, resulting in poor detection results in practical applications.
Build a generator network with memory matching mechanism and a dual receptive field feature fusion network, and combine the encoder-decoder-encoder structure discriminator network to improve the performance of multi-class anomaly detection through feature selection, information compression and image reconstruction.
It improves the accuracy and efficiency of multi-category anomaly detection, can effectively handle multi-category anomaly situations, generate high-quality reconstructed data, and adapt to the detection needs of complex and high-dimensional data.
Smart Images

Figure BDA0005315758650000031 
Figure BDA0005315758650000032 
Figure BDA0005315758650000033
Abstract
Description
Technical Field
[0001] The present invention belongs to multiple types of anomaly detection methods and can be applied to the field of industrial defect detection, involving technical fields such as memory matching mechanisms, computer vision, and image feature extraction. Background Art
[0002] Anomaly detection, also known as outlier detection, aims to identify data points in a dataset that are significantly different from the normal pattern. In many fields such as industrial quality inspection, medical image analysis, and network security, anomaly detection plays a crucial role. Traditional anomaly detection methods usually rely on manually defined features or statistical models, but these methods often show limitations when dealing with complex, high-dimensional data. In recent years, with the rise of deep learning technology, anomaly detection methods based on deep generative models have made significant progress. However, existing deep learning-based anomaly detection methods still face some challenges. For example, how to effectively learn robust representations from complex normal data and distinguish the subtle differences between normal and abnormal remains a difficult problem. Many methods usually only focus on learning single-category normal data during anomaly detection and are difficult to handle multi-category anomaly situations that may occur in practical applications. How to ensure that the reconstructed data generated by the generative model has high quality is also a key factor affecting the performance of anomaly detection.
[0003] To solve the above problems, researchers are actively exploring new methods. Some studies attempt to improve the model's ability to model normal patterns by introducing attention mechanisms or memory networks, while other studies focus on designing more effective feature extraction networks or more robust loss functions. However, there is still a lack of an adaptive anomaly detection method that can effectively handle multi-category anomaly situations and at the same time ensure the reconstruction quality of the generative model. Aiming at the deficiencies of the above existing technologies, the present invention proposes an adaptive memory matching multi-class anomaly detection method. This method constructs a generator network with a memory matching mechanism to complete the feature selection of different deep memory features and assist the decoding network in the generator to achieve high-quality decoding of the latent vector. At the same time, a dual receptive field feature fusion network is used to reconstruct the information compression of the image and assist in training the latent vector in the generator network. An encoder-decoder-encoder structure discriminator network is constructed to enhance the discrimination ability. By integrating these components, an adaptive memory matching multi-class anomaly detection method architecture is formed, thereby effectively improving the performance of multi-class anomaly detection. Summary of the Invention
[0004] The object of the present invention is to solve the problem that the anomaly detection effect of multi-class anomaly detection in the current research methods is not good, resulting in the difficulty of generalizing the existing anomaly detection methods to practical application fields such as industrial defect detection. Aiming at the problems of insufficient feature selection ability in the current neural network architecture and single feature extraction path for images, the present invention respectively proposes a generator network with a memory matching mechanism and a dual receptive field feature fusion network to further enhance the feature extraction ability of anomaly detection information, and at the same time expand the feature extraction path, so as to achieve the purpose of realizing high-performance multi-class anomaly detection.
[0005] The technical solution adopted by the present invention to solve the above technical problems is:
[0006] S1. Construct a generator network with a memory matching mechanism to complete the feature selection of different deep memory features and assist the decoding network in the generator to achieve high-quality decoding of the latent vector.
[0007] S2. Construct a dual receptive field feature fusion network to reconstruct the information compression of the image and assist in training the latent vector in the generator network.
[0008] S3. Construct an encoder-decoder-encoder structure discriminator network to distinguish the authenticity of the input image and the generated image.
[0009] S4. Combine the generator network with a memory matching mechanism in S1, the dual receptive field feature fusion network in S2, and the encoder-decoder-encoder structure discriminator network in S3 to construct an adaptive memory matching multi-class anomaly detection method architecture.
[0010] S5. Training of an adaptive memory matching multi-class anomaly detection method architecture and multi-class anomaly detection test.
[0011] 1. According to an adaptive memory matching multi-class anomaly detection method as described in claim 1, characterized in that the specific process of S1 is:
[0012] The generator network with a memory matching mechanism compresses the image by an encoder network and a multi-layer memory network, completes the feature selection of different deep memory features by the memory matching mechanism, and completes the high-quality decoding of the compressed image information (latent vector) by the decoder network.
[0013] The encoder network is composed of a combination block of three layers of convolution, batch normalization and activation function. After passing through the encoder network, the input image with an original image dimension of H×W×C x is compressed to Denoted as query. Subsequently, query is input into a multi-layer memory network to extract memory features. In the multi-layer memory network, I is used to represent the memory pool. A total of P = 10 memory pools are set in this network. Each memory pool contains N = 16 memory items MemItems, and each memory item is denoted as mitem n . The multi-layer memory network takes query as the initial node and obtains the next memory node by reading the memory items in each memory pool. After reading, the memory items in the memory pool are updated, and finally a multi-level memory network structure is generated. Specifically, query obtains the memory node MemNode1 by reading the memory items in memory pool I1, and obtains the memory node MemNode2 by reading the memory items in memory pool I2; MemNode1 reads memory pools I3 and I4 respectively to obtain memory nodes MemNode3 and MemNode4; MemNode2 reads memory pools I5 and I6 respectively to obtain memory nodes MemNode5 and MemNode6; finally, MemNode3, MemNode4, MemNode5, and MemNode6 read memory pools I7, I8, I9, and I10 respectively to obtain memory nodes MemNode7, MemNode8, MemNode9, and MemNode10. All memory nodes together constitute the multi-layer memory network structure. In this process, the reading method and the updating method are the same, and the sizes of the memory nodes obtained by the reading operation are all Here, taking query reading MemNode1 as an example, the reading and updating methods of the memory node are introduced.
[0014] Reading: When using query to read MemNode1, query will be preprocessed into vectors along the channel direction, and the dimension of each vector is 1×1×C. q k represents the k-th segmented vector. During the reading process, each q k is used to read all the memory items in memory pool I1 to represent each q k with memory items. Since matrix calculations are used in the neural network, first calculate the cosine similarity between the query composed of all q k and all memory items MemItems to construct a correlation matrix of size N×K. Subsequently, the softmax function is applied along the vertical direction in the correlation matrix to probabilize the values, obtaining the matching probability MatPrV.
[0015]
[0016] Through formula (1), each q kThe matching probabilities with all memory items are regarded as the weight parameters for each memory item. All weight parameters are multiplied by the corresponding memory items and summed to obtain the mnode at the corresponding position of q in MemNode1 k in the corresponding position of mnode k .
[0017] Subsequently, the matching probability is used as the weight coefficient to re - represent Q. Specifically, each vector is transformed into:
[0018]
[0019] Through formula (2), all weight parameters are multiplied by the corresponding memory items and summed to obtain the mnode at the corresponding position of q in MemNode1 k in the corresponding position of mnode k , and then the K mnodes k are concatenated to obtain MemNode1.
[0020] Update: The core principle in the update process is to use all q k to perform correlation ranking on each memory item and establish an Index as an index record. Imitating the calculation of the matching probability along the vertical direction in formula (1), in the update process, the correlation between each memory item mitem n and all q k is calculated to obtain the following matching probability graph.
[0021]
[0022] Formula (4) normalizes the calculated matching probability, and finally each memory item is updated through formula (5). The update of the memory item does not change the original dimension. Finally, the information of these memory nodes is fused with the query to obtain the multi - level memory feature fusion information denoted as z.
[0023] In the multi - layer memory network, the memory features are divided into three layers according to the reading level. The first layer includes MemNode1 and MemNode2; the second layer includes MemNode3, MemNode4, MemNode5, and MemNode6; the third layer includes MemNode7, MemNode8, MemNode9, and MemNode10. An adaptive memory matching method is proposed for each level to select important memory feature blocks. In this process, two of the most important feature blocks are selected and upsampled for feature fusion with the subsequent decoder network, and the multi - level memory information is used to guide the decoder network to decode the compressed information z. There are two in total in the first layer, and memory matching is not required for this layer. Both the second and third - layer memory features contain four memory feature blocks, and the process of the memory matching mechanism is as follows.
[0024] Taking the second - layer features as an example, first, a global pooling operation is performed on the memory block features (Mem2Level) with a size of , which compresses them into a vector with a size of 1×1×4C. Each numerical point can reflect the global information of each memory block sliced along the channel direction.
[0025] Mem2Level
[0026] = Concat(MemNode3, MemNode4, MemNode5, MemNode6) (6)
[0027]
[0028] Subsequently, this vector is input into two fully - connected layers to further learn the importance of the memory features for the decoded information. Here, a compression ratio r = 16 is used to compress and restore the vector dimension. Then, the activation function tanh is used to adjust the importance of its numerical points. Finally, this activated vector is grouped in the channel direction, divided into 1 - C, C + 1 - 2C, 2C + 1 - 3C, 3C + 1 - 4 according to the four memory blocks respectively. Then, the matching values are calculated according to the grouping. After the top two memory block features with the highest calculated values are feature - aligned with the decoder structure unit through up - sampling, they are used in the decoding process of the decoder network to guide image reconstruction.
[0029] In the decoder network, four stacked transposed - convolution blocks are used to upsample the compressed information z until the original size of the image is restored. In the memory matching mechanism, three - layer memory matching features, namely Mem1Level, Mem2Level, and Mem3Level, are collected. The selected memory block features are fused with each layer of the decoding network in a feature - concatenation manner. The fusion process is as follows.
[0030] Decoder 256 = ConvTrans(BN(RELU(z))) (8)
[0031] Decoder 128 = ConvTrans(BN(RELU(Concat(Decoder 256 ,Mem1Level)))) (9)
[0032] Decoder 64 = ConvTrans(BN(RELU(Concat(Decoder 128 ,Mem2Level)))) (10)
[0033]
[0034] The decoder finally restores the compressed information z to the reconstructed image Under the action of the memory matching mechanism, the decoder can understand the prototype pattern of normal images, so as to better restore the images.
[0035] 2. An adaptive memory matching multi-class anomaly detection method according to claim 1, wherein the specific process of S2 is as follows:
[0036] A dual receptive field feature fusion network is constructed, and in this network, the local features and multi-scale features of the reconstructed image are fused in three steps. The local features are extracted using a 3×3 convolutional kernel; the multi-scale features have a larger receptive field than the local features to obtain the correlation relationship between the local blocks of the image. The multi-scale feature F MCB is extracted by the multi-scale convolutional block, and the multi-scale convolutional block MCB fuses the coarse-grained feature F coarse , the channel feature F channel and the dilated convolutional feature F dilated . The coarse-grained feature F coarse enhances the sensitivity to the coarse-grained features, establishes the relationship between these features, and promotes the integration of global context information. The channel feature F channel is used to compensate for the loss of channel-specific information that may occur during feature fusion. The dilated convolutional feature F dilated further expands the receptive field of the input feature map, thereby establishing the connection between pixel blocks. The following takes the first step as an example to introduce the feature extraction process of MCB.
[0037]
[0038] Subsequently, the features are concatenated along the channel direction, and 4×4 convolution is used for feature extraction to obtain the fused feature F fusion .
[0039]
[0040] The fused features are further refined by 1×1 convolution to make them more suitable for the anomaly detection task.
[0041] F out = LeakyReLU(Conv 1×1 (F fusion )) (16)
[0042] The above is the specific implementation process of the multi-scale convolutional block MCB. In each step, the multi-scale features extracted by MCB will be fused with the detailed features F local extracted by 3×3 convolution, and the fusion method is as follows.
[0043] F dual = F out + F local (17)
[0044] After three-step fusion, the reconstructed image is finally compressed into a reconstructed compressed information with the same dimension as the compressed information z The image information of the double receptive field is compressed in the reconstructed compressed information. By calculating the loss with the compressed information z, the training objective function is optimized.
[0045] 3. An adaptive memory matching multi-class anomaly detection method according to claim 1, wherein the specific process of S3 is as follows:
[0046] Designed with reference to the architecture of the generator network, an encoder-decoder-encoder structure discriminator network D is constructed. This discriminator network is divided into three sub-networks: an encoder network Encoder1, a decoder network, and an encoder network Encoder2. The discriminator network inputs the input image and the reconstructed image respectively to feedback the authenticity of the image.
[0047] The encoder network Encoder1 is composed of a combination block of four layers of convolution, batch normalization, and LeakyReLU activation function, and sequentially converts the input with dimensions of H×W×C x into where nz is the dimension of the compressed information z. The decoder network is composed of a combination block of four layers of transposed convolution, batch normalization function, and tanh activation function, and upsamples the 1×1×nz to the output with dimensions of H×W×C x ; The encoder network Encoder2 is composed of a combination block of four layers of convolution, batch normalization, and LeakyReLU activation function. Different from the encoder network Encoder1, the output of this network is a judgment on the authenticity of the content input to the discriminator rather than the compressed information z.
[0048] 4. An adaptive memory matching multi-class anomaly detection method according to claim 1, wherein the specific process of S4 is as follows:
[0049] In S1, S2, and S3, a generator network with a memory matching mechanism, a double receptive field feature fusion network, and an encoder-decoder-encoder structure discriminator network are respectively constructed. On this basis, three loss functions are set to train the architecture of the adaptive memory matching multi-class anomaly detection method.
[0050] The first loss function is the adversarial loss L for training the generative adversarial network adv, through the setting of this loss function, the generator can capture the data distribution p of the training samples x x , so that the generator can generate a reconstructed sample G(x) with a similar data distribution. The adversarial loss L adv feeds back the training of the network through the prediction results of the discriminator by inputting the training sample x and the reconstructed sample G(x) into the discriminator.
[0051]
[0052] The second loss function is the context loss function L between the reconstructed sample G(x) and the training sample con , which constrains the distance between samples through the L1 distance, so as to generate a reconstructed sample G(x) with similar structure.
[0053]
[0054] The third loss function is the latent vector loss function. The generator network with the memory matching mechanism and the dual receptive field feature fusion network will compress the training sample x and the reconstructed sample G(x) into compressed information z and further improve the reconstruction effect of the reconstructed sample by constraining the distance in the compressed information feature space.
[0055]
[0056] According to the above three loss functions, determine the training objective L of a final adaptive memory matching multi-class anomaly detection method, and each loss function is assigned by a different weight w:
[0057] L = w adv L adv + w con L con + w dis-z L dis-z (21)
[0058] 5. An adaptive memory matching multi-class anomaly detection method according to claim 1, wherein the specific process of S5 is:
[0059] The architecture of an adaptive memory matching multi-class anomaly detection method is deployed in the environment of Pytorch 1.0 and optimized using the Adam optimizer. Experiments are conducted on a multi-class anomaly detection dataset. During the training phase, the method is input with 32×32 images for training. Through the generator network with a memory matching mechanism, the images are first compressed into compressed information z with memory features, and then z is restored to the original image size of 32×32 to complete image reconstruction. During training, training samples and reconstructed samples are input into the discriminator network with an encoder-decoder-encoder structure to promote the training of the entire generative adversarial network.
[0060] The specific implementation details during training are as follows: The weight parameter of the training objective is w con as the main training loss, set to 50, and other weight parameters are set to 1. The learning rate is set to 0.003. The multi-class anomaly detection dataset contains 10 classes, and each class is used as the anomaly class for experiments in turn, with each class trained for 15 epochs. The dimension of the compressed information is set to 512, and the training batch size is set to 64.
[0061] Compared with the existing technologies, the beneficial effects of the present invention are:
[0062] 1. The present invention proposes an adaptive memory matching multi-class anomaly detection method. The network uses a generator network with a memory matching mechanism to extract prototype features of anomaly detection images using the memory matching mechanism, and fuses the memory features with the decoder network through feature fusion.
[0063] 2. The present invention proposes a dual receptive field feature fusion network. This method overcomes the traditional single-path feature extraction method, extracts features of the reconstructed image through two paths to obtain a larger receptive field, and also proposes a corresponding loss function to assist image reconstruction, thereby enriching the representation of the compressed information of the input image. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is a schematic structural diagram of an adaptive memory matching multi-class anomaly detection method.
[0065] Figure 2 and Figure 3 It is a result comparison diagram of multi-class anomaly detection of an adaptive memory matching multi-class anomaly detection method.
[0066] Figure 4 It is a result comparison diagram of image defect detection of an adaptive memory matching multi-class anomaly detection method.
[0067] Figure 5 It is a visualization result diagram of image defect detection of an adaptive memory matching multi-class anomaly detection method. Detailed implementation manners
[0068] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the patent.
[0069] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0070] Figure 1 It is a schematic structural diagram of an adaptive memory matching multi-class anomaly detection method. As Figure 1 shown, the whole structure includes three parts: a generator network with a memory matching mechanism, a dual receptive field feature fusion network, and an encoder-decoder-encoder structure discriminator network.
[0071] In Figure 1 , first, the input image is fed into the generator network with a memory matching mechanism. The generator network includes an encoder network, a multi-layer memory network, and a decoder network. The encoder network is composed of a combination block of three layers of convolution, batch normalization, and activation functions. After passing through the encoder network, the original input image with dimensions of H×W×C x is compressed to denoted as query. Subsequently, the query is input into the multi-layer memory network for extracting memory features. In the multi-layer memory network, I is used to represent the memory pool. A total of P = 10 memory pools are set in this network, and each memory pool contains N = 16 memory items MemItems, and each memory item is denoted as mitem n . The multi-layer memory network uses the query as the initial node, and obtains the next memory node by reading the memory items in each memory pool. After reading, the memory items in each memory pool are updated, and finally a multi-level memory network structure is generated. Specifically, the query obtains the memory node MemNode1 by reading the memory items in the memory pool I1, and obtains the memory node MemNode2 by reading the memory items in the memory pool I2; MemNode1 reads the memory pools I3 and I4 respectively to obtain the memory nodes MemNode3 and MemNode4; MemNode2 reads the memory pools I5 and I6 respectively to obtain the memory nodes MemNode5 and MemNode6; finally, MemNode3, MemNode4, MemNode5, and MemNode6 read the memory pools I7, I8, I9, and I10 respectively to obtain the memory nodes MemNode7, MemNode8, MemNode9, and MemNode10. All the memory nodes together constitute the multi-layer memory network structure. In this process, the reading method and the updating method are the same, and the sizes of the memory nodes obtained by the reading operation are all The following takes the query reading MemNode1 as an example to introduce the reading and updating methods of the memory node.
[0072] Reading: When querying to read MemNode1, the query will be preprocessed into vectors along the channel direction, and the dimension of each vector is 1×1×C. q k represents the k-th segmented vector. During the reading process, each q k is used to read all the memory items in the memory pool I1 to represent each q k in terms of memory items. Since matrix calculations are used in the neural network, first calculate the cosine similarity between the query composed of all q k and all memory items MemItems to construct a correlation matrix of size N×K. Subsequently, apply the softmax function along the vertical direction in the correlation matrix to probabilize the values and obtain the matching probability MatPrV.
[0073]
[0074] After formula (1), the matching probabilities between each q k and all memory items are calculated, and these matching probabilities are regarded as the weight parameters of each memory item. Multiply all the weight parameters by the corresponding memory items and sum them to obtain the mnode k at the corresponding position in MemNode1 for q k .
[0075] Subsequently, use the matching probability as the weight coefficient to re-represent Q. Specifically, each of its vectors is transformed into:
[0076]
[0077] After formula (2), multiply all the weight parameters by the corresponding memory items and sum them to obtain the mnode k at the corresponding position in MemNode1 for q k , and then splice the K mnode k to obtain MemNode1.
[0078] Update: The core principle in the update process is to use all q k to sort the relevance of each memory item and establish an Index as an index record. Imitating the calculation of the matching probability along the vertical direction in formula (1), calculate the relevance between each memory item mitem n and all q k along the horizontal direction in the update process to obtain the following matching probability graph.
[0079]
[0080] Formula (4) normalizes the calculated matching probability, and finally completes the update of each memory item through formula (5). The update of the memory item does not change the original dimension. Finally, the information of these memory nodes is fused with the query to obtain the multi-level memory feature fusion information denoted as z.
[0081] In the multi-level memory network, the memory features are divided into three layers according to the reading levels. The first layer includes MemNode1 and MemNode2; the second layer includes MemNode3, MemNode4, MemNode5, and MemNode6; the third layer includes MemNode7, MemNode8, MemNode9, and MemNode10. An adaptive memory matching method is proposed for each level to select important memory feature blocks. In this process, two of the most important feature blocks are selected and upsampled for feature fusion with the subsequent decoder network. The multi-level memory information is used to guide the decoder network to decode the compressed information z. There are two in total in the first layer, and memory matching is not required for this layer. Both the second and third layer memory features contain four memory feature blocks, and the process of the memory matching mechanism is as follows.
[0082] Taking the second layer features as an example, first, for the memory block features (Mem2Level) with a size of a global pooling operation is performed to compress it into a vector with a size of 1×1×4C. Each numerical point can reflect the global information of the memory block sliced along the channel direction for each layer.
[0083] Mem2Level
[0084] = Concat(MemNode3, MemNode4, MemNode5, MemNode6) (6)
[0085]
[0086] Subsequently, this vector is input into two fully connected layers to further learn the importance of the memory features for the decoded information. Here, a compression ratio r = 16 is used to compress and restore the vector dimension. Then, the activation function of tanh is used to adjust the importance of its numerical points. Finally, the activation vector is grouped in the channel direction and divided into 1-C, C+1-2C, 2C+1-3C, 3C+1-4 according to the four memory blocks respectively. Then, the matching values are calculated according to the grouping. The two memory block features with the highest calculated values are used for feature alignment with the decoder structural unit through upsampling and are used in the decoding process of the decoder network to guide image reconstruction.
[0087] In the decoder network, four stacked transposed convolution blocks are used to upsample the compressed information z until the original size of the image is restored. In the memory matching mechanism, three levels of memory matching features are collected, namely Mem1Level, Mem2Level, and Mem3Level. The selected memory block features are fused with each layer of the decoding network by feature concatenation, and the fusion process is as follows.
[0088] Decoder 256 = ConvTrans(BN(RELU(z))) (8)
[0089] Decoder 128 = ConvTrans(BN(RELU(Concat(Decoder 256 ,Mem1Level)))) (9)
[0090] Decoder 64 = ConvTrans(BN(RELU(Concat(Decoder 128 ,Mem2Level)))) (10)
[0091]
[0092] The decoder finally restores the compressed information z to the reconstructed image Under the action of the memory matching mechanism, the decoder can understand the prototype pattern of the normal image, so as to better restore the image.
[0093] As Figure 1 shown, this method constructs a dual receptive field feature fusion network. In this network, three steps are used to fuse the local features and multi-scale features of the reconstructed image. The local features are extracted using a 3×3 convolutional kernel; the multi-scale features have a larger receptive field than the local features to obtain the correlation between local blocks of the image. The multi-scale feature F MCB is extracted by a multi-scale convolutional block, and the multi-scale convolutional block MCB fuses the coarse-grained feature F coarse , the channel feature F channel and the dilated convolutional feature F dilated . The coarse-grained feature F coarse enhances the sensitivity to the coarse-grained features, establishes the relationship between these features, and promotes the integration of global context information. The channel feature F channel is used to compensate for the loss of channel-specific information that may occur during feature fusion. The dilated convolutional feature F dilated further expands the receptive field of the input feature map, thus establishing the connection between pixel blocks. The feature extraction process of MCB is introduced below taking the first step as an example.
[0094]
[0095] Subsequently, the features are concatenated along the channel direction, and 4×4 convolution is used for feature extraction to obtain the fused feature F fusion 。
[0096]
[0097] The fused feature is further refined through 1×1 convolution to make it more suitable for the anomaly detection task.
[0098] F out =LeakyReLU(Conv 1×1 (F fusion )) (16)
[0099] The above is the specific implementation process of the multi-scale convolution block MCB. In each step, the multi-scale features extracted by MCB will be fused with the detailed feature F extracted by 3×3 convolution local as follows.
[0100] F dual =F out +F local (17)
[0101] After three-step fusion, finally, the reconstructed image is compressed into the reconstructed compressed information with the same dimension as the compressed information z The reconstructed compressed information compresses the image information of the double receptive field. By calculating the loss with the compressed information z, the training objective function is optimized.
[0102] Designed with reference to the architecture of the generator network, a discriminator network D with an encoder-decoder-encoder structure is constructed. The discriminator network is divided into three sub-networks: the encoder network Encoder1, the decoder network, and the encoder network Encoder2. The discriminator network inputs the input image and the reconstructed image respectively to feedback the authenticity of the image.
[0103] The encoder network Encoder1 is composed of a combination block of four layers of convolution, batch normalization, and LeakyReLU activation function, and sequentially transforms the input with dimension H×W×C x into where nz is the dimension of the compressed information z. The decoder network is composed of a combination block of four layers of transposed convolution, batch normalization function, and tanh activation function, and upsamples 1×1×nz to the dimension of H×W×C xOutput; The encoder network Encoder2 is composed of a combined block of four layers of convolution, batch normalization, and the LeakyReLU activation function. Different from the encoder network Encoder1, the output of this network is a judgment on the authenticity of the content of the input discriminator rather than the compressed information z.
[0104] Figure 2 and Figure 3 is a result comparison chart of multi-class anomaly detection for an adaptive memory matching multi-class anomaly detection method. Figure 2 Compared with previous advanced methods on the MNIST dataset, it can be seen from the figure that an adaptive memory matching multi-class anomaly detection method achieves the best results in nine out of the ten classes 0-9 in MNIST. Similarly, Figure 3 This method was compared on the CIFAR10 dataset and achieved the best results in ten different natural classes of the CIFAR10 dataset.
[0105] Figure 4 is a result comparison chart of image defect detection for an adaptive memory matching multi-class anomaly detection method. This method was compared with SPADE using a pre-trained network, PaDiM using a large number of parameters, and WinCLIP incorporating a cross-modal model of text and images in the VisA industrial defect detection target dataset. It can be seen from the comparison results that the present invention achieves excellent results in the VisA dataset.
[0106] Figure 5 is a visualization result chart of image defect detection for an adaptive memory matching multi-class anomaly detection method in the VisA dataset. The present invention selected eight target classes in the VisA dataset for visualization. The defects of the target are marked by the upper frame, and the lower part shows the repair situation of the defective parts of the present invention for this dataset. From Figure 5 it can be seen that the present invention can repair the image defect area and achieve image defect detection.
[0107] The present invention proposes an adaptive memory matching multi-class anomaly detection method, which is divided into three parts: a generator network with a memory matching mechanism, a dual receptive field feature fusion network, and an encoder-decoder-encoder structure discriminator network. The generator network with a memory matching mechanism consists of an encoder network, a multi-layer memory network, and a decoder network. The encoder network and the multi-layer memory network are used to compress the image, the memory matching mechanism is used to select features of different deep memory features, and the decoder network is used to perform high-quality decoding of the image compression information (latent vector). The dual receptive field feature fusion network uses three steps to fuse the local features and multi-scale features of the reconstructed image to obtain the correlation relationship between local blocks of the image. The encoder-decoder-encoder structure discriminator network is designed with reference to the architecture of the generator network, and the discriminator network inputs the input image and the reconstructed image respectively to feedback the authenticity of the image. Finally, the present invention conducts multi-class anomaly detection experiments on the MNIST and CIFAR10 datasets. The results show that this method has great advantages over previous advanced methods. At the same time, the visualization results on the VisA industrial defect target detection dataset also demonstrate the strong recovery ability of the present invention for defect parts.
[0108] Finally, the details of the above examples of the present invention are only examples for explaining the present invention. For those skilled in the art, any modifications, improvements, substitutions, etc. to the above embodiments shall be included in the protection scope of the claims of the present invention.
Claims
1. An adaptive memory matching multi-class anomaly detection method, characterized in that, The method includes the following steps: S1. Construct a generator network with a memory matching mechanism to complete the feature selection of different deep memory features, and assist the decoding network in the generator to achieve high-quality decoding of the latent vector; S2. Construct a dual receptive field feature fusion network for information compression of the reconstructed image and assist in training the latent vector in the generator network; S3. Construct an encoder-decoder-encoder structure discriminator network to distinguish the authenticity of the input image and the generated image; S4. Combine the generator network with a memory matching mechanism in S1, the dual receptive field feature fusion network in S2, and the encoder-decoder-encoder structure discriminator network in S3 to construct an adaptive memory matching multi-class anomaly detection method architecture; S5. Training of an adaptive memory matching multi-class anomaly detection method architecture and multi-class anomaly detection tests.
2. The adaptive memory matching multi-class anomaly detection method according to claim 1, wherein The specific process of S1 is as follows: The generator network with a memory matching mechanism compresses the image by the encoder network and the multi-layer memory network, selects the features of different deep memory features by the memory matching mechanism, and completes the high-quality decoding of the image compression information (latent vector) by the decoder network. The encoder network consists of a combination of three layers of convolution, batch normalization, and activation functions. After passing through the encoder network, the original input image with dimensions H×W×C x is compressed into denoted as query. Subsequently, the query is input into a multi-layer memory network to extract memory features. In the multi-layer memory network, I is used to represent the memory pool. A total of P = 10 memory pools are set in this network. Each memory pool contains N = 16 memory items MemItems, and each memory item is denoted as mitem n . The multi-layer memory network takes the query as the initial node and obtains the next memory node by reading the memory items in each memory pool. After reading, the memory items in each memory pool are updated, and finally a multi-level memory network structure is generated. Specifically, the query obtains the memory node MemNode1 by reading the memory items in the memory pool I1, and obtains the memory node MemNode2 by reading the memory items in the memory pool I2; MemNode1 reads the memory pools I3 and I4 respectively to obtain the memory nodes MemNode3 and MemNode4; MemNode2 reads the memory pools I5 and I6 respectively to obtain the memory nodes MemNode5 and MemNode6; finally, MemNode3, MemNode4, MemNode5, and MemNode6 read the memory pools I7, I8, I9, and I10 respectively to obtain the memory nodes MemNode7, MemNode8, MemNode9, and MemNode10. All the memory nodes together form a multi-layer memory network structure. In this process, the reading method and the updating method are the same, and the dimensions of the memory nodes obtained by the reading operation are all Next, take the query reading MemNode1 as an example to introduce the reading and updating methods of the memory node. Reading: When query is used to read MemNode1, the query will be preprocessed into vectors along the channel direction, and the dimension of each vector is 1×1×C. q k represents the k-th segmented vector. During the reading process, each q k is used to read all memory items in the memory pool I1 to represent each q k with memory items. Since matrix calculations are used in the neural network, first calculate the cosine similarity between the query composed of all q k and all memory items MemItems to construct a correlation matrix of size N×K. Subsequently, apply the softmax function along the vertical direction in the correlation matrix to probabilize the values and obtain the matching probability MatPrV. Through formula (1), the matching probability of each q with all memory items is calculated, and these matching probabilities are regarded as the weight parameters of each memory item. All weight parameters are multiplied by the corresponding memory items and summed to obtain the mnode at the corresponding position of q in MemNode1. k k k Subsequently, the matching probability is used as a weight coefficient to re-represent Q. Specifically, each of its vectors is converted to: Through formula (2), all weight parameters are multiplied by the corresponding memory items and summed to obtain the mnode at the corresponding position in MemNode1 for q k corresponding to k , and then the K mnodes k are concatenated to obtain MemNode1. Update: The core principle during the update process is to utilize all q k to perform relevance ranking on each memory item and establish an Index as the index record. Imitating the calculation of the matching probability along the vertical direction in formula (1), during the update process, calculate the relevance between each memory item mitem n and all q k to obtain the following matching probability graph. Formula (4) normalizes the calculated matching probability, and finally completes the update of each memory item through formula (5). The update of the memory item does not change the original dimension. Finally, the information of these memory nodes is fused with the query to obtain multi-level memory feature fusion information denoted as z. In the multi-layer memory network, the memory features are divided into three layers according to the reading level. The first layer includes MemNode1 and MemNode2; the second layer includes MemNode3, MemNode4, MemNode5, and MemNode6; the third layer includes MemNode7, MemNode8, MemNode9, and MemNode10. An adaptive memory matching method is proposed for each layer to select important memory feature blocks. In this process, two of the most important feature blocks are selected and upsampled for feature fusion with the subsequent decoder network, and the multi-level memory information is used to guide the decoder network to decode the compressed information z. There are two in total in the first layer, and memory matching is not required in this layer. Both the second and third layer memory features contain four memory feature blocks, and the process of the memory matching mechanism is as follows. Taking the second-layer features as an example, first, perform a global pooling operation on the memory block features (Mem2Level) with a size of to compress it into a vector with a size of 1×1×4C. Each numerical point can reflect the global information of the slices of each layer of memory blocks along the channel direction. Subsequently, the vector is input into two fully connected layers to further learn the importance of the memory features for the decoded information. Here, the compression ratio r = 16 is used to compress and restore the vector dimension. Then, the activation function of tanh is used to adjust the importance of its numerical points. Finally, the activation vector is grouped in the channel direction and divided into 1-C, C+1-2C, 2C+1-3C, 3C+1-4 according to the four memory blocks respectively. Then, the matching values are calculated according to the grouping, and the features of the top two memory blocks with the highest calculated values are used for feature alignment with the decoder structure unit through upsampling and are used in the decoding process of the decoder network to guide image reconstruction. In the decoder network, four stacked transposed convolution blocks are adopted to upsample the compressed information z until the original size of the image is restored. In the memory matching mechanism, three levels of memory matching features are collected, namely Mem1Level, Mem2Level, and Mem3Level. The features of the selected memory blocks are fused with each layer of the decoding network by feature concatenation, and the fusion process is as follows. Decoder 256 = ConvTrans(BN(RELU(z))) (8) Decoder 128 = ConvTrans(BN(RELU(Concat(Decoder 256 , Mem1Level)))) (9) Decoder 64 = ConvTrans(BN(RELU(Concat(Decoder 128 , Mem2Level)))) (10) The decoder finally restores the compressed information z to the reconstructed image Under the action of the memory matching mechanism, the decoder can understand the prototype pattern of normal images, so as to better restore the images.
3. An adaptive memory matching multi-class anomaly detection method according to claim 1, characterized in that The specific process of S2 is as follows: A dual receptive field feature fusion network is constructed, and in this network, three steps are adopted to fuse the local features and multi-scale features of the reconstructed image. The local features are extracted using a 3×3 convolutional kernel; the multi-scale features have a larger receptive field than the local features to obtain the correlation relationships between local blocks of the image. The multi-scale feature F MCB is extracted by a multi-scale convolutional block, and the multi-scale convolutional block MCB fuses the coarse-grained feature F coarse , the channel feature F channel and the dilated convolutional feature F dilated . The coarse-grained feature F coarse enhances the sensitivity to the coarse-grained features, establishes the relationships between these features, and promotes the integration of global context information. The channel feature F channel is used to compensate for the loss of channel-specific information that may occur during feature fusion. The dilated convolutional feature F dilated further expands the receptive field of the input feature map, thereby establishing the connections between pixel blocks. The following takes the first step as an example to introduce the feature extraction process of MCB. Subsequently, the features are concatenated along the channel direction, and a 4×4 convolution is used for feature extraction to obtain the fused feature F fusion . The fused features are further refined by 1×1 convolution to make them more suitable for the anomaly detection task. F out = LeakyReLU(Conv 1×1 (F fusion )) (16) The above is the specific implementation process of the multi-scale convolutional block MCB. At each step, the multi-scale features extracted by the MCB will be fused with the detailed features F extracted by the 3×3 convolution, and the fusion method is as follows. local The fusion is carried out as follows. F dual = F out + F local (17) After three steps of fusion, the reconstructed image will ultimately be compressed into reconstructed compressed information with the same dimension as the compressed information z The image information of the double receptive field is compressed in the reconstructed compressed information, and the loss is calculated with the compressed information z, thereby optimizing the training objective function.
4. An adaptive memory matching multi-class anomaly detection method according to claim 1, characterized in that The specific process of S3 is as follows: Designed with reference to the architecture of the generator network, an encoder-decoder-encoder structure discriminator network D is constructed. This discriminator network is divided into three sub-networks: encoder network Encoder1, decoder network, and encoder network Encoder2. The discriminator network inputs the input image and the reconstructed image respectively to feedback the authenticity of the image. The encoder network Encoder1 consists of a combination block of four layers of convolution, batch normalization, and the LeakyReLU activation function, which sequentially transforms the input with dimensions of H×W×C x into , where nz is the dimension of the compressed information z. The decoder network consists of a combination block of four layers of transposed convolution, batch normalization function, and the tanh activation function, which upsamples the 1×1×nz to the output with dimensions of H×W×C x ; The encoder network Encoder2 consists of a combination block of four layers of convolution, batch normalization, and the LeakyReLU activation function. Different from the encoder network Encoder1, the output of this network is a judgment on the authenticity of the content of the input discriminator rather than the compressed information z.
5. An adaptive memory matching multi-class anomaly detection method according to claim 1, characterized in that The specific process of S4 is as follows: In S1, S2, and S3, a generator network with a memory matching mechanism, a dual receptive field feature fusion network, and an encoder-decoder-encoder structure discriminator network are respectively constructed. On this basis, the architecture of the adaptive memory matching multi-class anomaly detection method is trained by setting three loss functions. The first loss function is the adversarial loss L used to train the generative adversarial network adv , through the setting of this loss function, the generator can capture the data distribution p of the training samples x x , so that the generator generates reconstructed samples G(x) with a similar data distribution. The adversarial loss L adv feeds back the training of the network by inputting the training samples x and the reconstructed samples G(x) into the discriminator and using the prediction results of the discriminator. The second loss function is the context loss function L between the reconstructed sample G(x) and the training samples con , which constrains the distance between samples through the L1 distance, thereby generating reconstructed samples G(x) with similar structures. The third loss function is the latent vector loss function. The generator network with the memory matching mechanism and the dual receptive field feature fusion network will compress the training sample x and the reconstructed sample G(x) into compressed information z and By constraining the distance in the feature space of the compressed information, the reconstruction effect of the reconstructed sample is further improved. According to the above three loss functions, the training objective L of a final adaptive memory matching multi-class anomaly detection method is determined, and each loss function is assigned by a different weight w: L = w adv L adv + w con L con + w dis-z L dis-z (21) 6. An adaptive memory matching multi-class anomaly detection method according to claim 1, characterized in that The specific process of S5 is as follows: The architecture of an adaptive memory matching multi-class anomaly detection method is deployed in the environment of Pytorch 1.0, optimized by the Adam optimizer, and experiments are conducted on multi-class anomaly detection datasets. In the training stage, the method is input with 32×32 images for training. The generator network with a memory matching mechanism will first compress the images into compressed information z with memory features, and then z will be restored to the original image size of 32×32, thus completing the image reconstruction. During training, the training samples and reconstructed samples are input into the discriminator network with an encoder-decoder-encoder structure, thereby promoting the training of the entire generative adversarial network. The specific implementation details in training are as follows: The weight parameter of the training objective is w con as the main training loss, which is set to 50, and the other weight parameters are all set to 1. The learning rate is set to 0.
003. The multi-class anomaly detection dataset contains 10 classes in total. Each class is used as the anomaly class for experiments in turn, and each class is trained for 15 epochs. The dimension of the compressed information is set to 512, and the training batch size is set to 64.