Small sample image anomaly detection method based on multi-scale attention and efficient convolution
By constructing a small sample image anomaly detection method of ResNet network and multi-scale attention mechanism, the problem of failing to fully utilize early features in the prior art is solved, and efficient anomaly detection under small sample conditions is achieved.
Patent Information
- Application Number
- CN202510357469.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing small sample image abnormality detection methods fail to make full use of features extracted in the early stage of encoding, resulting in insufficient accuracy of abnormality detection.
A small sample image anomaly detection method based on multi-scale attention and efficient convolution is adopted. By constructing a ResNet network encoder, combining MSCAM, EUCB and LGAG modules, semantic information at different levels of the image is integrated, and feature mapping is enhanced by using a multi-scale convolution attention mechanism, and anomaly detection is performed by combining a probability judgment classifier of multivariate Gaussian distribution.
When the number of training samples is limited and the test sample label is unknown, the model's understanding of normal sample distribution is improved, the sensitivity to abnormal points and detection accuracy is enhanced, and the semantic information at different levels of the image is effectively integrated, which improves the accuracy of abnormal detection.
Smart Images

Figure CN120298770A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a small-sample image anomaly detection method based on multi-scale attention and efficient convolution. Background Art
[0002] Anomaly Detection (AD) is to predict the results that do not conform to the expected behavior. In an industrial environment, the "anomalies" on the product surface are vaguely defined and there are various forms of abnormal defects on the production line.
[0003] Denis Gudovskiy et al. solved the problem of infeasible labels based on conditional normalizing flow and combined with a multi-scale generative decoder; Yao et al. proposed an explicit boundary-guided semi-push-pull contrast learning mechanism, and the boundary only depends on the normal feature distribution, thus alleviating the bias problem caused by a small number of known anomalies; however, these traditional unsupervised methods need to rely on a large amount of unlabeled data, often assuming that there are significant statistical differences between normal data and abnormal data. In the absence of prior knowledge, it is easy to misclassify data noise or irrelevant data instances as anomalies, resulting in inaccurate detection results.
[0004] For few-shot anomaly detection (FSAD) such as Xie et al., only a small number of normal images are used as the training data set to detect anomalies in the test samples. Visual isometric invariant features (VIIF) are introduced as the features for anomaly detection, and the isomorphism of the graph neural network (GNN) is used to replace data augmentation; Huang et al. proposed a few-shot anomaly detection (FSAD) method that does not require retraining or parameter fine-tuning. By using image registration as a proxy task, an anomaly detection model that is universal across categories is trained; during testing, anomalies are identified by comparing the registration features of the test image with the corresponding support (normal) image; however, this type of method fails to fully utilize the features extracted in the early stage of encoding, and the accuracy of anomaly detection needs to be further improved. Summary of the Invention
[0005] Aiming at the deficiencies of the existing methods, the present invention solves the problem that the existing methods fail to fully utilize the features extracted in the early stage of encoding, and the accuracy of anomaly detection needs to be further improved.
[0006] The technical solution adopted by the present invention is that the small-sample image anomaly detection method based on multi-scale attention and efficient convolution includes the following steps:
[0007] Step 1: Collect small-sample image data;
[0008] Step 2: Construct a ResNet network encoder and decoders for the MSCAM module, EUCB module, and LGAG module; extract image features through the encoder, construct a layer output feature map from low-level to high-level, and learn different-level feature information to integrate different-level semantic information of the image; in the decoder, use a multi-scale convolutional attention mechanism to enhance the feature map; use channel, spatial, and grouped gating attention mechanisms to integrate complex spatial relationships and local attention; and splice the output vectors of different levels of the network.
[0009] As a preferred embodiment of the present invention, the first residual block of ResNet is connected to the first MSCAM module; the second residual block of ResNet is connected to the first LGAG module. After the first MSCAM module is connected to the first EUCB module, the first branch is connected to the first LGAG module. The first LGAG module performs an addition operation with the second branch of the first EUCB module and inputs it into the second MSCAM module; the third residual block of ResNet is connected to the third LGAG module. After the second MSCAM module is connected to the second EUCB module, the first branch is connected to the second LGAG module. The second LGAG module performs an addition operation with the second branch of the second EUCB module and inputs it into the third MSCAM module.
[0010] As a preferred embodiment of the present invention, the MSCAM module includes: a CHA module, a SPA module, and an MSCB module. The first branch of the CHA module includes Global average pooling, 1x1Conv, ReLU, and 1x1Conv; the second branch includes Global max pooling, 1x1Conv, ReLU, and 1x1Conv, and the third branch is the input feature vector; after the addition operation of the first branch and the second branch, it passes through the Sigmoid function and then multiplies with the third branch to output the feature vector as the input of the SPA module.
[0011] As a preferred embodiment of the present invention, the SPA module performs Channel Max and Channel Ave on the output features of the CHA module, then performs a Concat operation, followed by 7x7Conv and the Sigmoid function, and finally multiplies with the output features of the CHA module.
[0012] As a preferred embodiment of the present invention, the MSCB module performs 1x1Conv, BN, MSDC module, and BN operations on the output features of the SPA module and then adds them to the output features of the SPA module.
[0013] As a preferred embodiment of the present invention, the MSDC module separately performs three DWC convolutions, BN, and ReLU6 on the input features and then adds them together, and then performs a Channel shuffle operation.
[0014] As a preferred embodiment of the present invention, the EUCB module includes: performing upsampling on the input features and then successively performing DWC convolution, BN, Relu, and 1x1Conv.
[0015] As a preferred embodiment of the present invention, the LGAG module includes: separately performing two 3x3Group Conv and BN operations on the input features, and successively performing ReLU, 1x1Conv, and Sigmoid functions on the two output features and then multiplying them with the input features.
[0016] As a preferred embodiment of the present invention, a probability decision classifier using a multivariate Gaussian distribution is employed to calculate the covariance matrix and Mahalanobis distance of the normal samples in the decoder output.
[0017] As a preferred embodiment of the present invention, a few-shot image anomaly detection system based on multi-scale attention and efficient convolution includes: a memory for storing instructions executable by a processor; and a processor for executing the instructions to implement a few-shot image anomaly detection method based on multi-scale attention and efficient convolution.
[0018] Advantages of the present invention:
[0019] 1. Under the condition that the number of training samples is limited and the labels of test samples are unknown, the present invention enhances the model's understanding ability of the normal sample distribution, thereby improving the sensitivity and detection accuracy of the model to the outliers deviating from the normal distribution;
[0020] 2. The present invention processes images by an encoder to extract features, constructs output feature maps of multiple layers from low level to high level, and learns different levels of feature information to integrate semantic information of different levels of the image; by using the multi-level features generated by the encoder, it avoids relying on additional extraction components to obtain multi-scale features for anomaly detection;
[0021] 3. In the decoder, a multi-scale convolutional attention mechanism is constructed to enhance the feature maps. Through channel, spatial, and grouped gating attention mechanisms, complex spatial relationships and local attention are effectively integrated; then, by concatenating the output vectors of different levels of the network and calculating the covariance matrix, and comparing it with the covariance matrix of the test samples, the anomaly detection result is output. Description of the Drawings
[0022] Figure 1 is an application framework diagram of the few-shot image anomaly detection based on multi-scale attention and efficient convolution of the present invention;
[0023] Figure 2 is the structural diagram of the multi-scale convolutional attention module (MSCAM) of the present invention;
[0024] Figure 3 is the structural diagram of the multi-scale depthwise separable convolution (MSDC) of the present invention;
[0025] Figure 4 is the structural diagram of the efficient upsampling convolutional block (EUCB) of the present invention;
[0026] Figure 5 is the structural diagram of the large kernel grouped attention gating (LGAG) of the present invention;
[0027] Figure 6 is the schematic diagram of the multivariate Gaussian distribution feature alignment of the present invention. Detailed implementation manners
[0028] The present invention will be further described below with reference to the accompanying drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner. Therefore, it only shows the components related to the present invention.
[0029] As Figure 1 shown, the small-sample image anomaly detection method based on multi-scale attention and efficient convolution includes the following steps:
[0030] Step 1: Collect small-sample image data;
[0031] Download the MVTec dataset from the official image data, and further create a training set, a validation set, and a test set, which contain information such as the label and path of each image, and read the images according to the batch size;
[0032] In the training stage, assume that the training set contains normal samples from several categories where the subset T i includes k samples belonging to the category c i ∈ C, and the category set c i = {c i |i = 1, 2, 3…n};
[0033] In the test stage, given a support set s t consisting of k normal samples from n categories belonging to the target category c t ; where c t ∈ C, perform anomaly evaluation on k test samples (which may be normal or abnormal) from the target category c t .
[0034] Step 2: Construct a probability decision classifier based on a feature extractor, a multi-scale convolutional attention module (MSCAM), an efficient upsampling convolutional module (EUCB), a large kernel grouped attention gating (LGAG), and a multivariate Gaussian distribution;
[0035] Train the pre-trained network structure. In the pre-training stage, perform forward propagation, calculate the loss, backpropagation, and optimization, record and update the best accuracy and the corresponding model parameters; in the testing stage, load the saved best model parameters, construct test data for the final test, and complete the small sample image anomaly detection task;
[0036] The feature extractor is a ResNet-18 residual network. In the training stage: use the ResNet-18 residual network pre-trained on the large dataset ImageNet for feature extraction, and output feature vectors of different network layers where x i represents the i-th sample in the training set, x i ∈ T train , l represents the output layer of the ResNet different convolutional residual block neural network. Use the first three convolutional residual modules C1, C2, and C3 of ResNet-18. Therefore, l ∈ {1, 2, 3}.
[0037] Use the feature vector output by the first residual block C1 of the ResNet-18 network as the input of the first MSCAM module;
[0038] As Figure 2 shown, the MSCAM module includes: a channel attention module (CHA), a spatial attention module (SPA), and a multi-scale convolutional module (MSCB);
[0039] The CHA module consists of three branches. In the first branch, the feature vector successively passes through Global averagepooling, 1x1Conv, ReLU, and 1x1Conv; in the second branch, the feature vector successively passes through Global maxpooling, 1x1Conv, ReLU, and 1x1Conv. The third branch is the feature vector After the addition operation of the first branch and the second branch, it passes through the Sigmoid function and then multiplies with the third branch, and outputs the feature vector as the input of the SPA module;
[0040] The SPA module includes three branches. The output feature vector of the SPA module in the first branch passes through Channel Max, and the output feature vector of the SPA module in the second branch passes through Channel Ave. After the Concat operation on the first and second branches, the result is input into a 7x7 Conv. After padding = 3 and passing through the Sigmoid function, it is added to the output feature vector of the SPA module in the third branch and serves as the input to the MSCB module;
[0041] The MSCB module includes two branches. The output feature vector of the SPA module in the first branch sequentially passes through 1x1 Conv, BN, the MSDC module, and BN, and then is added to the output feature of the SPA module in the second branch to output the feature vector
[0042]
[0043] where, represents the output feature vector of the MSCAM module, and CHA(·) is the channel attention module.
[0044] The MSDC module includes three branches. The three branches include DWC convolution, BN, and Relu6. After the addition operation of the three branches, a ChannelShuffle operation is performed.
[0045] Specifically, different features and information are captured through parallel adaptive average pooling (p aa (·)) and adaptive max pooling (p am (·)) paths. Subsequently, a 1x1 convolutional kernel Conv 1×1 (·) performs dot product calculations independently for each channel to reduce the number of channels, thereby reducing the feature dimension and computational cost while retaining key features; then, after passing through the ReLU activation function R e (·), a point convolution Conv(·) is used to restore the original number of channels, defined as Equation (2):
[0046]
[0047] where, σ represents the Sigmoid function.
[0048] Finally, after adding the feature maps of the parallel channels, the Sigmoid activation function is used to normalize the attention weights, thereby realizing the information integration between channels.
[0049] SPA(·) is the spatial attention mechanism. First, max-pooling and average-pooling operations are respectively performed on the channel dimension to generate two spatial feature maps. After these two feature maps are concatenated on the channel dimension, feature extraction is carried out through a convolutional layer with a kernel size of 7×7. Subsequently, the sigmoid activation function is used to map the convolutional output to the range of [0,1], generating the spatial attention weight map, which is defined as Equation (3):
[0050]
[0051] where MaxPool(·) and AvgPool(·) respectively represent max-pooling and average-pooling;
[0052] The MSCB module expands the channels of the above output through 1x1 convolution, and the definition formula is (4):
[0053] x imscb = BN(MSDC(BN(Conv(x ispa )))) + x ispa (4)
[0054] where BN(·) represents batch normalization;
[0055] The definition formula of MSDC is (5):
[0056] x imsdc = ChannelShuffle(Concat(R e (BN(DWC p×p (x))))) (5)
[0057] where R e is the activation layer processing, ChannelShuffle is channel shuffling, and DWC is depthwise separable convolution.
[0058] As Figure 3 shown in the MSDC module; first, the input passes through depthwise separable convolutions (DWC) with different kernel sizes, and then sequentially passes through the batch normalization layer and the activation layer for processing. The outputs of different scales are combined in a concatenated manner; then, the number of channels is compressed through 1x1 convolution, and channel shuffling and normalization processing are performed.
[0059] As Figure 4 shown, it is the efficient upsampling convolutional block (EUCB). The current feature map is upsampled through the EUCB, and the scaling factor is set to 2. After the EUCB module upsamples the input features, a 3×3 depth convolution DWC(·) is used. After passing through the batch normalization BN(·) and the ReLU(·) activation function, a 1×1 convolution is used to reduce the number of channels to match the feature map of the next stage, and the definition formula is (6):
[0060] x ieucb = Conv(R e (BN(DWC 3×3 (UP(x))))) (6)
[0061] Such as Figure 5 , large kernel group attention gating (LGAG), the LGAG module processes the input feature map separately on different channels through the combination of large kernel convolution and attention mechanism, realizes the capture and fusion of multi-scale information, and improves the performance of image segmentation and visual tasks;
[0062] The input feature map first undergoes feature extraction through two independent 3×3 grouped convolutions. Each grouped convolution processes different channels of the input feature map respectively, thereby capturing multi-scale local features; the feature map after the convolution operation is normalized through a batch normalization layer to stabilize the training process and improve the generalization ability of the model; the two groups of normalized feature maps are fused through element-wise addition to integrate the multi-channel features into a unified feature representation, and the defining formula is (7):
[0063]
[0064] Among them, GC p (·) and GC q (·) represent the grouped convolution operation.
[0065] The fused feature map undergoes a non-linear transformation through the ReLU activation function R e (·) to further enhance the expression ability of effective features; then, it passes through a 1×1 convolution layer to compress the features in the channel dimension, generating a single-channel feature map; the single-channel feature map generates attention coefficients through the Sigmoid activation function to measure the importance of features at each position; finally, these attention coefficients act on the initial input feature map through element-wise multiplication to generate the final attention gating feature map; this mechanism effectively enhances the feature expression related to anomalies and suppresses irrelevant or redundant information by dynamically adjusting the weights at different positions in the feature map, providing accurate feature output for image anomaly detection, and the defining formula is (8):
[0066]
[0067] The output after feature concatenation of the first three convolutional residual modules C1, C2, and C3 of the residual network after the same operation is set as where X k is a matrix of d × k, and each column corresponds to a feature vector x i .
[0068] In the test phase:
[0069] As Figure 6 The construction process of the probability decision classifier for the multivariate Gaussian distribution. To learn and test the normal features of the image, first, the image is divided into a grid, where each position (m, n) is defined as a grid cell with a feature resolution of [1, W] × [1, H]; at position (m, n), a set of embedding vectors X is extracted from N normally trained support images mn , Assume that the embedding vectors follow a multivariate Gaussian distribution N(μ mn , ∑ mn ), where μ mn is the mean vector of the normal samples at position (m, n), and ∑ mn is the covariance matrix of the normal samples at position (m, n), and the defining formula is as (9):
[0070]
[0071] Among them, the regularization term ∈I ensures that the covariance matrix is full-rank and invertible.
[0072] The set of features extracted through the multi-scale attention mechanism and efficient convolution not only combines the feature information of the image at different levels but also can effectively capture the local and global detailed features of the image.
[0073] In the test stage, the Mahalanobis distance is used to evaluate the deviation degree of the features of each position (m, n) in the sample to be tested from the normal sample distribution, so as to calculate the anomaly score at this position; when the anomaly scores of the test sample at certain positions are significantly higher than the normal sample distribution, these positions will be regarded as anomaly regions; the definition of the Mahalanobis distance is (10):
[0074]
[0075] Among them, is the inverse of the covariance matrix.
[0076] Experimental process:
[0077] To verify the effectiveness of the model of the present invention, experiments were carried out on the MVTec and MPDD datasets. The following is a brief introduction to each dataset.
[0078] MVTec: The MVTec AD dataset is a standard dataset specifically for evaluating anomaly localization algorithms in industrial quality control and is widely used in the fields of machine learning and computer vision. This dataset contains industrial images of 15 categories, covering 10 object categories and 5 texture categories, with 73 types of defects (scratches, dents, dirt, deformations, material shortages, etc., all artificially created), a total of 3466 unannotated images and 1888 annotated images (pixel-level segmentation annotations), all of size 700×700 or 1024×1024. The training set in the MVTec AD dataset consists of normal images without defects, a total of 3629 images. The test set includes 1725 images with different types of defects and normal images, and each category has an average of 5 different defect types on average. The test set also provides per-pixel ground truth labels for anomaly region localization.
[0079] The MPDD dataset is a new dataset specifically proposed for defect detection in the manufacturing process of painted metal parts, containing 6 categories of metal parts. The image shooting conditions in this dataset are complex, including multiple objects in different spatial directions, positions and distances. In addition, there are different lighting intensities and non-uniform backgrounds, making the anomaly detection task very challenging. The training set of the MPDD dataset contains 888 normal images, and the test set includes 176 normal images and 282 images with defects, and the image resolution is 1024×1024 for all. In addition, the dataset also provides per-pixel ground truth labels for each defect region for anomaly localization.
[0080] Table 1 conducts k-shot anomaly detection experiments on the MPDD dataset and compares with the latest methods;
[0081]
[0082] Table 1 reports the average AUROC (%) after 10 experimental runs and bolds the best-performing results.
[0083] Table 2 conducts k-shot anomaly detection experiments on the MPDD dataset and compares with the latest methods
[0084]
[0085] Table 2 reports the average AUROC (%) after 10 experimental runs and bolds the best-performing results.
[0086] Based on the above-mentioned ideal embodiments of the present invention as inspiration, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. A few-shot image anomaly detection method based on multi-scale attention and efficient convolution, characterized in that It includes the following steps: Step 1, collect small sample image data; Step 2, construct a ResNet network encoder and a decoder based on the MSCAM module, EUCB module, and LGAG module; extract image features through the encoder, construct a layer output feature map from low level to high level, and learn different-level feature information to integrate different-level semantic information of the image; in the decoder, use a multi-scale convolutional attention mechanism to enhance the feature map, use channel, spatial, and grouped gating attention mechanisms to integrate complex spatial relationships and local attention, and splice the output vectors of different levels of the network.
2. The few-shot image anomaly detection method based on multi-scale attention and efficient convolution according to claim 1, characterized in that The first residual block of ResNet is connected to the first MSCAM module; the second residual block of ResNet is connected to the first LGAG module. After the first MSCAM module is connected to the first EUCB module, the first branch is connected to the first LGAG module. The first LGAG module performs an addition operation with the second branch of the first EUCB module and inputs it into the second MSCAM module; the third residual block of ResNet is connected to the third LGAG module. After the second MSCAM module is connected to the second EUCB module, the first branch is connected to the second LGAG module. The second LGAG module performs an addition operation with the second branch of the second EUCB module and inputs it into the third MSCAM module.
3. The few-shot image anomaly detection method based on multi-scale attention and efficient convolution according to claim 1, characterized in that The MSCAM module includes: a CHA module, a SPA module, and an MSCB module. The first branch of the CHA module includes Global average pooling, 1x1Conv, ReLU, and 1x1Conv; the second branch includes Global max pooling, 1x1Conv, ReLU, and 1x1Conv, and the third branch is the input feature vector; after the addition operation of the first branch and the second branch, it passes through the Sigmoid function and then multiplies with the third branch to output the feature vector as the input of the SPA module.
4. The few-shot image anomaly detection method based on multi-scale attention and efficient convolution according to claim 3, wherein The SPA module performs Channel Max and Channel Ave on the output features of the CHA module, then performs a Concat operation, followed by 7x7Conv and the Sigmoid function, and then multiplies with the output features of the CHA module.
5. The few-shot image anomaly detection method based on multi-scale attention and efficient convolution according to claim 4, wherein The MSCB module performs 1x1Conv, BN, MSDC module, and BN operations on the output features of the SPA module and then adds them to the output features of the SPA module.
6. The few-shot image anomaly detection method based on multi-scale attention and efficient convolution according to claim 5, wherein The MSDC module performs three DWC convolutions, BN, and ReLU6 on the input feature respectively and then adds them, and then performs a Channel shuffle operation.
7. The few-shot image anomaly detection method based on multi-scale attention and efficient convolution according to claim 1, wherein The EUCB module includes: performing upsampling on the input feature and then sequentially performing DWC convolution, BN, ReLU, and 1x1Conv.
8. The few-shot image anomaly detection method based on multi-scale attention and efficient convolution according to claim 1, characterized in that, The LGAG module includes: performing two 3x3Group Conv and BN operations on the input feature respectively, and sequentially performing ReLU, 1x1Conv, and Sigmoid function on the two output features and then multiplying with the input feature.
9. The few-shot image anomaly detection method based on multi-scale attention and efficient convolution according to claim 1, wherein Use a probability decision classifier with a multivariate Gaussian distribution to calculate the covariance matrix and Mahalanobis distance of the normal samples output by the decoder.
10. A small-sample image anomaly detection system based on multi-scale attention and efficient convolution, characterized in that, It includes: A memory for storing instructions executable by a processor; A processor for executing the instructions to implement the few-shot image anomaly detection method based on multi-scale attention and efficient convolution according to any one of claims 1-9.
Citation Information
Cited By
Intelligent positioning and grabbing method for crosshead part assembly
CN120816498A
Intelligent positioning and grabbing method for crosshead part assembly
CN120816498B