Multi-class unsupervised anomaly detection method and device
By combining the feature fusion method of Vision Transformer and CNN, the problem of insufficient utilization of global and local information in the prior art is solved, and a multi-class unsupervised anomaly detection with higher accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202510623243.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-07-29
AI Technical Summary
The existing multi-class unsupervised anomaly detection methods are not fully integrated when utilizing global and local information, which makes it difficult for the model to fully capture anomaly features of different scales in complex scenarios, and the calculation cost is high and the generalization ability is weak.
The pre-trained Vision Transformer is used to extract semantic features of different scales, and through adaptive multi-feature fusion within and between features, combining depth separation convolution and point-by-point convolution for intra-feature attention integration, and multi-level reconstruction is used to achieve weighted fusion of global and local features.
It improves the accuracy and robustness of the model in multiple unsupervised anomaly detection, can understand the global structure and local anomalies of the data more comprehensively, and reduces the calculation cost.
Smart Images

Figure CN120388238A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of anomaly detection, and more specifically, to a multi-class unsupervised anomaly detection method and device. Background Art
[0002] The anomaly detection (AD) task aims to identify abnormal data that is significantly different from most of the data. With the continuous development of deep learning, anomaly detection technology has been widely applied in various fields. For example, in industrial manufacturing, anomaly detection technology can replace manual labor to quickly identify defective products, thus avoiding missed and misdetected inspections caused by human factors and reducing operating costs.
[0003] Due to the variety and unpredictability of product defects in the industrial production process, constructing a high-quality defect dataset usually requires a large amount of time and labor costs. Therefore, unsupervised anomaly detection methods have become a research hotspot; however, the single-class setting not only increases the computing and storage costs but also reduces the efficiency of the model, making its deployment in practical applications complex and difficult; thus, multi-class unsupervised anomaly detection (MUAD) has become the main research direction.
[0004] Currently, existing MUAD methods are mainly divided into three categories:
[0005] Enhancement-based methods, which enhance model performance by artificially synthesizing anomalies. However, the complexity and unpredictability of anomalies in real-world scenarios result in weak generalization ability of the model;
[0006] Embedding-based methods, which distinguish anomalies by mapping features to a high-dimensional space. However, the computational cost is too high, and the efficiency is low when dealing with large-scale data and complex models, making it difficult to apply in resource-constrained environments;
[0007] In contrast, reconstruction-based methods utilize the characteristics of the data itself for reconstruction without artificially synthesizing anomalies. Therefore, they have strong adaptability to various types of anomalies and greatly improve the generalization ability; at the same time, they avoid the complex calculations of high-dimensional mapping, reduce the computational cost, and can also maintain high efficiency when dealing with large-scale data.
[0008] However, existing reconstruction methods usually rely solely on Transformer with global modeling ability or CNN with local perception ability, resulting in the existing anomaly detection models being unable to utilize global and local information simultaneously, having weak perception ability for the complex structure and subtle anomalies of the data. Even if global and local features are extracted, they are only fused by concatenation. This insufficient fusion makes the model unable to fully utilize the advantages of global and local information, resulting in the model being difficult to comprehensively capture anomaly features at different scales when facing complex scenarios. Summary of the Invention
[0009] In view of this, to at least partially solve the above technical problems, the present invention provides a multi-class unsupervised anomaly detection method and device;
[0010] To achieve the above object, the present invention adopts the following technical solutions:
[0011] On the one hand, the present application provides a multi-class unsupervised anomaly detection method, and the steps include:
[0012] Obtain the input image and extract semantic features of different scales;
[0013] Perform adaptive multi-feature fusion on the semantic features of different scales; including:
[0014] Based on the channel dimension, determine the intra-feature attention of each scale feature;
[0015] Integrate and separate the intra-feature attention in the feature map dimension to obtain the weights corresponding to different scale features;
[0016] Perform weighted fusion on each scale feature according to the weights;
[0017] Perform multi-level decoding and reconstruction on the fused features, calculate the residuals between the reconstruction results of each level and the semantic features of the corresponding scales, and add the obtained residual images to obtain the anomaly detection result.
[0018] Preferably, use the pre-trained Vision Transformer to extract semantic features of different scales.
[0019] The Vision Transformer has 12 Transformer layers, with every 3 Transformer layers as a stage, and outputs semantic features of one scale.
[0020] Preferably, based on the channel dimension, determine the intra-feature attention of each scale feature; including:
[0021] Adopt depthwise separable convolution to extract information in the spatial dimension of each scale feature;
[0022] Fuse the information between channels through pointwise convolution to obtain the intra-feature attention.
[0023] Preferably, integrate and separate the intra-feature attention in the feature map dimension, including:
[0024] Concatenate the intra-feature attention, and then perform non-linear transformation through the ReLU activation function;
[0025] Send the transformation result into the Softmax layer for adjustment and then separate it in the feature map dimension.
[0026] Preferably, each - level decoding and reconstruction sequentially includes feature extraction and adaptive multi - feature fusion.
[0027] Preferably, feature extraction includes:
[0028] Performing global feature extraction through multiple cascaded Transformer layers;
[0029] And performing local feature extraction through multiple parallel atrous convolution layers, where the first atrous convolution layer extracts the first local feature according to the fused feature, and the second and the remaining atrous convolution layers extract the current local feature according to the fused feature and the extraction result of the previous atrous convolution layer respectively.
[0030] On the other hand, the present application also provides a multi - class unsupervised anomaly detection device. This device applies the detection method described above. The device includes:
[0031] An encoder, configured to obtain an input image and extract semantic features of different scales;
[0032] A feature fusion module, configured to perform adaptive multi - feature fusion on semantic features of different scales; including:
[0033] Based on the channel dimension, determining the intra - feature attention of each scale feature;
[0034] Integrating and separating the intra - feature attention in the feature map dimension to obtain the weights corresponding to different scale features;
[0035] Performing weighted fusion on each scale feature according to the weights;
[0036] A decoder, configured to perform multi - level decoding and reconstruction on the fused feature, calculating the residual between each level of reconstruction result and the semantic feature of the corresponding scale, and adding the obtained residual images to obtain the anomaly detection result.
[0037] Preferably, the decoder includes multiple cascaded decoding and reconstruction modules, and the decoding and reconstruction module includes a feature extraction unit and a feature fusion module;
[0038] The feature extraction unit includes a global feature extraction branch and a local feature extraction branch:
[0039] The global feature extraction branch includes multiple cascaded Transformer layers;
[0040] The local feature extraction branch includes multiple parallel atrous convolution layers, where the first atrous convolution layer extracts the first local feature according to the fused feature, and the second and the remaining atrous convolution layers extract the current local feature according to the fused feature and the extraction result of the previous atrous convolution layer respectively.
[0041] By combining the global modeling ability of Transformer with the advantage of CNN in extracting local features, the present invention can more comprehensively understand the global structure and local anomalies of data, thereby improving the performance of the model in multi-class unsupervised anomaly detection;
[0042] At the same time, through normalization operations, the mutual correlation and weight adjustment within and between features are achieved, and then the global and local feature information from different scales is effectively integrated. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0044] Figure 1 It is the architecture diagram of the multi-class unsupervised anomaly detection network of the present invention;
[0045] Figure 2 It is the structure diagram of the feature fusion module of the present invention;
[0046] Figure 3 It is the structure diagram of the feature extraction unit in the decoder of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0048] The embodiments of the present invention disclose a multi-class unsupervised anomaly detection method and device, aiming to simultaneously and fully utilize global information and local information to improve the anomaly detection performance of the model in complex data sets; at the same time, solve the problems of insufficient utilization and insufficient fusion of global and local information in existing models, and achieve higher-precision anomaly detection and positioning.
[0049] This application can be applied to industrial production, such as monitoring the product quality on the production line, detecting surface defects, cracks or irregular shapes of products, etc.;
[0050] It can also be applied to the medical field, such as disease screening, especially for the analysis of medical images such as X-ray films, CT images, MRI scans, etc.;
[0051] In addition, in the field of video surveillance, the anomaly detection technology of the present application can be used to identify abnormal behaviors or events in a scene. By analyzing the video stream in real time, an alarm is immediately issued when an abnormal situation occurs, helping security personnel to respond in a timely manner.
[0052] In one embodiment, a multi-class unsupervised anomaly detection method includes the following steps:
[0053] Obtain an input image and extract semantic features at different scales;
[0054] Perform adaptive multi-feature fusion on the semantic features at different scales, including:
[0055] Based on the channel dimension, determine the intra-feature attention of each scale feature;
[0056] Integrate and separate the intra-feature attention in the feature map dimension to obtain the weights corresponding to the features at different scales;
[0057] Weightedly fuse the features at each scale according to the weights;
[0058] Perform multi-level decoding and reconstruction on the fused features. Calculate the residual between the reconstruction result of each level and the semantic feature of the corresponding scale, and add the obtained residual images to get the anomaly detection result.
[0059] In an exemplary embodiment, the overall detection network architecture is as Figure 1 shown, mainly consisting of three major parts: an encoder, a bottleneck, and a decoder.
[0060] (1) Encoder, used to obtain an input image and extract semantic features at different scales;
[0061] In this embodiment, the encoder adopts a pre-trained Vision Transformer (ViT), which is set with 12 Transformer layers, and each 3 Transformer layers output a semantic feature of one scale. In one embodiment, the input image is an RGB image with a length of 256 pixels, a width of 256 pixels, and 3 channels. Assuming the input image is X, when it is sent into the encoder, four features output by the 3rd, 6th, 9th, and 12th layers of the encoder can be obtained. The process is as follows:
[0062] F i = Encoder(X),
[0063] where i = 1, 2, 3, 4.
[0064] In the present application, the encoder can efficiently extract semantic features at different scales from the input data.
[0065] (2) Feature Fusion Module ((Adaptive Multi-Feature Fusion, AMFF)), also known as the bottleneck in this embodiment, is used to perform adaptive multi-feature fusion on the semantic features of different scales output by the four stages of the encoder; in this embodiment, in order to better fuse multiple features effectively, a simple but efficient feature fusion module is designed, and the structure is referred to Figure 2 ;
[0066] This module receives the feature F output by the encoder i , and performs intra-feature attention and inter-feature attention on it respectively.
[0067] For intra-feature attention, double convolution (depthwise separable convolution and pointwise convolution) operations are respectively applied to each feature to achieve intra-feature attention, and a weight representation y of the channel dimension of each feature is obtained i , that is, depthwise separable convolution is used to extract information in the spatial dimension of each scale feature; then pointwise convolution is used to fuse the information between channels to obtain intra-feature attention.
[0068] In one embodiment, the specific process is as follows:
[0069] Use depthwise separable convolution with a kernel size of 3 to extract information in the spatial dimension of the feature:
[0070] y depthwise = DepthwiseConv k=3 (F i ),
[0071] where i = 1, 2, 3, 4;
[0072] Then use pointwise convolution with a kernel size of 1 to fuse the information between channels:
[0073] y i = Conv 1×1 (y depthwise ),
[0074] where i = 1, 2, 3, 4;
[0075] Through the above operations, intra-feature attention is achieved, and it has lower computational overhead and higher flexibility.
[0076] For inter-feature attention, in this application, the intra-feature attention is concatenated, and then a non-linear transformation is performed through the ReLU activation function; then the transformation result is sent to the Softmax layer for adjustment and separated in the feature map dimension to obtain the weights corresponding to different scale features; furthermore, each scale feature is weighted and fused according to the weights;
[0077] In one embodiment, the specific process is as follows:
[0078] Concatenate y i on the second dimension (feature dimension). After concatenation, the features first undergo a non - linear transformation through the ReLU activation function, enabling the model to learn more complex patterns and relationships in the data. At the same time, it normalizes the feature values to a relatively small range. Then, it is sent to the Softmax layer to adjust the weights between features:
[0079]
[0080] where i = 1, 2, 3, 4. Then Separate the weights of each feature map on the second dimension again
[0081] Finally Multiply element - by - element with the corresponding feature map and then sum to obtain the finally fused feature. The process is as follows:
[0082]
[0083] where i = 1, 2, 3, 4.
[0084] This module realizes the mutual correlation and weight adjustment within and between features through normalization operations, thereby effectively integrating multi - source feature information and enhancing the model's ability to process and express complex information. This module can be used not only for multi - feature fusion in the decoder but also for fusing multi - features output by the encoder.
[0085] (3) The decoder includes multiple cascaded decoding and reconstruction modules. The decoding and reconstruction module includes a feature extraction unit and a feature fusion module; it is used for multi - level decoding and reconstruction of the fused features, and performs residual calculation between each level of reconstruction result and the semantic features of the corresponding scale. The obtained residual images are added together to obtain the anomaly detection result.
[0086] In one embodiment, the features obtained after passing through the bottleneck are sent to the decoder after adding Gaussian noise.
[0087] Adding Gaussian noise is to improve the robustness of the model, reduce overfitting to specific data, and at the same time help the model better identify abnormal data. In the anomaly detection task, after adding noise, the model can effectively distinguish normal and abnormal samples, improving the detection accuracy.
[0088] Furthermore, as Figure 3 , in order for the model to be able to simultaneously focus on global information and local information during the feature reconstruction process, Transformer and convolutional operations are applied to the decoder of the model. Specifically,
[0089] The feature extraction unit includes a global feature extraction branch and a local feature extraction branch:
[0090] The global feature extraction branch includes multiple cascaded Transformer layers; that is, the input features are passed through three layers of Transformer to extract the global feature z1. The calculation process of the Transformer is as follows:
[0091] z1 = MLP(LN(Attention(LN(F)) + F)) + Attention(LN(F)) + F
[0092] Where,
[0093] The local feature extraction branch includes multiple parallel dilated convolutional layers. Among them, the first dilated convolutional layer extracts the first local feature according to the fused features, and the second and subsequent dilated convolutional layers extract the current local features according to the fused features and the extraction results of the previous dilated convolutional layer respectively.
[0094] The calculation process is as follows:
[0095] z2 = Conv k=3,r=1 (F),
[0096] z3 = Conv k=3,r=2 (F + z2),
[0097] z4 = Conv k=3,r=4 (F + z3).
[0098] Where, k represents the convolutional kernel size, and r represents the dilation rate of the convolutional operation. Finally, the four obtained features are fed into the AMFF for feature fusion to obtain the final result.
[0099] In this embodiment, the decoder has a total of three stages, and each stage is composed of a feature extraction unit and a feature fusion module, and finally outputs the features of the three stages. Then, the output features of the 3rd, 6th, and 9th layers of the encoder are respectively subjected to residual calculation with the corresponding three features of the decoder, and finally the three obtained residual images are added together to obtain the final anomaly prediction map.
[0100] Compared with the prior art, the present invention effectively fuses global and local information, utilizes the advantages of Transformer and CNN, and solves the problem of insufficient information utilization and fusion in the existing methods, thereby significantly improving the accuracy and robustness of anomaly detection.
[0101] Furthermore, the detection effect of this application is demonstrated and verified.
[0102] Determine the training data set:
[0103] 1. MVTec-AD:
[0104] The MVTec AD dataset is the most widely used anomaly detection dataset, containing 15 industrial products of 2 types, 3,629 normal images for training, and 467 / 1,258 normal / anomaly images for testing (a total of 53,354 images).
[0105] 2. VisA:
[0106] The VisA dataset covers 12 objects of 3 types, 8,659 normal images for training, and 962 / 1,200 normal / anomaly images for testing (a total of 10,821 images).
[0107] 3. Real-IAD:
[0108] Real-IAD includes objects from 30 different categories, collecting 150K high-resolution images, making it larger than previous anomaly detection datasets. It consists of 99,721 normal images and 51,329 anomaly images.
[0109] The specific implementation environment of the present invention is based on the NVIDIA GeForce RTX 4090 GPU. For model training, the AdamW optimizer is adopted, with a batch size of 8 and a learning rate of 1e-4. Training is performed for 500 epochs using 1 GeForce RTX 4090 GPU. During training, the decoder learns and reconstructs the intermediate layer features of the encoder by maximizing the cosine similarity between the feature maps. The calculation formula for the cosine similarity loss is as follows:
[0110]
[0111] where F is the encoder feature, is the decoder feature; · represents the dot product, and ||F||2 represents the L2 norm of F. We calculate the cosine similarity loss for each stage and then sum them up to obtain the total loss:
[0112]
[0113] During testing, the anomaly map is also obtained by calculating the cosine similarity for each stage and then summing them up.
[0114] Then, a normal RGB image with a width of 256 pixels, a height of 256 pixels, and 3 channels is input into the detection network of the present application. The output image after passing through the network is a prediction map with a width of 256 pixels, a height of 256 pixels, and 1 channel, which indicates the anomaly probability at each pixel position.
[0115] Furthermore, image-level and pixel-level anomaly predictions are performed on the obtained prediction map. At the image level, the maximum value of all anomaly probabilities in the prediction map is selected to determine whether there is an anomaly in this image. If the maximum value is greater than 0.5, there is an anomaly; otherwise, there is none. Then, all the predicted labels are compared with the true labels, and the AUROC, AP, and F1-score at the image level are calculated.
[0116] At the pixel level, a binarization operation is performed on the prediction probability of each pixel, that is, if the probability value is greater than 0.5, the value at that position is set to 1; otherwise, it is set to 0. The specific calculation process is as follows:
[0117]
[0118] Among them, f(x, y) represents the prediction map output by the network, (x, y) represents the coordinates of the image, and g(x, y) represents the binarized prediction map. Through the above steps, a binarized anomaly localization prediction map with a width of 256 pixels, a height of 256 pixels, and a channel number of 1 will be obtained. Then, the predicted map of anomaly localization is compared with the true map pixel by pixel, and the AUROC, AP, and F1-score at the pixel level are calculated.
[0119] This application compares with the prior art in a total of 6 metrics at the image level and pixel level respectively, and the comparison results are as follows:
[0120]
[0121]
[0122] The results prove that the detection method of this application has achieved the best in each metric.
[0123] On the MVTec-AD dataset, the detection method of this application is superior to all comparison methods and has achieved the best results of 99.3 / 99.7 / 98.3 and 98.3 / 57.9 / 61.9 in multi-class anomaly detection and segmentation. Specifically, compared with the previous best result MambaAD, the detection method of this application has improved by 0.7 / 0.1 / 0.5 at the image level and 0.6 / 1.6 / 2.7 at the pixel level.
[0124] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0125] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-class unsupervised anomaly detection method, characterized in that Obtain the input image and extract semantic features at different scales; Perform adaptive multi-feature fusion on the semantic features at different scales; including: Based on the channel dimension, determine the intra-feature attention of each scale feature; Integrate and separate the intra-feature attention in the feature map dimension to obtain the weights corresponding to the features at different scales; Perform weighted fusion on the features at each scale according to the weights; Perform multi-level decoding and reconstruction on the fused features. Calculate the residuals between the reconstruction results at each level and the semantic features at the corresponding scale, and add the obtained residual images to obtain the anomaly detection result.
2. The multi-class unsupervised anomaly detection method according to claim 1, wherein Use the pre-trained Vision Transformer to extract semantic features at different scales.
3. The multi-class unsupervised anomaly detection method according to claim 1, wherein Based on the channel dimension, determine the intra-feature attention of each scale feature; including: Use depthwise separable convolution to extract information in the spatial dimension of each scale feature; Fuse the information between channels through pointwise convolution to obtain the intra-feature attention.
4. The multi-class unsupervised anomaly detection method according to claim 1, characterized in that, Integrate and separate the intra-feature attention in the feature map dimension, including: Concatenate the intra-feature attention, and then perform non-linear transformation through the ReLU activation function; Send the transformation result into the Softmax layer for adjustment and then separate it in the feature map dimension.
5. The multi-class unsupervised anomaly detection method according to claim 1, characterized in that Each level of decoding and reconstruction sequentially includes feature extraction and adaptive multi-feature fusion.
6. The multi-class unsupervised anomaly detection method according to claim 5, wherein Feature extraction includes: Perform global feature extraction through multiple cascaded Transformer layers; And perform local feature extraction through multiple parallel atrous convolution layers. Among them, the first atrous convolution layer extracts the first local feature according to the fused feature, and the second and subsequent atrous convolution layers extract the current local feature according to the fused feature and the extraction result of the previous atrous convolution layer respectively.
7. A detection device applying the multi-class unsupervised anomaly detection method according to any one of claims 1-6, characterized in that, Including: An encoder for obtaining the input image and extracting semantic features at different scales; A feature fusion module for performing adaptive multi-feature fusion on the semantic features at different scales; including: Based on the channel dimension, determine the intra-feature attention of each scale feature; Integrate and separate the intra-feature attention in the feature map dimension to obtain the weights corresponding to the features at different scales; Perform weighted fusion on the features at each scale according to the weights; A decoder for performing multi-level decoding and reconstruction on the fused features. Calculate the residuals between the reconstruction results at each level and the semantic features at the corresponding scale, and add the obtained residual images to obtain the anomaly detection result.
8. The multi-class unsupervised anomaly detection device according to claim 7, characterized in that, The decoder includes multiple cascaded decoding and reconstruction modules, and the decoding and reconstruction module includes a feature extraction unit and a feature fusion module; The feature extraction unit includes a global feature extraction branch and a local feature extraction branch: The global feature extraction branch includes multiple cascaded Transformer layers; The local feature extraction branch includes multiple parallel atrous convolution layers. Among them, the first atrous convolution layer extracts the first local feature according to the fused feature, and the second and subsequent atrous convolution layers extract the current local feature according to the fused feature and the extraction result of the previous atrous convolution layer respectively.