Unsupervised anomaly detection method based on multi-scale features
By introducing the multi-scale feature adaptive unsupervised detection framework MSFA, a global dynamic transformation unit GDT and a local area interaction module LRI, the problems of high storage and computing costs and learning shortcuts in multi-category anomaly detection are solved, and the abnormal detection performance and recognition capabilities are improved.
Patent Information
- Application Number
- CN202510583403.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-09-02
AI Technical Summary
The existing unsupervised anomaly detection methods have problems such as high storage and computing costs, difficulty in adapting to the needs of multiple industrial scenarios, and prone to learning shortcuts in multi-category anomaly detection, resulting in degradation of detection performance.
MSFA is adopted to introduce the global dynamic transformation unit GDT and the local area interaction module LRI to enhance feature expression and abnormal recognition capabilities, and to optimize feature aggregation using the global dynamic transformation unit, the local area interaction module promotes information interaction and generates joint feature representations.
It significantly improves the abnormal detection performance, improves the model's ability to recognize multiple categories of abnormalities, reduces the dependence on local redundant information, enhances the recognition ability of critical abnormal areas, and provides intuitive visual support.
Smart Images

Figure CN120580407A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to an unsupervised anomaly detection method based on multi-scale features. Background Art
[0002] Anomaly detection (AD), a key technology in data analysis and computer vision, is widely used in scenarios such as industrial defect detection, medical image analysis, and video surveillance. Its core goal is to identify data points or regions that deviate from normal patterns, thereby enabling early warning and diagnosis of potential problems. However, traditional anomaly detection methods rely on supervised learning and require a large number of labeled anomaly samples. In practical applications, the scarcity and high diversity of anomaly samples, coupled with the high cost of labeling, significantly limits the practical application of such methods.
[0003] To overcome these limitations, unsupervised anomaly detection (UAD) has become a research hotspot. UAD methods use only normal samples during training, learning their typical features and expression patterns to identify image regions that differ significantly from normal patterns during testing. However, as application requirements continue to expand and become more complex, existing UAD methods face unique challenges in multi-category anomaly detection. On the one hand, traditional methods mostly adopt a "single model for a single category" training strategy, which increases storage and computational costs and is difficult to adapt to the needs of multi-category industrial scenarios. On the other hand, even using a unified model to simultaneously process multiple image types is prone to the "shortcut learning" problem. Specifically, during training, the model may tend to choose a simpler mapping method to avoid in-depth learning of the features of normal images across categories. For example, when presented with an abnormal input image, rather than repairing it based on its understanding of normal morphology, the model may uniformly blur all inputs to make them appear "normal." Although this fuzzy, featureless reconstruction strategy can reduce reconstruction errors, it masks abnormal signals in the original image, making it difficult for the model to correctly judge the location and nature of abnormalities during the inference stage, thereby significantly weakening the detection performance. Summary of the Invention
[0004] Purpose of the invention: In response to the problems pointed out in the background technology, the present invention discloses an unsupervised anomaly detection method based on multi-scale features, and proposes a multi-scale feature adaptive unsupervised anomaly detection framework MSFA. By introducing the global dynamic transformation unit GDT and the local region interaction module LRI, the feature expression and anomaly recognition capabilities of the model are enhanced. MSFA performs unified dimensional mapping on the multi-layer intermediate feature maps, and retains the spatial structure information by superimposing position encoding. It adopts a weighted fusion strategy to integrate multi-layer information to generate a joint feature representation for reconstruction.
[0005] Technical solution: The present invention discloses an unsupervised anomaly detection method based on multi-scale features, comprising the following steps:
[0006] Obtain an image dataset to be detected and preprocess the image dataset;
[0007] Construct a multi-scale feature adaptive unsupervised detection framework MSFA, hereinafter referred to as the MSFA framework, which includes a feature extraction module, a feature reconstruction module and anomaly detection module;
[0008] The feature extraction module processes the input image through the pre-trained network and extracts multi-scale intermediate feature maps. All feature maps are flattened and uniformly mapped to a fixed dimension, and position codes are superimposed.
[0009] The feature reconstruction module processes the features extracted by the feature extraction module by introducing a global dynamic transformation unit GDT and a local region interaction module LRI. The two are executed collaboratively in the encoder, and the fusion result is combined with the guided reference representation to generate the final reconstructed feature map;
[0010] The anomaly detection module compares the reconstructed feature map with the original feature map layer by layer, calculates the reconstruction error and cosine similarity, generates a multi-scale anomaly score, and all score maps are upsampled to the size of the original image, and finally outputs an anomaly localization map.
[0011] Furthermore, the pre-trained network in the feature extraction module adopts a pre-trained convolutional neural network (CNN) as a feature extractor.
[0012] Furthermore, all feature maps are flattened and uniformly mapped to a fixed dimension, and position encoding is superimposed, as follows:
[0013] Assume that the extracted feature map is: Where j represents the level index of the feature map, C j 、H j 、W jare the number of channels, height, and width of the j-th layer feature map, respectively. First, the feature map is flattened in the spatial dimension so that the feature vector corresponding to each spatial position is represented as an independent unit, and we get: Among them L j =H j ×W j ;
[0014] Then, the dimension of each feature vector is transformed from C to j Mapped to a unified dimension D to ensure the uniformity of features:
[0015] Introducing a set of trainable position features for each position And add it element by element to the corresponding eigenvector to get:
[0016] Finally, the feature map of each layer is converted into a representation Z with uniform dimension (j) .
[0017] Furthermore, the reconstruction encoder module introduces a global dynamic transformation unit GDT and a local region interaction module LRI in the encoder, and inputs the extracted features processed by the feature extraction module and the guided reference representation R from the auxiliary knowledge base into the modified encoder. The modified encoder receives the input features and the reference representation R, and the input features and the reference representation R respectively participate in attention interaction. The fusion result is normalized and processed by the feedforward network FFN and then output; the modified encoder contains four layers and the structure remains consistent.
[0018] Furthermore, the reference representation R is a learnable parameter matrix, which provides the model with common information across images through the cross-image regularities learned during the training process, and is used to enhance feature selection in the attention mechanism.
[0019] Furthermore, the global dynamic transformation unit GDT introduces a competitive inhibition mechanism and introduces a learnable relationship strength B when measuring the relationship between positions. pos1,pos2 , works together with the key vector to adjust the contribution of each position to the current query position. The formula is as follows:
[0020] F pos1,pos2 =f(K pos1 ,B pos1,pos2 )
[0021] Among them, B pos1,pos2 It describes the relative influence between position pos1 and position pos2, and f(·) is the fusion weight function used to integrate the information of each position;
[0022] GDT first calculates the F corresponding to each position pos1pos1,pos2 Normalize and get the fusion coefficient α pos1,pos2 On this basis, we further introduce the query vector Q i The relevant adjustment factors are used to weight the fusion results, and the final output is as follows:
[0023]
[0024] Among them, σ(·) is the Sigmoid function, which is used to adjust the intensity of the output at each position.
[0025] Furthermore, the specific operations of the local region interaction module LRI are:
[0026] The input features are first divided into multiple S×S non-overlapping sub-regions X (r) , where r represents the region number. Within each region, a local self-attention mechanism is performed using three trainable linear transformation matrices The input is projected into the Query, Key, and Value spaces respectively, and the attention weight is calculated by scaling the dot product to achieve information aggregation within the region:
[0027] Q,K,V=X (r) W Q ,X (r) W K ,X (r) W V
[0028] In each sub-region X (r) In the code, the local feature a obtained by applying the self-attention mechanism is updated (r) , the global vector representation of each region is obtained by average pooling Concatenate the vector representations of all regions according to the region number to obtain the region representation matrix
[0029] Design a cross-region information scheduling mechanism, transform the regional feature matrix through a learnable mapping function Φ(·), and calculate the scheduling weight matrix W between regions d :
[0030]
[0031] Among them, W d The rth row in represents the correlation weight between the rth region and all other regions. The weight matrix is used to perform weighted fusion of regional features to obtain the feature representation after fusion of cross-region information:
[0032]
[0033] Among them, the rth row represents the final feature representation of the rth region after fusing information from other regions;
[0034] After completing the cross-region interaction and obtaining the updated regional feature representation, the information of each local area is Re-merge into global features to ensure smooth subsequent operations.
[0035] Furthermore, when the MSFA framework is trained, normal image samples are randomly selected from the training data set of each category and used as input for training. The mean square error and cosine similarity are used to measure the loss between the reconstructed features and the original extracted features. The total loss function is:
[0036]
[0037] in, is the mean square error, is the cosine similarity, and λ is a hyperparameter that controls the balance between the two losses.
[0038] Furthermore, after the MSFA framework is trained, a pre-trained encoder is used to extract multi-scale feature maps from the input image. And generate the corresponding reconstructed feature map To restore the normal mode of the image; calculate the difference between the original feature map and the reconstructed feature map at each scale, using The two metrics of loss and cosine similarity are used to construct the anomaly score map S j , specifically defined as follows:
[0039]
[0040] Subsequently, the anomaly score maps of all scales are upsampled to the original resolution and averaged to obtain the final anomaly map S:
[0041]
[0042] Among them, L represents the number of layers of feature extraction;
[0043] Furthermore, the maximum pooling operation is applied to the final anomaly map S to extract the score of the most significant abnormal area in the image as the anomaly indicator of the entire image.
[0044] Beneficial effects:
[0045] This paper proposes a multi-scale feature adaptive unsupervised detection framework MSFA. By introducing a global dynamic transformation unit (GDT) and a local region interaction module (LRI), the model significantly enhances the ability to identify key abnormal areas while reducing its reliance on local redundant information. GDT optimizes the feature aggregation process through position adjustment and competition suppression strategies; LRI combines the local self-attention mechanism with the cross-region scheduling mechanism to effectively promote information interaction between different regions. Experimental results show that on two industrial datasets, MVTecAD and VisA, our method not only significantly improves the performance of anomaly detection, but also provides intuitive visualization support for anomaly localization. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a schematic diagram of the overall process of the MSFA framework of the present invention;
[0047] Figure 2 Schematic diagram of the structure of the global dynamic transformation unit (GDT) of the present invention;
[0048] Figure 3 Schematic diagram of the encoder structure of the present invention;
[0049] Figure 4 Visualization results of anomaly localization on the MVTec-AD dataset;
[0050] Figure 5 Visualize the anomaly localization results on the VisA dataset;
[0051] Figure 6 The performance difference of the reconstruction results in the abnormal area. DETAILED DESCRIPTION
[0052] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.
[0053] This paper proposes an unsupervised detection framework based on multi-scale feature adaptation (MSFA). Its core strategy is to force the model to learn the normal mode of multi-scale feature association through the synergy of global feature integration and local feature interaction, thereby avoiding the "learning shortcut" phenomenon. Figure 1 As shown in the figure, the framework process is divided into four stages: feature extraction, region-aware reconstruction, training and loss function, and anomaly detection. Figure 1 The framework shown can be divided into several functional stages. The specific structure and function of each stage are described as follows:
[0054] (1) Feature extraction module: The input image is first passed through a pre-trained CNN backbone network to extract multi-scale intermediate feature maps. All feature maps are flattened and uniformly mapped to a fixed dimension, while position encoding is superimposed to preserve spatial structure information and provide a consistent representation for subsequent modeling.
[0055] (2) Feature reconstruction module: GDT dynamically adjusts the feature interaction strength between different positions based on the positional relationship, and directly establishes global dependency connections without using the traditional self-attention structure. LRI first restores the input features to a feature map and divides them into non-overlapping local regions. It uses regional attention to extract local structures, and then realizes cross-region information fusion through the regional scheduling mechanism and recombines them into a unified feature representation. The two are executed collaboratively in the encoder, and the fusion result is combined with the guided reference representation to generate the final reconstructed feature map.
[0056] (3) Anomaly Detection Module: The reconstructed feature map is compared with the original feature map layer by layer, and the reconstruction error and cosine similarity are calculated to generate a multi-scale anomaly score. All score maps are upsampled to the size of the original image, and the anomaly localization map is finally output.
[0057] This paper uses a convolutional neural network (CNN) pre-trained on ImageNet as a feature extractor to extract intermediate feature maps at multiple scales from the input image. This process does not involve further training of the network and is an offline feature extraction method. Assume that the extracted feature map is: Where j represents the level index of the feature map, C j 、H j 、W j are the number of channels, height, and width of the j-th layer feature map, respectively.
[0058] In order to uniformly represent the feature maps of different layers, we first flatten the feature maps in the spatial dimension so that the feature vector corresponding to each spatial position is represented as an independent unit, and we get: Among them L j =H j ×W j . The dimension of each eigenvector is then transformed from C to j Mapped to a unified dimension D to ensure the uniformity of features: In order to preserve the structural information of each spatial position in the image, we introduce a set of trainable position features for each position And add it element by element to the corresponding eigenvector to get:
[0059]
[0060] Finally, the feature maps of each layer are converted into a representation Z with uniform dimension (j),This representation not only maintains the consistency of spatial ,position information, but also lays the foundation for the subsequent effective ,fusion of features between different layers.
[0061] In order to alleviate the problem that the model is prone to falling into "learning shortcuts" in multi-category unsupervised anomaly detection, the present invention introduces a global dynamic transformation unit (GDT) and a localized region interaction module (LRI) in the encoder to replace the original self-attention mechanism.
[0062] In order to guide the model to learn the distribution of normal patterns, we input the processed extracted features together with the guided reference representation R from the auxiliary knowledge base into the modified encoder. The encoder consists of our proposed global dynamic transformation unit and local region interaction module. After multi-layer encoder processing, the output features are fed into the projection layer to obtain the final reconstructed representation In the standard Transformer, self-attention is calculated as follows:
[0063]
[0064] Among them, Q, K, and V represent query, key, and value vectors respectively, and Softmax is used to normalize the attention weights so that their sum is 1.
[0065] Although Softmax can normalize the attention weights, when QK T When the distribution of values in [ ] fluctuates significantly (for example, the similarity between some feature vectors is much higher than at other locations), Softmax exponentially amplifies the maximum values. This bias causes a small number of locations to dominate the weight, weakening the contribution of other features. This mechanism can lead to biased model training, where the model tends to rely on features at a few key locations for decision-making while neglecting the integration of global information, resulting in learning shortcuts. This bias is particularly pronounced when the input data contains significant local features, further exacerbating the formation of learning shortcuts.
[0066] In order to alleviate the above problems, the present invention proposes a global dynamic transformation unit (GDT) to replace the traditional multi-head self-attention mechanism. Figure 2 The GDT workflow is demonstrated. GDT introduces a competitive inhibition mechanism to limit the dominant influence of local maxima on the overall result. Compared to dot-product attention based on similarity calculations, GDT focuses more on the relative relationships between features. By incorporating structural relationships between positions, GDT can dynamically adjust the impact of features at different positions on the overall output.
[0067] Specifically, in order to optimize the interaction between different position features, GDT introduces a learnable relationship strength B when measuring the relationship between positions. pos1,pos2 , which works together with the key vector to adjust the contribution of each position to the current query position. Its value is automatically determined by the model during training. When a position is far from the query position, the model tends to assign a smaller adjustment value, thereby reducing its influence; while for closer positions, it assigns a larger adjustment value, thereby increasing their influence. In this way, the model can adjust based on the relative importance of each position. The formula is as follows:
[0068] F pos1,pos2 =f(K pos1 ,B pos1,pos2 )#(3)
[0069] Among them, B pos1,pos2 It describes the relative influence between position pos1 and position pos2, and f(·) is the fusion weight function used to integrate the information of each position.
[0070] To alleviate the weight concentration problem caused by extreme scores, GDT first calculates the F corresponding to each position pos1. pos1,pos2 Normalize and get the fusion coefficient α pos1,pos2 On this basis, in order to maintain the importance of the target position’s own characteristics, we further introduce the query vector Q i The relevant adjustment factors are used to weight the fusion results. The final output is as follows:
[0071]
[0072] Among them, σ(·) is the Sigmoid function, which is used to adjust the intensity of the output at each position.
[0073] GDT introduces a competitive inhibition strategy into the attention mechanism to mitigate the over-concentration of attention on a small number of locations, thereby alleviating the problem of attention drift caused by highly similar inputs. This design reduces the model's reliance on local content, thereby inhibiting the formation of learning shortcuts.
[0074] While global weighting enhances feature representation, over-reliance on global information can mask key changes in local regions in anomaly detection tasks. To address this, we introduce a Local Region Interaction (LRI) module to force the model to interact with information within a local region, thereby improving the discriminability and generalization of feature representation. The core idea of this module is to break global dependencies, restricting information interaction within each region to construct independent representations at the regional level.
[0075] Specifically, the input features are first spatially divided into multiple S×S non-overlapping sub-regions X (r) , where r represents the region number. Within each region, a local self-attention mechanism is performed. We use three trainable linear transformation matrices Project the input into the Query, Key, and Value spaces respectively, and calculate the attention weight by scaling the dot product to achieve information aggregation within the region: Q, K, V = X (r) W Q ,X (r) W K ,X (r) W V #(5)
[0076] In each sub-region X (r) In the code, the local feature A obtained by applying the self-attention mechanism is updated (r) Then, to obtain the overall representation at the regional level, the global vector representation of each region is obtained by average pooling. Concatenate the vector representations of all regions according to the region number to obtain the region representation matrix It should be noted that this process is limited to the local area and does not involve cross-regional information interaction.
[0077] After extracting features from each local region, we designed a cross-region information scheduling mechanism that guides information flow between regions by calculating the correlation between pairs of regions. We transform the region-level feature matrix through a learnable mapping function Φ(·) and calculate the scheduling weight matrix W between regions. d :
[0078]
[0079] Among them, W d The rth row in represents the correlation weight between the rth region and all other regions. Subsequently, the weight matrix is used to perform weighted fusion of regional features to obtain the feature representation after fusion of cross-region information:
[0080]
[0081] Among them, the rth row represents the final feature representation of the rth region after fusing information from other regions.
[0082] After completing the cross-region interaction and obtaining the updated regional feature representation, we will Re-merge into global features to ensure smooth subsequent operations.
[0083] The encoder receives input features and reference representation R. The input and reference participate in attention interaction respectively, and the fusion result is normalized and processed by feedforward network (FFN) before output. In this embodiment, the encoder contains four layers with the same structure, such as Figure 3 As shown in Figure 2, the reference representation R is a learnable parameter matrix that provides the model with cross-image commonality information through the cross-image regularities learned during training, which is used to enhance feature selection in the attention mechanism.
[0084] During training, we randomly select normal image samples from the training dataset for each category and use them as input for the model. Each batch of images passes through the encoder network to extract deep features. These extracted features are then used as input to the decoder for subsequent reconstruction. We use mean squared error and cosine similarity to measure the loss between the reconstructed features and the original extracted features.
[0085] The mean square error (MSE) formula measures the difference between the predicted value and the true value, which is expressed as:
[0086]
[0087] in, and Denote the true value and predicted value respectively, and n is the number of samples. This formula measures the accuracy of feature reconstruction by calculating the square error of each sample.
[0088] The cosine similarity formula is used to measure the directional similarity between two vectors. The formula is as follows
[0089]
[0090] in, and is the i-th component of the two feature vectors. This loss function evaluates the directional consistency between the predicted vector and the true vector. The higher the similarity between the two, the closer the cosine similarity value is to 1.
[0091] The total loss function is constructed as follows:
[0092]
[0093] where λ is a hyperparameter that controls the balance between the two losses.
[0094] In the inference phase, we use the pre-trained encoder to extract multi-scale feature maps from the input image And generate the corresponding reconstructed feature map To restore the normal mode of the image. At this point, the reconstructed feature map should be as close as possible to the original normal feature map. Any deviation may indicate an abnormality. In order to quantify this deviation, we calculate the difference between the original feature map and the reconstructed feature map at each scale, using Loss and cosine similarity are two metrics. Construct anomaly score map S j , specifically defined as follows:
[0095]
[0096] Subsequently, the anomaly score maps of all scales are upsampled to the original resolution and averaged to obtain the final anomaly map S:
[0097]
[0098] Where L represents the number of feature extraction layers. To further generate image-level anomaly scores, we apply a max pooling operation to the final anomaly map S to extract the scores of the most significant anomaly regions in the image as the anomaly indicators for the entire image.
[0099] Experimental verification:
[0100] In order to verify the effectiveness of the proposed unsupervised detection framework based on multi-scale feature adaptation, we conducted a [3] The model is evaluated on the dataset and the VisA dataset.
[0101] The MVTec-AD dataset simulates anomaly detection scenarios in real industrial production. As the first comprehensive, multi-target, multi-defect anomaly detection dataset with pixel-level accurate labeling, it contains 15 categories, including 5 texture classes and 10 object classes such as bottles and leather. The dataset has a total of 3,629 normal images as training and 1,725 test images, including 467 normal images and 1,258 anomaly images.
[0102] VisA consists of 9,621 normal images and 1,200 abnormal images, covering 12 different object categories. These objects are further divided into three different structure types: complex structure, multiple instances, and single instance.
[0103] To evaluate the model's performance in anomaly detection and localization tasks, this paper uses the area under the ROC curve (AUROC) as the primary metric. Image-level AUROC (I-AUROC) measures the model's ability to discriminate whether anomalies exist in the entire image, while pixel-level AUROC (P-AUROC) assesses its accuracy in localizing anomalies within a local area.
[0104] The calculation of AUROC is based on the true positive rate (TPR) and false positive rate (FPR) under different decision boundaries, which are defined as follows:
[0105]
[0106] Where TP and TN represent the abnormal and normal samples correctly identified by the model, respectively, while FP and FN correspond to the normal and abnormal samples misclassified. Calculate the corresponding TPR and FPR, and plot the ROC curve based on them. The area of the curve is calculated as follows:
[0107]
[0108] Where n represents the total number of cutoff points used for division, TPR(k) and FPR(k) are the true positive rate and false positive rate corresponding to the kth cutoff point, respectively.
[0109] Image-level calculations are based on the predicted values and image-level labels of each image; pixel-level calculations compare the predicted heat map with the pixel labels pixel by pixel, and the overall process is consistent.
[0110] To compare with existing methods, we trained the model within a unified anomaly detection model. All training and test images were resized to 256×256 pixels. This experiment used EfficientNet-B6 as the feature extraction network, with the feature layer configuration and output feature index length consistent. The encoder and decoder layers of the feature reconstruction module each had four layers. During model training, the Adam optimizer had an initial learning rate of 1e-4 and a training epoch of 200.
[0111] Quantitative results analysis:
[0112] MVTec-AD: Table 1 shows the comparison results of the model and other methods in image-level AUROC and pixel-level AUROC, where all results are dataset-level average results of their respective data subsets.
[0113] Table 1 Comparison results of image-level and pixel-level AUROC on the MVTec-AD dataset
[0114]
[0115]
[0116] Specifically, our model outperforms PatchCore and UniAD by an average of 2.4% and 1.6%, respectively. PatchCore is based on a one-to-one detection approach, while UniAD employs a one-shot approach for anomaly detection. In particular, under the one-shot setting, our model outperforms UniAD by 5.4% and 7.8% on the Cable and Screw categories, respectively. We attribute this improvement to the design of a multi-level attention fusion architecture.
[0117] VisA: Compared to MVTecAD, VisA faces greater challenges due to its more complex structure and scenes containing multiple misalignment instances. Table 2 shows the performance comparison results of our model with the other three methods under a unified experimental setting.
[0118] Table 2 Comparison of image-level and pixel-level AUROC on the VisA dataset
[0119]
[0120]
[0121] Qualitative results analysis:
[0122] Qualitative experiments are conducted on the MVTec-AD and VisA datasets to intuitively demonstrate the performance of our method in the anomaly localization task. Figure 4 As shown in the figure, for the MVTec-AD dataset, the present invention can more accurately detect abnormal areas compared to DiAD in terms of positioning of abnormal areas, while effectively reducing false detections in normal areas.
[0123] On the VisA dataset, Figure 5 As shown in the figure, this method exhibits higher anomaly localization capability than UniAD in complex backgrounds and multi-instance scenarios. In the fryum category, it can detect both edge and internal defects, while UniAD mainly responds to edge areas.
[0124] Ablation experiment:
[0125] To verify the effectiveness of each module, an ablation experiment was designed to gradually introduce the global dynamic transformation unit and the local area interaction module to evaluate their independent contributions to the model performance. All experiments were performed on the MVTec-AD dataset under the same settings. The results are shown in Table 3.
[0126] Table 3 Ablation experiment results on the MVTec-AD dataset
[0127]
[0128] By introducing these two modules at the same time, Image-AP and Pixel-AP are improved by 2.8% and 0.4% respectively compared with the case without adding these modules, which proves the effectiveness of the method of fusing global and local multi-scale features. To intuitively demonstrate the impact of each module, we randomly select several samples in each group of experiments, and their reconstruction effects are shown in the figure below. Figure 6 When only GDT or LRI is added, the reconstruction result still differs from the original image, and the abnormal areas are not effectively restored. When both are introduced simultaneously, the reconstructed image is closer to the normal pattern overall.
[0129] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made in accordance with the spirit of the present invention are intended to be covered by the scope of protection of the present invention.
Claims
1. An unsupervised anomaly detection method based on multi-scale features, characterized in that: The steps include: Obtain an image dataset to be detected and preprocess the image dataset; Construct a multi-scale feature adaptive unsupervised detection framework MSFA, hereinafter referred to as the MSFA framework, which includes a feature extraction module, a feature reconstruction module and anomaly detection module; The feature extraction module processes the input image through the pre-trained network and extracts multi-scale intermediate feature maps. All feature maps are flattened and uniformly mapped to a fixed dimension, and position codes are superimposed. The feature reconstruction module processes the features extracted by the feature extraction module by introducing a global dynamic transformation unit GDT and a local region interaction module LRI. The two are executed collaboratively in the encoder, and the fusion result is combined with the guided reference representation to generate the final reconstructed feature map; The anomaly detection module compares the reconstructed feature map with the original feature map layer by layer, calculates the reconstruction error and cosine similarity, generates a multi-scale anomaly score, and all score maps are upsampled to the size of the original image, and finally outputs an anomaly localization map.
2. The unsupervised anomaly detection method based on multi-scale features according to claim 1, characterized in that: The pre-trained network in the feature extraction module uses a pre-trained convolutional neural network (CNN) as a feature extractor.
3. The unsupervised anomaly detection method based on multi-scale features according to claim 1, characterized in that: All feature maps are flattened and uniformly mapped to a fixed dimension, and position encoding is superimposed as follows: Assume that the extracted feature map is: Where j represents the level index of the feature map, C j 、H j 、W j are the number of channels, height, and width of the j-th layer feature map, respectively. First, the feature map is flattened in the spatial dimension so that the feature vector corresponding to each spatial position is represented as an independent unit, and we get: Among them L j =H j ×W j ; Then, the dimension of each feature vector is transformed from C to j Mapped to a unified dimension D to ensure the uniformity of features: Introducing a set of trainable position features for each position And add it element by element to the corresponding eigenvector to get: Finally, the feature map of each layer is converted into a representation Z with uniform dimension (j) .
4. The unsupervised anomaly detection method based on multi-scale features according to claim 1, characterized in that: The feature reconstruction module introduces a global dynamic transformation unit GDT and a local region interaction module LRI in the encoder, and inputs the extracted features processed by the feature extraction module and the guided reference representation R from the auxiliary knowledge base into the modified encoder. The modified encoder receives the input features and the reference representation R, and the input features and the reference representation R respectively participate in attention interaction. The fusion result is normalized and processed by the feedforward network FFN and then output; the modified encoder contains four layers and the structure remains consistent.
5. The unsupervised anomaly detection method based on multi-scale features according to claim 4, characterized in that: The reference representation R is a learnable parameter matrix that provides the model with common information across images through the cross-image regularities learned during training, which is used to enhance feature selection in the attention mechanism.
6. The unsupervised anomaly detection method based on multi-scale features according to claim 4, characterized in that: The global dynamic transformation unit GDT introduces a competitive inhibition mechanism and introduces a learnable relationship strength B when measuring the relationship between positions. pos1,pos2 , works together with the key vector to adjust the contribution of each position to the current query position. The formula is as follows: F pos1,pos2 =f(K pos1 ,B pos1,pos2 ) Among them, B pos1,pos2 It describes the relative influence between position pos1 and position pos2, and f(·) is the fusion weight function used to integrate the information of each position; GDT first calculates the F corresponding to each position pos1 pos1,pos2 Normalize and get the fusion coefficient α pos1,pos2 On this basis, we further introduce the query vector Q i The relevant adjustment factors are used to weight the fusion results, and the final output is as follows: Among them, σ(·) is the Sigmoid function, which is used to adjust the intensity of the output at each position.
7. The unsupervised anomaly detection method based on multi-scale features according to claim 4, characterized in that: The specific operations of the local region interaction module LRI are: The input features are first divided into multiple S×S non-overlapping sub-regions X (r) , where r represents the region number. Within each region, a local self-attention mechanism is performed using three trainable linear transformation matrices The input is projected into the Query, Key, and Value spaces respectively, and the attention weight is calculated by scaling the dot product to achieve information aggregation within the region: Q,K,V=X (r) W Q ,X (r) W K ,X (r) W V In each sub-region X (r) In the code, the local feature A obtained by applying the self-attention mechanism is updated (r) , the global vector representation of each region is obtained by average pooling Concatenate the vector representations of all regions according to the region number to obtain the region representation matrix Design a cross-region information scheduling mechanism, transform the regional feature matrix through a learnable mapping function Φ(·), and calculate the scheduling weight matrix W between regions d : Among them, W d The rth row in represents the correlation weight between the rth region and all other regions. The weight matrix is used to perform weighted fusion of regional features to obtain the feature representation after fusion of cross-region information: Among them, the rth row represents the final feature representation of the rth region after fusing information from other regions; After completing the cross-region interaction and obtaining the updated regional feature representation, the information of each local area is Re-merge into global features to ensure smooth subsequent operations.
8. The unsupervised anomaly detection method based on multi-scale features according to claim 1, characterized in that: When the MSFA framework is trained, normal image samples are randomly selected from the training data set of each category and used as the input of the MSFA framework for training. The mean square error and cosine similarity are used to measure the loss between the reconstructed features and the original extracted features. The total loss function is: in, is the mean square error, is the cosine similarity, and λ is a hyperparameter that controls the balance between the two losses.
9. The unsupervised anomaly detection method based on multi-scale features according to claim 8, characterized in that: After the MSFA framework is trained, the pre-trained encoder is used to extract multi-scale feature maps from the input image. And generate the corresponding reconstructed feature map To restore the normal mode of the image; calculate the difference between the original feature map and the reconstructed feature map at each scale, using The two metrics of loss and cosine similarity are used to construct the anomaly score map S j , specifically defined as follows: Subsequently, the anomaly score maps of all scales are upsampled to the original resolution and averaged to obtain the final anomaly map S: Among them, L represents the number of layers of feature extraction.
10. The unsupervised anomaly detection method based on multi-scale features according to claim 9, characterized in that: The maximum pooling operation is applied to the final anomaly map S to extract the score of the most significant abnormal area in the image as the anomaly indicator of the entire image.
Citation Information
Cited By
Visual anomaly detection method and system for rail transit operation and maintenance scene
CN121074823A
Gold core particle image generation method and system based on multi-scale feature comparison
CN121120837A
Unsupervised recommendation system anomaly detection method based on multi-scale behavior modeling and boundary perception enhancement
CN121188669A
An unsupervised recommendation system anomaly detection method based on multi-scale behavior modeling and boundary perception enhancement
CN121188669B
In-transit transportation abnormity intelligent early warning system for intelligent supply chain
CN121600486A