Intelligent detection method based on global and local attention and detail enhancement
By introducing global and local attention and detail enhancement technology into the intelligent detection method, the performance degradation of existing anomaly detection methods when the number of categories increases is solved, and efficient multi-category anomaly detection is achieved, reducing computing overhead and data preparation complexity.
Patent Information
- Application Number
- CN202510466455.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
Existing anomaly detection methods increase dramatically when the number of categories increases, and it is difficult to cope with large differences between categories, and rely on a large amount of available data and the requirement that one model be trained for each data set.
Using intelligent detection methods based on global and local attention and detail enhancement, multi-stage features are extracted through a pre-trained encoder with frozen parameters, multi-scale semantic information is fused with the bottleneck layer, and fed into the multi-stage reconstruction network, and reconstructed using the global and local attention module SGL and the detail enhancement module DEH.
A single model is realized that it can effectively detect multiple categories of items simultaneously, reducing dependence on a large number of defective samples, improving reconstruction quality and detection efficiency, and reducing computing overhead.
Smart Images

Figure CN119991528A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial visual intelligent detection, and in particular to an intelligent detection method based on global and local attention and detail enhancement. Background Art
[0002] The rise of smart manufacturing has greatly enhanced the key role of industrial visual defect anomaly detection in the production process. This technology is not only expected to significantly improve efficiency and reduce manual inspection expenses, but also significantly enhance product quality and the stability of production lines. However, in today's manufacturing field, abnormal samples occur less frequently and are difficult to collect, which poses a considerable challenge to traditional supervised training methods. Therefore, if an effective anomaly detection model can be trained based only on normal samples, it will greatly simplify the tedious process of data preparation, thereby saving the time and effort required to label abnormal samples.
[0003] With the continuous advancement of deep learning, anomaly detection technology has been increasingly improved, significantly improving production efficiency and product quality. This intelligent detection method not only greatly reduces the need for manual intervention, but also can accurately predict problems before they occur, thereby effectively reducing the risk of equipment failure and downtime, and providing solid support for the transformation and upgrading of enterprise intelligent manufacturing.
[0004] Currently, most anomaly detection methods mainly adopt a single-category setting, that is, a model needs to be trained and tested separately for each category. These methods cover a variety of techniques such as reconstruction, single-class classification, and knowledge extraction. Although these methods perform well in some scenarios, their reliance on large amounts of data and the requirement to train a model for each dataset causes a sharp increase in time and memory consumption when the number of categories increases, and it is difficult to cope with situations where there are large differences between categories. Although the recent introduction of multi-category anomaly detection technology has made some progress, there is still room for improvement in balancing accuracy and efficiency. Summary of the invention
[0005] To this end, the present invention provides an intelligent detection method based on global and local attention and detail enhancement to solve the problems raised in the background technology.
[0006] In order to achieve the above object, the present invention provides the following technical solutions: based on the intelligent detection method of global and local attention and detail enhancement, multi-stage features are extracted by inputting samples through a pre-trained encoder with frozen parameters, and then multi-scale semantic information is fused using a bottleneck layer, and then fed into a multi-stage reconstruction network; Each stage of the reconstruction network consists of multiple FE modules connected in series, one of which is a reconstruction component, which consists of a global and local attention module SGL and a detail enhancement module DEH. The global and local attention module SGL receives the output of the previous FE module or the previous reconstruction stage as input, extracts global and local semantic information, and the output result is enhanced in detail by the detail enhancement module DEH to improve the reconstruction quality. The global and local attention module SGL contains parallel global GA and local LA branches. The global GA captures global semantic information by multiple series of cross-axis attention SCA, and the local LA contains two large-kernel deep convolutions for local enhancement. The core of the detail enhancement module DEH consists of a channel attention CA and a multi-dimensional differential convolution DEConv; Channel attention CA first passes through global average pooling AP and global maximum pooling MP to obtain channel global information, learns attention weight distribution through two convolution and activation functions, and finally obtains attention score through Sigmoid and multiplies it with input features; The multidimensional differential convolution DEConv contains one ordinary convolution and four differential convolutions, namely center differential convolution CDC, angular differential convolution ADC, horizontal differential convolution VDC, and vertical differential convolution HDC. Ordinary convolution is used to obtain intensity level information, and differential convolution is used to enhance gradient level information.
[0007] Preferably, the global and local attention modules SGL adopt a multi-branch parallel structure, in which the cross-axis attention SCA is used to capture global information, thereby effectively modeling global normal semantic information; in this structure, the query Q, key K and value V in the attention mechanism are calculated by strip kernel depth 1D convolution; compared with the traditional fully connected layer, this method significantly reduces the number of parameters. Longer 1D convolution kernels can more effectively capture the dependencies between long-range pixels and improve the model's ability to understand contextual information; set For the The input characteristics of the SCA module are as follows Axis and The query Q, key K, and value V of an axis are calculated as follows: (1) (2) in and Along Axis and The size of the strip kernel in the axial dimension is 1D depth convolution, is a hyperparameter, Representing the normalized layer, the single cross-axis attention SCA is calculated as follows: (3) (4) (5) All the formulas (1)-(4) Shared weight parameters, is a learnable scale parameter that controls the size of the Q, K matrix multiplication before applying the Softmax function; for Convolutional blocks; Set up The input characteristics of the FE module in the stage are , global attention module The calculation is as follows: (6) The local attention branch of the global and local attention module SGL contains two sets of convolution kernels with sizes , The depth convolution is used to focus on local details. The SGL calculation of the stage FE module is as follows: (7).
[0008] Preferably, the five parallel convolutions of the multi-dimensional differential convolution DEConv will inevitably lead to an increase in parameters and inference time. By utilizing the additivity of convolution, the parallel deployed convolutions are simplified to a single standard convolution. The multi-dimensional differential convolution DEConv is calculated as follows: (8) Among them is Input features, Represents five types of convolution structures, is the convolution operation, Represents the transformation kernel composed of parallel convolutions; the overall calculation of the detail enhancement module DEH is as follows: (9).
[0009] Preferably, the reconstruction of each stage of the reconstruction network is performed by stacking multiple FE modules to guide the reconstruction; (10).
[0010] Preferably, MSE loss is used to optimize the multi-scale feature reconstruction module, and the loss function is The definition is as follows: (11) in and They are the characteristics of each stage of the encoder and reconstruction network respectively.
[0011] The present invention has the following advantages: Most current defect detection methods rely on training a separate model for each category, which requires a large amount of available data, and as the number of categories increases, time and memory consumption will also increase significantly, and perform poorly when there is a large diversity between classes. The present invention introduces a new multi-class unsupervised defect framework, where a single model can effectively detect multiple categories of items at the same time. The framework is designed based on the concept of reconstruction. The model training process of the present invention only needs to rely on normal samples, which makes it more convenient on increasingly mature industrial production lines, without the need to collect a large number of defective samples for model learning.
[0012] The present invention introduces a new reconstruction component with simple structure and flexible use. It can achieve excellent reconstruction quality while still controlling the computational overhead within an acceptable range. In practical applications, whether it is real-time monitoring of production lines or equipment status detection, the method of the present invention can quickly identify potential anomalies and provide clear visualization results to help users quickly locate problems.
[0013] Therefore, the present invention divides the reconstruction into two stages, the global and local attention module SGL and the detail enhancement module DEH, and innovatively introduces deep strip kernel convolution in the global attention. Combined with cross attention, it effectively models the global information while effectively controlling parameters and computational complexity, and uses multi-scale square deep convolution in parallel to enhance the extraction of local information.
[0014] In response to the reconstruction paradigm's pursuit of quality, differential convolution is innovatively introduced to strengthen gradient-level information and enhance the details of global and local features. SGL and DEH complement each other. A large number of experiments have proved the rationality and effectiveness of the design of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A schematic diagram of the overall architecture provided by the present invention; Figure 2 A schematic diagram of a reconstruction assembly provided by the present invention; Figure 3 This is a visual comparison schematic diagram provided by the present invention. DETAILED DESCRIPTION
[0016] The following is a description of the implementation of the present invention by specific embodiments. People familiar with the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0017] The current defect detection paradigm focuses on building a model separately on each target dataset, which brings inconvenience to the real production environment as the number of categories increases. Therefore, a multi-class unsupervised defect anomaly detection framework is needed.
[0018] The present invention follows the following paradigm. Due to the unavailability of real abnormal labels, a feature extractor with strong generalization ability is used in the training stage to extract multi-stage features of non-abnormal input and transmit them to the reconstruction network. By optimizing the distance between the corresponding stage features of the feature extractor and the reconstruction network, the accuracy and robustness of the model in non-abnormal scenarios are improved. In the reasoning stage, when the model faces abnormal information, the feature extractor has excellent generalization ability and can effectively identify and retain abnormal semantics, while the reconstruction network is only trained under the condition of no abnormal information, so the features between the two show significant differences. This differentiated feature performance provides a reliable basis for anomaly detection. By using the similarity between features, abnormal images can be effectively extracted and potential abnormal information can be revealed. However, the excellent generalization ability of neural networks can still play an excellent role in the face of data that has never been seen, which makes the reconstruction network insensitive to some anomalies, so there is a potential risk of reconstruction anomalies, affecting the detection effect.
[0019] The present invention proposes a global and local attention module SGL to model long-distance dependencies and local features, effectively mask abnormal information, and simultaneously use the detail enhancement module DEH to sharpen details and improve reconstruction quality. Specifically, SGL contains cross-axis attention based on multi-scale strip kernel deep convolution, which can strike a good balance between feature capture and parameters. DEH uses grouped differential convolution to calculate pixel differences in feature maps, and then convolves with convolution kernels, which can effectively encode prior information explicitly into convolution.
[0020] The method of the present invention has excellent reconstruction quality and controls the number of parameters and the amount of calculation within an acceptable range. Through experiments on two benchmark data sets and comparison with some state-of-the-art methods, the method of the present invention has better performance.
[0021] The details are as follows: The multi-class defect detection framework proposed in this embodiment is as follows: Figure 1 As shown, it includes an encoder and a decoder.
[0022] First, the input sample is passed through a pre-trained encoder with frozen parameters to extract multi-stage features, and then a bottleneck layer is used to fuse multi-scale semantic information, and then fed into a multi-stage reconstruction network. Each stage of the reconstruction network contains multiple FE modules, among which the global and local attention modules SGL model global and local information, and the detail enhancement module DEH performs detail enhancement. During the training process, the sum of the MSE of the features of the encoder and decoder at different stages is used as the loss function, and the sum of the cosine similarity is used to obtain the abnormal map in the inference stage. Figure 2 Shows some details of the reconstruction components, including the module shown on the far right. Convolutional blocks, use Convolution, if Depthwise convolutional blocks are used Depthwise convolution, and so on.
[0023] The global and local attention module SGL adopts a multi-branch parallel structure, in which the cross-axis attention SCA is used to capture global information, thereby effectively modeling global normal semantic information; in this structure, the query Q, key K and value V in the attention mechanism are calculated through strip kernel depth 1D convolution; compared with the traditional fully connected layer, this method significantly reduces the number of parameters. Longer 1D convolution kernels can more effectively capture the dependencies between long-range pixels, improving the model's ability to understand contextual information; set For the The input characteristics of the SCA module are as follows Axis and The query Q, key K, and value V of an axis are calculated as follows: (1) (2) in, , Used as Axis and The Q, K, and V of the axis, that is, the Q, K, and V values in one axis direction are the same. For normal convolution, and Along Axis and The size of the strip kernel in the axial dimension is 1D depth convolution, is a hyperparameter, Representing layer normalization, the single cross-axis attention SCA is calculated as follows: (3) (4) (5) All the formulas (1)-(4) Shared weight parameters, is a learnable scale parameter, , It is the result of the cross attention calculation, which is used to control the size of the Q, K matrix multiplication before applying the Softmax function; for Convolutional blocks, including Convolutional layer, normalization layer and activation function, refer to Figure 2 The far right side includes an instance normalization. The goal of this embodiment is to establish a multi-class reconstruction component. It only needs to learn the mean and variance of a single category instance, so batch normalization is not suitable. The same applies to the following.
[0024] Set up The input characteristics of the FE module in the stage are , global attention module The calculation is as follows: (6) in, It represents the number of SCA stacked in a FE module. The local attention branch of the global and local attention module SGL contains two sets of convolution kernel sizes. , The depth convolution is used to focus on local details. The SGL calculation of the stage FE module is as follows: (7); in, , They are , The depth convolution blocks of each include normalization layers and activation functions.
[0025] The real abnormal area is usually smaller than the global one. The previous strip convolution is long, and the convolution kernel in the local attention is large for the deep layer of the network, which may cause the loss of resolution and details. Figure 2 As shown in the figure, the five parallel convolutions of the multi-dimensional differential convolution DEConv will inevitably lead to an increase in parameters and inference time. By utilizing the additivity of convolution, the parallel deployed convolutions are simplified to a single standard convolution. The multi-dimensional differential convolution DEConv is calculated as follows: (8) Among them is Input features, Represents five types of convolution structures, is the convolution operation, represents the transformation kernel composed of parallel convolutions; the detail enhancement module DEH also contains a channel attention CA, such as Figure 2 As shown on the far left, AP and MP are global average pooling and global maximum pooling, respectively, which are used to dynamically adjust the importance of each channel, strengthen feature expression, and highlight the features with the most effective normal information, thereby suppressing noise and irrelevant information; the overall calculation of the detail enhancement module DEH is as follows: (9).
[0026] in, , They are Ordinary convolutional blocks, and The differential convolution blocks all contain normalization layers and activation functions.
[0027] The reconstruction of each stage of the reconstruction network is composed of a stack of multiple FE modules to guide the reconstruction; (10).
[0028] Use MSE loss to optimize the multi-scale feature reconstruction module, the loss function The definition is as follows: (11) in and They are the characteristics of each stage of the encoder and reconstruction network respectively.
[0029] The evaluation indicators are set as follows: 1. AUROC First, we use image-level AUROC and pixel-level AUROC. AUROC (Area Under Receiver Operating Characteristic Curve) is a widely used indicator, mainly used to evaluate the performance of binary classification models. It quantifies the classification ability of the model by calculating the area under the receiver operating characteristic curve (ROC curve).
[0030] The horizontal axis of the ROC curve is the "False Positive Rate" (FPR), which is defined as the ratio of the number of negative samples that are mistakenly classified as positive to the number of all negative samples; the vertical axis is the "True Positive Rate" (TPR), also known as the recall rate, which is defined as the ratio of the number of positive samples that are correctly classified as positive to the number of all positive samples.
[0031] (12); 2. AP AP is the area under the PR curve, P is the precision, and R is the recall. Image-level and pixel-level AP are also used to evaluate the model of the present invention.
[0032] (13) 3. F1_max In a dataset with imbalanced categories, the F1 score is very useful. F1 is the harmonic mean of P and R and is calculated as follows: (14) In multi-classification or when predicting with different thresholds, multiple F1 scores can be calculated, and F1_maax is calculated as follows: (15) 4. AU-PRO Compared to AUROC, AU-PRO usually provides more meaningful evaluation when dealing with an imbalance of positive and negative samples, especially when there is a trade-off between precision and recall.
[0033] (16) here Represents the corresponding P under given R conditions.
[0034] Experimental environment: The experimental environment is the basic condition for conducting experiments. The experimental environment of this embodiment is described in detail as follows: Table 1 Experimental environment
[0035] Parameter settings: Table 2 Experimental settings
[0036] Experimental results: This embodiment uses two data sets, MVTec-AD and VisA, to conduct experiments. By comparing with some of the most advanced methods, as shown in Table 3, mAD is the mean of various indicators. The model of this embodiment has achieved significant improvements in various indicators. These methods include RD4AD, UniAD, SimpleNet, DeSTSeg, DiAD. These methods include traditional convolutional neural networks, visual Transformer, diffusion model, etc. The model of this embodiment is not only simple in structure, but also performs well.
[0037] Thanks to the effectiveness of strip kernel convolution in capturing long-range dependencies and its own lightness, this embodiment cleverly uses ordinary square convolution and enhances the reconstruction quality by the detail enhancement module.
[0038] In addition, this embodiment makes a detailed comparison of the model parameter quantity and the amount of calculation, as shown in Table 4. Although this embodiment has not made outstanding progress in terms of model parameters, it can control the amount of calculation at a low level, and this embodiment also significantly improves the effect of defect detection.
[0039] like Figure 3 As shown, this embodiment compares the defect location effect of the model of this embodiment (OURS) on two data sets. Other methods have the problem of false detection. The method of this embodiment is very sensitive to defects and has a high accuracy in complex cables and circuit boards. Figure 3 The data set given at the bottom is well positioned, which reflects the effectiveness and high quality of the reconstructed network in this embodiment.
[0040] Table 3 Comparison results of seven evaluation index experiments Table 4 Parameters and computational complexity of other methods and performance comparison on MVTec-AD
[0041] MVTec-AD and contains 5354 images from 15 categories, covering 10 object categories and 5 texture categories. VisA contains 9621 normal images and 1200 abnormal images from 12 kinds of industrial metal products, which have different lighting conditions, uneven backgrounds, and multiple products in each image. The texture structure of some categories is complex and challenging. The specific information of MVTec-AD and VisA is shown in Tables 5 and 6.
[0042] Table 5 The number of normal and abnormal numbers in the training set and test set in MVTec-AD Table 6 The number of normal and abnormal numbers in the training set and test set in VisA
[0043] Although the present invention has been described in detail above by general description and specific embodiments, it is obvious to those skilled in the art that some modifications or improvements can be made to the present invention. Therefore, these modifications or improvements made without departing from the spirit of the present invention all belong to the scope of protection claimed by the present invention.
Claims
1. An intelligent detection method based on global and local attention and detail enhancement, characterized by: The multi-stage features are extracted by inputting the sample through a pre-trained encoder with frozen parameters, and then the multi-scale semantic information is fused using a bottleneck layer, which is then fed into a multi-stage reconstruction network. Each stage of the reconstruction network consists of multiple FE modules connected in series, one of which is a reconstruction component, which consists of a global and local attention module SGL and a detail enhancement module DEH. The global and local attention module SGL receives the output of the previous FE module or the previous reconstruction stage as input, extracts global and local semantic information, and the output result is enhanced in detail by the detail enhancement module DEH to improve the reconstruction quality. The global and local attention module SGL contains parallel global GA and local LA branches. The global GA captures global semantic information by multiple series of cross-axis attention SCA, and the local LA contains two large-kernel deep convolutions for local enhancement. The core of the detail enhancement module DEH consists of a channel attention CA and a multi-dimensional differential convolution DEConv; Channel attention CA first passes through global average pooling AP and global maximum pooling MP to obtain channel global information, learns attention weight distribution through two convolution and activation functions, and finally obtains attention score through Sigmoid and multiplies it with input features; The multidimensional differential convolution DEConv contains one ordinary convolution and four differential convolutions, namely center differential convolution CDC, angular differential convolution ADC, horizontal differential convolution VDC, and vertical differential convolution HDC. Ordinary convolution is used to obtain intensity level information, and differential convolution is used to enhance gradient level information.
2. The intelligent detection method based on global and local attention and detail enhancement according to claim 1, characterized in that: The global and local attention module SGL adopts a multi-branch parallel structure, in which the cross-axis attention SCA is used to capture global information, thereby effectively modeling global normal semantic information; in this structure, the query Q, key K and value V in the attention mechanism are calculated through strip kernel depth 1D convolution; let For the The characteristics of the SCA module input are as follows Axis and The query Q, key K, and value V of an axis are calculated as follows: (1) (2) in, , Used as Axis and The Q, K, and V of the axis, that is, the Q, K, and V values in one axis direction are the same. For normal convolution, and Along Axis and The size of the strip kernel in the axial dimension is 1D depthwise convolution, is a hyperparameter, Representing the normalized layer, the single cross-axis attention SCA is calculated as follows: (3) (4) (5) All the formulas (1)-(4) Shared weight parameters, , is the result of cross attention calculation, is a learnable scale parameter that controls the size of the Q, K matrix multiplication before applying the Softmax function; for Convolutional blocks, including Convolutional layers, normalization layers, and activation functions; Set up The input characteristics of the FE module in the stage are , global attention module The calculation is as follows: (6) in, It represents the number of SCA stacked in a FE module. The local attention branch of the global and local attention module SGL contains two sets of convolution kernel sizes. , The depth convolution is used to focus on local details. The SGL calculation of the stage FE module is as follows: (7); in, , They are , The depth convolution blocks of each include normalization layers and activation functions.
3. The intelligent detection method based on global and local attention and detail enhancement according to claim 1, characterized in that: The five parallel convolutions of multi-dimensional differential convolution DEConv will inevitably lead to an increase in parameters and inference time. By utilizing the additivity of convolution, the parallel deployed convolutions are simplified to a single standard convolution. The calculation of multi-dimensional differential convolution DEConv is as follows: (8) Among them is Input features, Represents five types of convolutional structures, is the convolution operation, Represents the transformation kernel composed of parallel convolutions; the overall calculation of the detail enhancement module DEH is as follows: (9); in, , They are Ordinary convolutional blocks, and The differential convolution blocks all contain normalization layers and activation functions.
4. The intelligent detection method based on global and local attention and detail enhancement according to claim 1, characterized in that: The reconstruction of each stage of the reconstruction network is composed of a stack of multiple FE modules to guide the reconstruction; (10)。 5. The intelligent detection method based on global and local attention and detail enhancement according to claim 1, characterized in that: Use MSE loss to optimize the multi-scale feature reconstruction module, the loss function The definition is as follows: (11) in and They are the characteristics of each stage of the encoder and reconstruction network respectively.
Citation Information
Patent Citations
Double-resolution real-time semantic segmentation method based on detail enhancement
CN117409412A
Medical image automatic segmentation method of U-shaped network based on fusion convolution and attention mechanism
CN117474866A
Expression recognition method based on attention-modulated contextual spatial information
WO2023185243A1
Cited By
Defect detection framework based on differential convolution attention and staged feature reconstruction
CN121074051A
Defect detection framework based on differential convolution attention and phased feature reconstruction
CN121074051B