Intelligent Detection Method Based on Global and Local Attention and Detail Enhancement
By adopting intelligent detection methods with global and local attention and detail enhancement in industrial vision anomaly detection, the existing methods' performance degradation and cumbersome data preparation problems when the number of categories increases, and efficient multi-category anomaly detection and simplified data preparation process are achieved.
Patent Information
- Application Number
- CN202510466455.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-04-15
AI Technical Summary
When the number of categories increases, the time and memory consumption of existing industrial vision anomaly detection methods increase sharply, and it is difficult to cope with large differences between categories. They rely on a large number of defective samples for model learning, making data preparation cumbersome.
Using intelligent detection methods based on global and local attention and detail enhancement, multi-stage features are extracted through a pre-trained encoder with frozen parameters, multi-scale semantic information is fused using a bottleneck layer, and fed back into the multi-stage reconstruction network. Each stage of the reconstruction network is connected in series by multiple FE modules, including the global and local attention module SGL and the detail enhancement module DEH.
A single model is realized to effectively detect multiple categories of items simultaneously, reducing dependence on a large number of defective samples, simplifying the data preparation process, improving detection efficiency and reconstruction quality, and controlling computing overhead within an acceptable range.
Smart Images

Figure CN119991528B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial vision intelligent detection, and specifically relates to an intelligent detection method based on global and local attention and detail enhancement. Background Art
[0002] The rise of intelligent manufacturing has greatly enhanced the key role of industrial vision defect anomaly detection in the production process. This technology not only promises to significantly improve efficiency and cut labor inspection costs, but also can significantly enhance product quality and the stability of the production line. However, in today's manufacturing field, the occurrence frequency of abnormal samples is low and the acquisition is difficult, which poses a great challenge to traditional supervised training methods. Therefore, if an effective anomaly detection model can be trained only relying on normal samples, it will greatly simplify the cumbersome process of data preparation, thus saving the time and effort required for labeling abnormal samples.
[0003] With the continuous progress of deep learning, anomaly detection technology has been increasingly improved, significantly improving production efficiency and product quality. This intelligent detection method not only greatly reduces the need for manual intervention, but also can accurately predict before problems occur, thus effectively reducing the risk of equipment failure and downtime, providing a solid support for the transformation and upgrading of enterprise intelligent manufacturing.
[0004] Currently, most anomaly detection methods mainly adopt a single-class setting, that is, a model needs to be trained and tested separately for each class. These methods cover a variety of technologies such as reconstruction, one-class classification, and knowledge extraction. Although these methods perform outstandingly in some scenarios, their dependence on a large amount of data and the requirement of training a model for each data set make the time and memory consumption increase sharply when the number of classes increases, and it is difficult to handle the situation where the differences between classes are large. Although the introduction of multi-class anomaly detection technology has made some progress recently, there is still room for further improvement in balancing accuracy and efficiency. Summary of the Invention
[0005] Therefore, the present invention provides an intelligent detection method based on global and local attention and detail enhancement to solve the problems raised in the background art.
[0006] To achieve the above object, the present invention provides the following technical solution: An intelligent detection method based on global and local attention and detail enhancement, which extracts multi-stage features through a pre-trained encoder with frozen parameters for the input samples, then uses a bottleneck layer to fuse multi-scale semantic information, and then feeds it back to the multi-stage reconstruction network;
[0007] Each stage of the reconstruction network is composed of multiple FE modules connected in series. One FE module is a reconstruction component, which consists of a global and local attention module SGL and a detail enhancement module DEH. The global and local attention module SGL receives the output of the previous FE module or the previous reconstruction stage as input, extracts global and local semantic information, and the output result is enhanced in detail by the detail enhancement module DEH to improve the reconstruction quality.
[0008] The global and local attention module SGL contains parallel global GA and local LA branches. The global GA captures global semantic information through multiple cross-axis attention SCAs connected in series, and the local LA contains two large-kernel depth convolutions for enhancing the local area.
[0009] The core of the detail enhancement module DEH contains a channel attention CA and a multi-dimensional differential convolution DEConv.
[0010] The channel attention CA first undergoes global average pooling AP and global max pooling MP to obtain channel global information, learns the attention weight distribution through two convolutions and activation functions, and finally obtains the attention score through Sigmoid and multiplies it with the input features.
[0011] The multi-dimensional differential convolution DEConv contains an ordinary convolution and four differential convolutions, namely the central differential convolution CDC, the angular differential convolution ADC, the horizontal differential convolution VDC, and the vertical differential convolution HDC. The ordinary convolution is used to obtain intensity-level information, and the differential convolutions are used to enhance gradient-level information.
[0012] Preferably, the global and local attention module SGL adopts a multi-branch parallel structure, where the cross-axis attention SCA is used to capture global information, thereby effectively modeling global normal semantic information. In this structure, the query Q, key K, and value V in the attention mechanism are calculated through strip-kernel depth 1D convolution. Compared with the traditional fully connected layer, this method significantly reduces the number of parameters. The longer 1D convolution kernel can more effectively capture the long-range pixel dependencies and improve the model's ability to understand context information. Let E i be the feature input to the i-th SCA module. The query Q, key K, and value V along the x-axis and y-axis are calculated as follows:
[0013]
[0014] where and are 1D depth convolutions with strip-kernel size x along the s axis and y-axis respectively, s is a hyperparameter, and LN(·) represents layer normalization. Then the calculation of a single cross-axis attention SCA is as follows:
[0015]
[0016] All Conv in Formulas (1)-(4) 1×1 (·) share weight parameters, τ is a learnable scale parameter used to control the size of the Q, K matrix multiplication before applying the Softmax function; ConvBlock 1×1 is a 1×1 convolutional block;
[0017] Let the input feature of the FE module in the k-th stage be Global attention module GA k is calculated as follows:
[0018]
[0019] The local attention branch of the global and local attention module SGL contains two groups of depthwise convolutions with kernel sizes of 5×5 and 7×7 respectively for focusing on local details. The SGL calculation of the FE module in the k-th stage is as follows:
[0020]
[0021] Preferably, the five convolutions in parallel of the multi-dimensional differential convolution DEConv will necessarily lead to an increase in parameters and inference time. Utilizing the additivity of convolutions, the convolutions deployed in parallel are simplified to a single standard convolution. The multi-dimensional differential convolution DEConv is calculated as follows:
[0022]
[0023] where F is the input feature, K i=1:5 represents five convolutional structures, * is the convolution operation, K tran represents the conversion kernel formed by combining the parallel convolutions; The overall calculation of the detail enhancement module DEH is as follows:
[0024] DEH(F) = Conv 1×1 (CA(F) + ConvBlock 3x3 (CA(F) + DEConvBlock 3×3 (CA(F)))) (9).
[0025] Preferably, the reconstruction of each stage of the reconstruction network is achieved by stacking multiple FE modules to guide the reconstruction;
[0026]
[0027] Preferably, the mean squared error (MSE) loss is used to optimize the multi-scale feature reconstruction module. The loss function L is defined as follows:
[0028]
[0029] where I i and R i are the features of each stage of the encoder and the reconstruction network respectively.
[0030] The present invention has the following advantages:
[0031] Currently, most defect detection methods mainly rely on training a separate model for each category, which requires a large amount of available data. Moreover, as the number of categories increases, the time and memory consumption also increase significantly, and they perform poorly in the case of large inter-class diversity. The present invention introduces a new multi-class unsupervised defect framework, where a single model can effectively detect multiple categories of items simultaneously. This framework is designed based on the reconstruction concept. The model training process of the present invention only needs to rely on normal samples, which makes it more convenient in the increasingly mature industrial production line, without the need to collect a large number of defective samples for model learning.
[0032] The present invention introduces a new reconstruction component, which has a simple structure and flexible use. While achieving excellent reconstruction quality, it still controls the computational overhead within an acceptable range. In practical applications, whether it is real-time monitoring of the production line or equipment status detection, the method of the present invention can quickly identify potential anomalies and provide clear visualization results to help users quickly locate problems.
[0033] Therefore, the present invention divides the reconstruction into two stages, the global and local attention module SGL and the detail enhancement module DEH. It innovatively introduces depth strip kernel convolution in the global attention, combines cross-attention, effectively models the global information while effectively controlling the parameters and computational amount, and uses multi-scale square depth convolution in parallel to strengthen the extraction of local information.
[0034] In response to the pursuit of quality in the reconstruction paradigm, differential convolution is innovatively introduced to strengthen gradient-level information and enhance the details of global and local features. SGL and DEH complement each other, and through a large number of experiments, the rationality and effectiveness of the design of the present invention are proved. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is the overall architecture schematic diagram provided by the present invention;
[0036] Figure 2 is the schematic diagram of the reconstruction component provided by the present invention;
[0037] Figure 3 is the schematic diagram of visual comparison provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in this technology can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0039] The current defect detection paradigm focuses on separately constructing a model for each target dataset. With the increase in the number of categories, it brings inconvenience to the actual production environment. Therefore, a multi-class unsupervised defect anomaly detection framework is needed.
[0040] The present invention follows the following paradigm. Due to the unavailability of real anomaly labels, in the training stage, a feature extractor with strong generalization ability is used to extract multi-stage features of anomaly-free inputs and send them to the reconstruction network. By optimizing the distance between the corresponding stage features of the feature extractor and the reconstruction network, the accuracy and robustness of the model in the anomaly-free scenario are improved. In the inference stage, when the model faces anomaly information, because the feature extractor has excellent generalization ability and can effectively identify and retain anomaly semantics, while the reconstruction network is only trained under the condition of anomaly-free information, there is a significant difference in the features between the two. This differential feature performance provides a reliable basis for anomaly detection. By using the similarity between features, abnormal images can ultimately be effectively extracted to reveal potential anomaly information. However, nowadays, the excellent generalization ability of neural networks can still play excellent performance when facing unseen data, which makes the reconstruction network not sensitive to some anomalies, thus there is a potential risk of reconstructing anomalies, affecting the detection effect.
[0041] The present invention proposes a global and local attention module SGL to model long-range dependencies and local features, effectively mask anomaly information, and at the same time use a detail enhancement module DEH to sharpen details and improve the reconstruction quality. Specifically, SGL includes cross-axis attention based on multi-scale strip kernel depth convolution, which can well balance between feature capture and parameters. DEH uses grouped differential convolution to calculate the pixel differences in the feature map and then convolves with the convolution kernel, which can effectively encode prior information explicitly into the convolution.
[0042] The method of the present invention has excellent reconstruction quality and controls the number of parameters and the amount of computation within an acceptable range. By experimenting on two benchmark datasets and comparing with some state-of-the-art methods, the method of the present invention has better performance.
[0043] Specifically as follows:
[0044] The multi-class defect detection framework proposed in this embodiment is as Figure 1As shown, it includes an encoder and a decoder.
[0045] First, the input samples are passed through a pre-trained encoder with frozen parameters to extract multi-stage features. Then, a bottleneck layer is used to fuse multi-scale semantic information, which is then fed back to a multi-stage reconstruction network. Each stage of the reconstruction network contains multiple FE modules, where the global and local attention module SGL models global and local information, and the detail enhancement module DEH enhances details. During the training process, the sum of the MSE of the features in different stages of the encoder and decoder is used as the loss function. In the inference stage, the sum of cosine similarities is used to obtain the anomaly map. Figure 2 Some details of the reconstruction components are shown. For the module shown on the far right, if it is a 3×3 convolutional block, 3×3 convolution is used; if it is a 5×5 depth convolutional block, 5×5 depth convolution is used, and so on.
[0046] The global and local attention module SGL adopts a multi-branch parallel structure, where the cross-axis attention SCA is used to capture global information, thus effectively modeling global normal semantic information. In this structure, the query Q, key K, and value V in the attention mechanism are calculated through strip kernel depth 1D convolution. Compared with the traditional fully connected layer, this method significantly reduces the number of parameters. The longer 1D convolution kernel can more effectively capture the long-range pixel dependencies, improving the model's ability to understand context information. Let E i be the feature input to the i-th SCA module. The query Q, key K, and value V along the x-axis and y-axis are calculated as follows:
[0047]
[0048] Among them, are used as Q, K, and V along the x-axis and y-axis respectively, that is, the Q, K, and V values in one axis direction are the same. Conv1 1×1 (·) is a normal 1×1 convolution, and are 1D depth convolutions with strip kernel sizes of x along the s axis and y-axis respectively. s is a hyperparameter, LN(·) represents layer normalization. Then the calculation of a single cross-axis attention SCA is as follows:
[0049]
[0050] All Conv 1×1 (·) in formulas (1)-(4) share weight parameters. τ is a learnable scale parameter. is the result of cross-attention calculation, which is used to control the size of the Q, K matrix multiplication before applying the Softmax function. ConvBlock 1×1is a 1×1 convolution block, including a 1×1 convolution layer, a normalization layer and an activation function. Refer to Figure 2 On the far right, it includes an instance normalization. The goal of this embodiment is to establish a multi-class reconstruction component, which only needs to learn the mean and variance of a single class instance, and batch normalization is not suitable. The same applies hereinafter.
[0051] Let the input feature of the FE module in the k-th stage be , and the global attention module GA k is calculated as follows:
[0052]
[0053] Among them, N represents the number of FE modules stacked with SCA. The local attention branch of the global and local attention module SGL contains two groups of depthwise convolutions with kernel sizes of 5×5 and 7×7 respectively for focusing on local details. The SGL of the FE module in the k-th stage is calculated as follows:
[0054]
[0055] Among them, DWConvBlock 5×5 (·), DWConvBlock 7×7 (·) are depthwise convolution blocks of 5×5 and 7×7 respectively, both of which include a normalization layer and an activation function.
[0056] The real abnormal area is usually relatively small compared to the global area. The previous strip-shaped convolution is long, and the convolution kernel in the local attention is relatively large for the deep layer of the network, resulting in the problem of losing resolution and details. As Figure 2 shown, the five convolutions in parallel of the multi-dimensional difference convolution DEConv will surely lead to an increase in parameters and inference time. Using the additivity of convolution, the convolutions deployed in parallel are simplified to a single standard convolution. The multi-dimensional difference convolution DEConv is calculated as follows:
[0057]
[0058] Among them, F is the input feature, K i=1:5 represents five convolution structures, * is the convolution operation, and K tran represents the conversion kernel formed by combining parallel convolutions; the detail enhancement module DEH also includes a channel attention CA, as Figure 2 shown on the far left, where AP and MP are global average pooling and global max pooling respectively, which are used to dynamically adjust the importance of each channel, strengthen feature expression, highlight the features with the most effective normal information, and thus suppress noise and irrelevant information; the overall calculation of the detail enhancement module DEH is as follows:
[0059] DEH(F) = Conv 1×1(CA(F)+ConvBlock 3x3 (CA(F)+DEConvBlock 3×3 (CA(F)))) (9).
[0060] Among them, ConvBlock 3×3 (·) and DEConvBlock 3×3 (·) are respectively a 3×3 ordinary convolution block and a 3×3 differential convolution block, both of which contain a normalization layer and an activation function.
[0061] The reconstruction of each stage of the reconstruction network is stacked by multiple FE modules to guide the reconstruction;
[0062]
[0063] The multi-scale feature reconstruction module is optimized using the MSE loss, and the loss function L is defined as follows:
[0064]
[0065] Among them, I i and R i are the features of each stage of the encoder and the reconstruction network respectively.
[0066] The evaluation metrics are set as follows:
[0067] 1. AUROC
[0068] First, image-level AUROC and pixel-level AUROC are used. AUROC (Area Under Receiver Operating Characteristic Curve) is a widely used metric, mainly used to evaluate the performance of binary classification models. It quantifies the classification ability of the model by calculating the area under the Receiver Operating Characteristic (ROC) curve.
[0069] The horizontal axis of the ROC curve is the "False Positive Rate" (FPR), defined as the ratio of the number of negative class samples misclassified as positive class to the total number of negative class samples; the vertical axis is the "True Positive Rate" (TPR), also known as the recall rate, defined as the ratio of the number of positive class samples correctly classified as positive class to the total number of positive class samples.
[0070]
[0071] 2. AP
[0072] AP is the area under the PR curve, P is precision, and R is recall. Image-level and pixel-level AP are also used to evaluate the model of the present invention.
[0073]
[0074] 3. F1_max
[0075] In an imbalanced dataset, the F1-score is very useful. F1 is the harmonic mean of P and R, and is calculated as follows:
[0076]
[0077] In multi-class classification or when making predictions with different thresholds, multiple F1-scores can be calculated. Then, F1_max is calculated as follows:
[0078] F1_max = max(F11, F12,..., F1 n ) (15)
[0079] 4. AU-PRO
[0080] Compared with AUROC, AU-PRO usually provides a more meaningful evaluation when dealing with imbalanced positive and negative samples, especially when a trade-off between accuracy and recall is required.
[0081]
[0082] Here P ( R ) represents the corresponding P given the condition of R.
[0083] Experimental environment:
[0084] The experimental environment is the basic condition for conducting experiments. The experimental environment of this embodiment is described in detail as follows:
[0085] Table 1 Experimental environment
[0086] Experimental environment Specification Processor Inter(R) Core(TM) i9-14900K Graphics card NVIDIA GeForce RTX 4090 24G Memory 128G DDR4 Development language Python 3.10.15 Development system Ubuntu 24.04.1 Development framework Pytorch 2.1.2 + cuda11.8
[0087] Parameter settings:
[0088] Table 2 Experimental related settings
[0089] Pre-trained encoder Resnet34 Number of FEs in each stage of the decoder [3,4,6,3] Striped kernel size s in each stage of the decoder [5,7,11,21] Feature dimension in each stage of reconstruction [64,128,256] Number of SCAs N in each GA N=2 Number of training batches 16 Learning rate (decay rate) <![CDATA[0.001(1×10 -4 )]]> Number of training epochs 500 Optimizer AdamW
[0090] Experimental results:
[0091] In this embodiment, experiments were conducted on two datasets, MVTec-AD and VisA. By comparing with some state-of-the-art methods, as shown in Table 3, where mAD is the mean of each indicator. The model in this embodiment has achieved significant improvements in all indicators. These methods include RD4AD, UniAD, SimpleNet, DeSTSeg, and DiAD. These methods include traditional convolutional neural networks, vision transformers, diffusion models, etc. The model in this embodiment not only has a simple structure but also performs excellently. Thanks to the effectiveness of the strip kernel convolution in capturing long-range dependencies and its own lightweight, including the ingenious use of ordinary square convolution in this embodiment and the enhancement of the reconstruction quality by the detail enhancement module.
[0092] In addition, the number of parameters and the amount of computation of the model in this embodiment were compared in detail, as shown in Table 4. Although this embodiment has not made outstanding progress in terms of model parameters, it can control the amount of computation at a low level, and this embodiment has significantly improved the effect of defect detection at the same time.
[0093] As Figure 3 shown, this embodiment compared the defect localization effect of the model (OURS) in this embodiment on two datasets. Other methods all have false detection problems. The method in this embodiment is very sensitive to defects and locates well in complex structures such as cables and circuit boards 4 ( Figure 3 the dataset given at the bottom), which reflects the effectiveness and high quality of the reconstruction network in this embodiment.
[0094] Table 3 Experimental comparison results of seven evaluation indicators
[0095]
[0096] Table 4 Comparison of parameters, computational amounts with other methods and performance comparison on MVTec-AD
[0097]
[0098] MVTec-AD contains 5354 images from 15 categories, covering 10 object categories and 5 texture categories. VisA contains 12 industrial metal products, with 9621 normal images and 1200 abnormal images. These products have different lighting conditions, uneven backgrounds, and multiple products in each image. Some category texture structures are relatively complex and challenging. The specific information of MVTec-AD and VisA is shown in Table 5 and Table 6.
[0099] Table 5 Number of normal and abnormal cases in the training set and test set of MVTec-AD
[0100]
[0101] Table 6 Number of normal and abnormal cases in the training set and test set of VisA
[0102]
[0103] Although the present invention has been described in detail with general descriptions and specific embodiments above, modifications or improvements can be made to it on the basis of the present invention, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection required by the present invention.
Claims
1. An intelligent detection method based on global and local attention and detail enhancement, characterized by: The multi-stage features are extracted by inputting the sample through a pre-trained encoder with frozen parameters, and then the multi-scale semantic information is fused using a bottleneck layer, which is then fed back to the multi-stage reconstruction network. Each stage of the reconstruction network consists of multiple FE modules connected in series, one of which is a reconstruction component, which consists of a global and local attention module SGL and a detail enhancement module DEH. The global and local attention module SGL receives the output of the previous FE module or the previous reconstruction stage as input, extracts global and local semantic information, and the output result is enhanced in detail by the detail enhancement module DEH to improve the reconstruction quality. The global and local attention module SGL contains parallel global GA and local LA branches. The global GA captures global semantic information by multiple series of cross-axis attention SCA, and the local LA contains two large-kernel deep convolutions for local enhancement. The core of the detail enhancement module DEH consists of a channel attention CA and a multi-dimensional differential convolution DEConv; Channel attention CA first passes through global average pooling AP and global maximum pooling MP to obtain channel global information, learns attention weight distribution through two convolution and activation functions, and finally obtains attention score through Sigmoid and multiplies it with input features; The multidimensional differential convolution DEConv contains one ordinary convolution and four differential convolutions, namely center differential convolution CDC, angle differential convolution ADC, horizontal differential convolution VDC, and vertical differential convolution HDC. Ordinary convolution is used to obtain intensity level information, and differential convolution is used to enhance gradient level information. During the training process, the sum of the MSE of the features of the encoder and decoder at different stages is used as the loss function. In the inference stage, the sum of the cosine similarities is used to obtain the anomaly map, and the set evaluation indicators are used to demonstrate the detection effect.
2. The intelligent detection method based on global and local attention and detail enhancement according to claim 1, characterized in that: The global and local attention modules SGL adopt a multi-branch parallel structure, in which the cross-axis attention SCA is used to capture global information, thereby effectively modeling global normal semantic information; in this structure, the query Q, key K and value V in the attention mechanism are calculated by strip kernel depth 1D convolution; let E i For the features input to the i-th SCA module, the query Q, key K, and value V along the x-axis and y-axis are calculated as follows: in, Used as Q, K, V of the x-axis and y-axis respectively, that is, the Q, K, V values of one axis are the same, Conv1 1×1 (·) is a normal 1×1 convolution, and Along x The size of the strip kernel in the x-axis and y-axis dimensions is s 1D depth convolution, s is a hyperparameter, LN(·) represents layer normalization, and the single cross-axis attention SCA is calculated as follows: All Conv in formula (1)-(4) 1×1 (·) shared weight parameters, is the result of the cross attention calculation, τ is a learnable scale parameter used to control the size of the Q, K matrix multiplication before applying the Softmax function; ConvBlock 1×1 It is a 1×1 convolution block, which includes a 1×1 convolution layer, a normalization layer and an activation function; Assume that the input feature of the k-th stage FE module is Global Attention Module GA k The calculation is as follows: Among them, N represents the number of SCA stacked in a FE module. The local attention branch of the global and local attention module SGL contains two sets of deep convolutions with kernel sizes of 5×5 and 7×7 respectively to focus on local details. The SGL of the k-th stage FE module is calculated as follows: Among them, DWConvBlock 5×5 (·), DWConvBlock 7×7 (·) are 5×5 and 7×7 depthwise convolution blocks, respectively, both of which contain normalization layers and activation functions.
3. The intelligent detection method based on global and local attention and detail enhancement according to claim 1, characterized in that: The five parallel convolutions of multi-dimensional differential convolution DEConv will inevitably lead to an increase in parameters and inference time. By utilizing the additivity of convolution, the parallel deployed convolutions are simplified to a single standard convolution. The calculation of multi-dimensional differential convolution DEConv is as follows: Where F is the input feature, K i=1:5 Represents five types of convolution structures, * is the convolution operation, K tran Represents the transformation kernel composed of parallel convolutions; the overall calculation of the detail enhancement module DEH is as follows: DEH(F)=Conv 1×1 (CA(F)+ConvBlock 3x3 (CA(F)+DEConvBlock 3×3 (CA(F)))) (9); Among them, ConvBlock 3×3 (·),DEConvBlock 3×3 (·) are 3×3 normal convolution blocks and 3×3 differential convolution blocks, respectively, both of which contain normalization layers and activation functions.
4. The intelligent detection method based on global and local attention and detail enhancement according to claim 1, characterized in that: The reconstruction of each stage of the reconstruction network is composed of a stack of multiple FE modules to guide the reconstruction; 5. The intelligent detection method based on global and local attention and detail enhancement according to claim 1, characterized in that: MSE loss is used to optimize the multi-scale feature reconstruction module, and the loss function L is defined as follows: Among them I i and R i They are the characteristics of each stage of the encoder and reconstruction network respectively.
Citation Information
Patent Citations
Double-resolution real-time semantic segmentation method based on detail enhancement
CN117409412A
Medical image automatic segmentation method of U-shaped network based on fusion convolution and attention mechanism
CN117474866A