Infrared small target detection system, training method and detection method
By designing an infrared small target detection system that hybridizes the self-attention module and the detail enhancement module, the problem of low accuracy of infrared small target detection in complex backgrounds is solved, and accurate perception and efficient detection of sparse targets are achieved.
Patent Information
- Application Number
- CN202510835174.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Existing infrared small target detection methods have difficulty in accurately distinguishing infrared small targets in complex backgrounds, especially in complex background noise where the signal-to-noise ratio is low. In addition, existing deep learning methods lack the ability to model global context, resulting in low detection accuracy.
A hybrid self-attention module and a detail enhancement module are adopted to enhance the attention to long-distance features through the expansion self-attention mechanism. Combining multi-scale feature extraction and differential thinking, an infrared small target detection system is designed, including an encoding module, a block embedding processing module, a hybrid self-attention module, a feature extraction module, a decoding module and a detection module. The receptive field design with a Gaussian shape in the self-attention mechanism is used to enhance the perception of sparse targets, and the detail enhancement module is used to alleviate the problem of small target detail loss caused by encoding feature compression.
It achieves accurate detection of small infrared targets in complex backgrounds, enhances the perception of sparse targets, improves the characterization ability of small targets, and improves detection accuracy and robustness.
Smart Images

Figure CN120747463A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of infrared target detection, and more specifically, relates to an infrared small target detection system, a training method and a detection method. Background Art
[0002] Infrared thermal imaging can provide clear target information, enhance detection accuracy and response time in complex backgrounds, and has low power consumption, high reliability, and stable working capabilities under all lighting conditions. It can play a vital role in anti-UAV monitoring, border patrol, field search and rescue, critical facility security and other fields. However, infrared small target detection still faces many challenges. First, due to the long distance between the infrared sensor and the infrared small target, the captured target image size is very small. In addition, the infrared radiation attenuates during long-distance transmission, making the target appear blurry and lacking clear texture and shape. Second, small infrared targets are usually present in complex background noise, which reduces the signal-to-noise ratio and increases the difficulty of target detection. Faced with these challenges, improving the accuracy of infrared small target detection in complex backgrounds has become an urgent need for technological development.
[0003] To address these issues, existing infrared small target detection methods can be broadly categorized into two types: model-driven and data-driven. Model-driven methods typically rely on handcrafted features and prior knowledge, which makes them less adaptable in complex environments and susceptible to noise and illumination variations. Common model-driven methods include filtering-based techniques, local contrast-based methods, and low-rank and sparse matrix recovery methods. While these methods perform well in some simple backgrounds, they often perform poorly in complex settings, suffer from high false alarm rates, and struggle to effectively handle small targets and those with varying shapes.
[0004] In response to the limitations of model-driven approaches, data-driven methods have gradually become a research focus in recent years. The application of deep learning models, in particular, has provided a new solution for infrared small target detection. These methods automatically learn and extract features from large amounts of data through end-to-end training, effectively overcoming the shortcomings of manual feature design. However, despite the progress made by deep learning methods, challenges in the practical application of infrared small target detection remain. Existing data-driven infrared small target detection methods often use convolutional neural networks (CNNs) for feature extraction and target detection. However, these methods lack the ability to model global context, making it difficult to accurately distinguish targets from complex backgrounds. Consequently, the accuracy of infrared small target detection in complex backgrounds is low. Summary of the Invention
[0005] In response to the above defects or improvement needs of the existing technology, the present invention provides an infrared small target detection system, training method and detection method to solve the technical problem that the existing technology cannot accurately detect infrared small targets in complex backgrounds.
[0006] In order to achieve the above objectives, in a first aspect, the present invention provides an infrared small target detection system, comprising:
[0007] The encoding module includes M cascaded encoding units; the encoding module is used to extract M visual feature maps of different scales of the input infrared small target image; M ≥ 2;
[0008] M block embedding processing modules are used to extract the embedded feature maps of the visual feature maps of M scales in a one-to-one correspondence; the embedding processing module is used to divide the input visual feature map into multiple local image blocks, linearly project them into the embedding space, and splice the obtained results in the channel dimension to obtain the corresponding embedded feature map;
[0009] A hybrid self-attention module, comprising: N cascaded hybrid self-attention units; N ≥ 1; each level of hybrid self-attention unit is used to receive the input of M feature maps, and perform 1×1 convolution operation, depth convolution operation and reshaping operation on each feature map in sequence to obtain the corresponding intermediate feature map; each intermediate feature map is sampled at the same sampling rate in the channel dimension and spliced to obtain the corresponding sampled feature map; all sampled feature maps are spliced in the channel dimension to obtain the total sampled feature map; each intermediate feature map is used as the Q matrix, and the total sampled feature map is used as the K matrix and the V matrix to perform attention mechanism calculation to obtain the corresponding spatial channel attention feature map; the M feature maps input to the first level hybrid self-attention unit are the M embedded feature maps obtained by the M block embedding processing modules; when N ≥ 2, the M feature maps input to the second to Nth spatial channel self-attention units are the M spatial channel attention feature maps obtained in the previous level hybrid self-attention unit;
[0010] A feature extraction module is used to extract features from the visual feature map output by the M-th level encoding unit to obtain a depth feature map;
[0011] The decoding module includes M cascaded decoding units. The input of the first-level decoding unit is connected to the output of the M-th-level encoding unit through the feature extraction module. The i-th-level decoding unit is also connected to the M-i+1-th-level encoding unit and the M-i+1-th-level hybrid self-attention unit, respectively. i = 1, 2, ..., M. The decoding module is used to decode the deep feature map layer by layer to obtain a decoded feature map.
[0012] The detection module is used to obtain the target segmentation map based on the decoded feature map.
[0013] Further preferably, the i-th level decoding unit is used to receive the visual feature map input by the corresponding decoding unit and the spatial channel attention feature map input by the M-i+1-th level mixed self-attention unit, and fuse the two to obtain a fused feature map; based on the fused feature map, a decoding operation is performed on the feature map input by the previous level to obtain the i-th level decoding feature map;
[0014] When i=1, the feature map of the previous level input is a depth feature map;
[0015] When i=2,3,…,M, the feature map of the previous level input is the i-1th level decoding feature map;
[0016] The output of the above decoding module is the M-th level decoding feature map.
[0017] Further preferably, the infrared small target detection system further comprises: M detail enhancement modules, configured to perform feature enhancement on the M visual feature maps of different scales in a one-to-one correspondence to obtain M visual enhancement feature maps of different scales;
[0018] The detail enhancement module is used to perform a 1×1 convolution operation on the input visual feature map, and then use G cascaded maximum pooling layers to perform downsampling operations to obtain G downsampled feature maps of different scales, which are input one-to-one into G large-core detail enhancement units to obtain G downsampled enhanced feature maps; the G downsampled enhanced feature maps are fused and then subjected to a 1×1 convolution operation to obtain the corresponding visual enhancement feature map; G ≥ 1;
[0019] The large core detail enhancement unit is used to perform an average pooling operation on the input downsampled feature map to obtain a pooled feature map; the input downsampled feature map is subtracted from the pooled feature map, and then a 1×1 convolution operation is performed, and the result is added to the input downsampled feature map to obtain a downsampled enhanced feature map;
[0020] The convolution kernel scale of the average pooling operation in the large-core detail enhancement unit corresponding to the downsampled feature map of the g-th scale is larger than the convolution kernel scale of the maximum pooling operation in the large-core detail enhancement unit corresponding to the downsampled feature map of the g+1-th scale, and is larger than the preset scale; g = 1, 2, ..., G-1;
[0021] The M-i+1th level encoding unit is connected to the i-th level decoding unit through the corresponding detail enhancement module;
[0022] The i-th level decoding unit is used to receive the visual enhancement feature map input by the corresponding detail enhancement module and the spatial channel attention feature map input by the M-i+1-th level mixed self-attention unit, and fuse the two to obtain a fused feature map; based on the fused feature map, the feature map input by the previous level is decoded to obtain the i-th level decoding feature map;
[0023] When i=1, the feature map of the previous level input is a depth feature map;
[0024] When i=2,3,…,M, the feature map of the previous level input is the i-1th level decoding feature map;
[0025] The output of the above decoding module is the M-th level decoding feature map.
[0026] Further preferably, the detection module is used to perform a 1×1 convolution operation, a Sigmoid operation and a binarization operation on the decoded feature map in sequence to obtain a target segmentation map.
[0027] In a second aspect, the present invention provides a training method for an infrared small target detection system, comprising:
[0028] Input each infrared small target image sample in the training set into the infrared small target detection system to obtain the corresponding target segmentation map;
[0029] Construct an overall training objective including the first training objective; the first training objective is to minimize the difference loss between the target segmentation map and the corresponding true mask map of each infrared small target image sample in the training set;
[0030] Based on the overall training objectives, the infrared small target detection system is trained;
[0031] Among them, the infrared small target detection system is the infrared small target detection system provided by the first aspect of the present invention.
[0032] Further preferably, the expression of the first training objective is:
[0033]
[0034] Among them, A p A is the pixel set of the target area in the target segmentation map of the infrared small target image sample; gt is the pixel set of the target area in the real mask image of the infrared small target image sample;
[0035] Further preferably, the above-mentioned total training goal further includes: a j-th training goal; the j-th training goal is: minimizing the difference loss between the j-th supplementary target segmentation map and the corresponding true mask map of each infrared small target image sample in the training set; j=2, 3, ..., M+2;
[0036] When j = 2, 3, ..., M, the method for obtaining the j-th supplementary target segmentation map includes:
[0037] Obtaining an M-j+1th level decoding feature map obtained in an M-j+1th level decoding unit of the infrared small target detection system when the infrared small target image sample is input into the infrared small target detection system;
[0038] The detection module in the infrared small target detection system is used to process the M-j+1th level decoded feature map, upsample it to the same scale as the corresponding real mask map, and obtain the jth supplementary target segmentation map;
[0039] When j=M+1, the method for obtaining the j-th supplementary target segmentation map includes:
[0040] Obtaining a depth feature map obtained in a feature extraction module of the infrared small target detection system when the infrared small target image sample is input into the infrared small target detection system;
[0041] The detection module in the infrared small target detection system is used to process the depth feature map, and then upsample it to the same scale as the corresponding real mask map to obtain the jth supplementary target segmentation map;
[0042] When j=M+2, the j-th supplementary target segmentation map is obtained by: obtaining a depth feature map obtained in a feature extraction module of the infrared small target detection system when the infrared small target image sample is input into the infrared small target detection system, and M decoding feature maps of different scales obtained in M cascaded decoding units;
[0043] The detection module in the infrared small target detection system is used to process the depth feature map and M decoding feature maps of different scales respectively to obtain M+1 supplementary target segmentation maps of different scales;
[0044] After uniformly transforming the scales of M+1 supplementary target segmentation maps of different scales to the same scale as the corresponding true mask map, they are spliced in the channel dimension to obtain a spliced map;
[0045] The detection module is used to process the spliced image to obtain the j-th supplementary target segmentation map.
[0046] Further preferably, the expression of the j-th training objective is:
[0047]
[0048] Among them, A pj A is the pixel set of the target area in the jth supplementary target segmentation map of the infrared small target image sample; gt is the pixel set of the target area in the real mask image of the infrared small target image sample;
[0049]
[0050] Further preferably, the above-mentioned overall training objective also includes: a position training objective; the position training objective is: minimizing the difference loss between the target center point in the target segmentation map of each infrared small target image sample in the training set and the target center point in the corresponding true mask map.
[0051] In a third aspect, the present invention provides a method for detecting small infrared targets, comprising:
[0052] The infrared small target to be detected is input into the infrared small target detection system provided by the first aspect of the present invention to obtain a corresponding target segmentation map, thereby realizing infrared small target detection.
[0053] In a fourth aspect, the present invention provides an electronic device comprising: a memory and a processor, wherein the memory stores a computer program, and the processor executes the method provided in the second aspect or the third aspect of the present invention when executing the computer program.
[0054] In a fifth aspect, the present invention further provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is executed by a processor, the device where the storage medium is located is controlled to execute the method provided in the second aspect or the third aspect of the present invention.
[0055] In a sixth aspect, the invention further provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the method provided in the second or third aspect of the invention.
[0056] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects:
[0057] 1. The present invention provides an infrared small target detection system. Considering that the self-attention mechanism presents a Gaussian-like shape and the receptive field is only sensitive to the central area, in order to enable the receptive field to focus on long-distance features, an expanded self-attention mechanism is designed in the hybrid self-attention module. Each intermediate feature map is sampled at the same sampling rate in the channel dimension and spliced to obtain the corresponding sampling feature map; all the sampling feature maps are spliced in the channel dimension to obtain the total sampling feature map; each intermediate feature map is used as the Q matrix, and the total sampling feature map is used as the K matrix and the V matrix, and the attention mechanism calculation is performed to obtain the corresponding spatial channel attention feature map, which can pay more attention to long-distance features, realize long-range dependency modeling and effective context aggregation, enhance the perception ability of sparse targets, effectively capture global context information, and accurately detect infrared small targets in complex backgrounds.
[0058] 2. Furthermore, in the infrared small target detection system provided by the present invention, the detail enhancement module first extracts multi-scale features by stacking multiple maximum pooling layers, removing redundant background information to capture details at different levels. Then, using a difference algorithm, the module calculates the difference between the downsampled feature map and the average pooled image to highlight the features of the small target area. This alleviates the problem of small target detail loss caused by the encoding feature compression process, further improving the small target representation capability.
[0059] 3. The present invention provides a training method for an infrared small target detection system, which takes minimizing the difference loss between the target segmentation map and the corresponding true mask map of each infrared small target image sample in the training set as the first training goal, and trains the infrared small target detection system so that the infrared small target detection system can accurately detect infrared small targets in complex backgrounds.
[0060] 4. Furthermore, in the training method for the infrared small target detection system provided by the present invention, the expression of the first training target is: Among them, A p A is the pixel set of the target area in the target segmentation map of the infrared small target image sample; gt is the pixel set of the target area in the real mask image of the infrared small target image sample; w2 is designed based on relative difference. p and A gt The greater the difference, the smaller w2 is, and vice versa. The larger it is, the more accurately it can reflect the difference loss between the target segmentation map of the infrared small target image sample and the corresponding true mask map, further improving the accuracy of the trained infrared small target detection system.
[0061] 5. Furthermore, in the training method of the infrared small target detection system provided by the present invention, the overall training goal also includes: a second training goal; the second training goal is to minimize the difference loss between the supplementary target segmentation map and the corresponding real mask map of each infrared small target image sample in the training set; wherein, the supplementary target segmentation map is based on the depth feature map obtained in the feature extraction module of the infrared small target detection system, and M decoding feature maps of different scales obtained in M cascaded decoding units. By fusing information of different scales, the sensitivity of the infrared small target detection system when the target scale changes is enhanced, and the accuracy of the infrared small target detection system for detecting infrared small targets in complex backgrounds is further enhanced.
[0062] 6. Furthermore, in the training method of the infrared small target detection system provided by the present invention, the expression of the second training objective is: Among them, A p' is the pixel set of the target area in the supplementary target segmentation map of the infrared small target image sample; A gt is the pixel set of the target area in the real mask image of the infrared small target image sample; w2' is designed based on relative difference. p ' and A gt The greater the difference, the smaller w2' is, and vice versa. The larger it is, the more accurately it can reflect the difference loss between the target segmentation map of the supplemented infrared small target image sample and the corresponding real mask map, further improving the accuracy of the trained infrared small target detection system. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 A schematic diagram of the structure of an infrared small target detection system provided by an embodiment of the present invention;
[0064] Figure 2 A schematic diagram of the structure of a hybrid attention module provided by an embodiment of the present invention;
[0065] Figure 3 A schematic diagram of the structure of a detail enhancement module provided in an embodiment of the present invention;
[0066] Figure 4 A schematic diagram of the training process of the infrared small target detection system provided by an embodiment of the present invention;
[0067] Figure 5 Schematic diagram of the comparison results of HSTNet in the present invention and other existing algorithms on the NUAA-SIRST, IRSTD-1K and IRSTD-Large datasets provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0068] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0069] In order to achieve the above objectives, in a first aspect, the present invention provides an infrared small target detection system, comprising:
[0070] The encoding module includes M cascaded encoding units; the encoding module is used to extract M visual feature maps of different scales of the input infrared small target image; M ≥ 2;
[0071] M block embedding processing modules are used to extract the embedded feature maps of the visual feature maps of M scales in a one-to-one correspondence; the embedding processing module is used to divide the input visual feature map into multiple local image blocks and linearly project them into the embedding space, thereby transforming the scale to the same scale, and then splicing the obtained results in the channel dimension to obtain the corresponding embedded feature map;
[0072] A hybrid self-attention module, comprising: N cascaded hybrid self-attention units; N ≥ 1; each level of hybrid self-attention unit is used to receive the input of M feature maps, and perform 1×1 convolution operation, depth convolution operation and reshaping operation on each feature map in sequence to obtain the corresponding intermediate feature map; each intermediate feature map is sampled at the same sampling rate in the channel dimension and spliced to obtain the corresponding sampled feature map; all sampled feature maps are spliced in the channel dimension to obtain the total sampled feature map; each intermediate feature map is used as the Q matrix, and the total sampled feature map is used as the K matrix and the V matrix to perform attention mechanism calculation to obtain the corresponding spatial channel attention feature map; the M feature maps input to the first level hybrid self-attention unit are the M embedded feature maps obtained by the M block embedding processing modules; when N ≥ 2, the M feature maps input to the second to Nth spatial channel self-attention units are the M spatial channel attention feature maps obtained in the previous level hybrid self-attention unit;
[0073] A feature extraction module is used to extract features from the visual feature map output by the M-th level encoding unit to obtain a depth feature map;
[0074] The decoding module includes M cascaded decoding units. The input of the first-level decoding unit is connected to the output of the M-th-level encoding unit through the feature extraction module. The i-th-level decoding unit is also connected to the M-i+1-th-level encoding unit and the M-i+1-th-level hybrid self-attention unit, respectively. i = 1, 2, ..., M. The decoding module is used to decode the deep feature map layer by layer to obtain a decoded feature map.
[0075] The detection module is used to obtain the target segmentation map based on the decoded feature map.
[0076] In an optional embodiment, the i-th level decoding unit is used to receive the visual feature map input by the corresponding decoding unit and the spatial channel attention feature map input by the M-i+1-th level mixed self-attention unit, and fuse the two to obtain a fused feature map; based on the fused feature map, a decoding operation is performed on the feature map input at the previous level to obtain an i-th level decoding feature map;
[0077] When i=1, the feature map of the previous level input is a depth feature map;
[0078] When i=2,3,…,M, the feature map of the previous level input is the i-1th level decoding feature map;
[0079] The output of the above decoding module is the M-th level decoding feature map.
[0080] It should be noted that there are many ways to fuse the visual feature map and the spatial channel attention feature map, such as pixel-by-pixel addition, channel splicing, element-by-element multiplication, etc., which are not limited here; preferably, in an optional implementation method, the fusion of the visual feature map and the spatial channel attention feature map is achieved by adding them pixel by pixel.
[0081] In an optional embodiment, the infrared small target detection system further includes: M detail enhancement modules, configured to perform feature enhancement on the M visual feature maps of different scales in a one-to-one correspondence to obtain M visual enhancement feature maps of different scales;
[0082] The detail enhancement module is used to perform a 1×1 convolution operation on the input visual feature map, and then use G cascaded maximum pooling layers to perform downsampling operations to obtain G downsampled feature maps of different scales, which are input one-to-one into G large-core detail enhancement units to obtain G downsampled enhanced feature maps; the G downsampled enhanced feature maps are fused and then subjected to a 1×1 convolution operation to obtain the corresponding visual enhancement feature map; G ≥ 1;
[0083] The large core detail enhancement unit is used to perform an average pooling operation on the input downsampled feature map to obtain a pooled feature map; the input downsampled feature map is subtracted from the pooled feature map, and then a 1×1 convolution operation is performed, and the result is added to the input downsampled feature map to obtain a downsampled enhanced feature map;
[0084] Among them, the convolution kernel scale of the average pooling operation in the large core detail enhancement unit corresponding to the down-sampling feature map of the g-th scale is larger than the convolution kernel scale of the maximum pooling operation in the large core detail enhancement unit corresponding to the down-sampling feature map of the g+1-th scale, and is larger than the preset scale; g=1,2,…,G-1; preferably, in an optional implementation, the preset scale is 5×5.
[0085] The M-i+1th level encoding unit is connected to the i-th level decoding unit through the corresponding detail enhancement module;
[0086] The i-th level decoding unit is used to receive the visual enhancement feature map input by the corresponding detail enhancement module and the spatial channel attention feature map input by the M-i+1-th level mixed self-attention unit, and fuse the two to obtain a fused feature map; based on the fused feature map, the feature map input by the previous level is decoded to obtain the i-th level decoding feature map;
[0087] When i=1, the feature map of the previous level input is a depth feature map;
[0088] When i=2,3,…,M, the feature map of the previous level input is the i-1th level decoding feature map;
[0089] The output of the above decoding module is the M-th level decoding feature map.
[0090] It should be noted that the encoding unit can be a residual module, a convolutional block, etc., without limitation here. The decoding unit can be an NCC model, an upsampling module, etc., without limitation here. The feature extraction module can be a residual module, an attention mechanism module, etc., without limitation here.
[0091] It should be noted that there are many ways to fuse the visual feature map and the spatial channel attention feature map, such as pixel-by-pixel addition, splicing in the channel dimension, element-by-element multiplication, etc., which are not limited here; preferably, in an optional implementation method, the fusion of the visual feature map and the spatial channel attention feature map is achieved by adding them pixel-by-pixel.
[0092] It should be noted that the fusion of the G downsampled enhanced feature maps, such as splicing in the channel dimension, pixel-by-pixel addition, element-by-element multiplication, etc., is not limited here; preferably, in an optional implementation, the fusion is achieved by splicing the G downsampled enhanced feature maps in the channel dimension.
[0093] It should be noted that there are many ways for the above-mentioned detection module to obtain the target segmentation map based on the decoded feature map, such as classifying whether each pixel point in the decoded feature map belongs to the target area, performing pixel-level regression on the decoded feature map, threshold-based binarization processing, and combining conditional random fields to obtain the target segmentation map, etc., which are not limited here.
[0094] Preferably, in an optional implementation, the above-mentioned detection module is used to perform a 1×1 convolution operation, a Sigmoid operation and a binarization operation on the decoded feature map in sequence to obtain a target segmentation map, with less computational effort.
[0095] In order to further illustrate the infrared small target detection system provided by the present invention, a specific embodiment 1 is described in detail below:
[0096] This embodiment provides an infrared small target detection system based on spatial and channel sparse self-attention mechanism, such as Figure 1 As shown in FIG, the system mainly includes: an encoder module (Encoder), a block embedding processing module (PE), a hybrid self-attention module (SCST), a detail enhancement module (MSDE), a feature extraction module, and a decoder module (Decoder). The encoder module in this embodiment includes four cascaded encoding units; the encoding units and feature extraction modules in this embodiment are both convolutional residual blocks. The decoding module includes four cascaded decoding units. There are also four detail enhancement modules.
[0097] The encoder module is used to extract the multi-level features of the input infrared small target image. The four-level convolution residual block can obtain the four-level features. The visual feature map output by the i-th convolution residual block is denoted as E i ; In the encoding module, as the layer becomes deeper, the feature scale becomes smaller and smaller.
[0098] Feature Map E i On the one hand, it is passed down to the i+1th convolutional residual block for further encoding;
[0099] On the other hand, the feature map E i After being processed by the detail enhancement module and the hybrid self-attention module respectively, the progressive interaction between the encoder and decoder at the same layer is realized: First, the feature map E i Processed by the detail enhancement module to obtain the visual enhancement feature map At the same time, the feature map E i The embedded feature map P is obtained by processing the block embedding processing module PE (Patch Embedding) of the corresponding layer i , and then output to the hybrid self-attention module. The hybrid self-attention module embeds feature maps P of four different layers i Perform fusion processing to obtain the features of fusion context information of different scales, recorded as O i At the decoding end, the i-th layer decoding unit first adds the interactive features of each layer encoder through the adder and O i The fused features are then combined with the output of the i+1th layer decoding unit for processing, and then passed to the i-1th layer decoding unit. Finally, the first layer decoding unit inputs the processed results into the Conv1×1 convolution layer for processing to obtain the target saliency map, which is then processed through the Sigmoid function and binarization to obtain the target segmentation map. In the decoding module, as the layers become deeper, the feature scale becomes larger.
[0100] In this embodiment, the encoder module uses a four-level convolutional residual block stack to extract image features. Each convolutional residual block has the same structure, consisting of a convolution, a batch normalization layer (BatchNormalization) layer, an activation layer (LeakyReLU function) and a residual connection.
[0101] like Figure 2 As shown in the example, the hybrid attention module improves the target detection performance of the system by adopting an improved expansion self-attention mechanism. The hybrid attention module first uses a 1×1 convolution module to process the embedded feature map P after the block embedding processing module. i(i=1,2,3,4), perform cross-channel pixel-level context information fusion; then use 3×3 depth convolution to extract local spatial features to obtain feature map I i Finally, the reshaped feature Q is obtained by the reshaping module R (the dimension C×H×W is transformed into C×(H×W), where C, H, and W represent the number of channels, height, and width respectively) i , and for feature Q i Perform expansion self-attention calculation to obtain feature map O i The above processing can be expressed by the following formula:
[0102] I i =φ dw [φ 1×1 (P i )]
[0103] Q i =Reshape(I i )
[0104] K=Cat[(Q1,Q2,Q3,Q4)]
[0105] in,
[0106] In traditional self-attention calculation, when calculating the query feature Q i and key feature K j When calculating the similarity between two pairs, each pair of i and j is involved in the calculation, which increases the amount of calculation:
[0107]
[0108] The present invention improves the above operation and implements the expansion self-attention mechanism. Specifically, for the query matrix After the dilation operation is introduced, it can be expanded into a sparse matrix
[0109]
[0110] Similarly, for the bond matrix Sum Matrix After the expansion operation, they can be expressed as:
[0111]
[0112] The calculation process of the inflated self-attention weight can be expressed as follows:
[0113]
[0114] in, σ and They are represented by Softmax activation function and instance normalization, C i (i=1,2,3,4) is the sequence length,
[0115] like Figure 3 As shown, this embodiment provides a detail enhancement module that first extracts multi-scale features by stacking multiple maximum pooling layers to capture detail information at different levels. Then, using the difference concept, the feature of small target areas is highlighted by calculating the difference between the original image and the image after average pooling. The specific process is as follows:
[0116] Output feature E for the i-th layer encoding module i , first processed through a 1×1 convolution layer to obtain the feature map Then it is processed by three cascaded 3×3 maximum pooling layers. The maximum pooling layer is used to extract detailed features in the image and remove redundant background information to obtain features after processing by different maximum pooling layers.
[0117]
[0118] where φ 1×1 and MP represent 1×1 convolutional layer and maximum pooling layer, respectively.
[0119] In addition, to further refine the features of small targets in the image, this paper proposes a large kernel detail enhancement unit (LKDE). This module can more effectively capture the details of the target by expanding the size of the convolution kernel. The processing process can be expressed as follows:
[0120]
[0121] in, Represented as a large core detail enhancement unit, For detail enhancement features, the structure of the large core detail enhancement unit is as follows Figure 3 As shown on the right. This unit is based on the difference idea, by calculating the difference between the original image and the image after average pooling, to highlight the details of small targets. Specifically, first the input feature map Perform large kernel average pooling to obtain low-frequency background feature maps. Then, the feature maps Perform differential operation with the background feature map to generate a differential feature map Then, the information is integrated in the channel dimension through the 1×1 convolution layer operation, and finally combined with the feature map Add element by element to get the feature map after detail enhancement The above process can be expressed by the following calculation formula:
[0122]
[0123] After the above processing, and The feature maps are concatenated along the channel dimension and fed into a 1×1 convolutional layer for processing:
[0124]
[0125] In this embodiment, the decoding unit includes: a cascaded convolution layer, a batch normalization layer (BatchNormalization) layer, an activation layer (LeakyReLU function) and a nearest neighbor interpolation upsampling.
[0126] In summary, this embodiment designs a hybrid self-attention module (SCST) and detail enhancement module (MSDE). The SCST module uses an expanded self-attention mechanism to enhance the perception of sparse targets, effectively capturing global contextual information and improving the detection of small targets. The MSDE module uses a multi-scale feature extraction method to alleviate the problem of small target detail loss caused by encoding feature compression, thereby improving the representation of small targets.
[0127] It should be noted that the training method for the above-mentioned infrared small target detection system can adopt a conventional training method, such as an end-to-end training method. Based on this, preferably, in a second aspect, the present invention provides a training method for an infrared small target detection system, comprising:
[0128] Input each infrared small target image sample in the training set into the infrared small target detection system to obtain the corresponding target segmentation map;
[0129] Construct an overall training objective including the first training objective; the first training objective is to minimize the difference loss between the target segmentation map and the corresponding true mask map of each infrared small target image sample in the training set;
[0130] Based on the overall training objectives, the infrared small target detection system is trained;
[0131] Among them, the infrared small target detection system is the infrared small target detection system provided by the first aspect of the present invention.
[0132] In an optional implementation manner, the expression of the first training objective is:
[0133]
[0134] Among them, A p A is the pixel set of the target area in the target segmentation map of the infrared small target image sample; gtis the pixel set of the target area in the real mask image of the infrared small target image sample;
[0135] In an optional embodiment, the above-mentioned total training goal further includes: a j-th training goal; the j-th training goal is: minimizing the difference loss between the j-th supplementary target segmentation map and the corresponding true mask map of each infrared small target image sample in the training set; j = 2, 3, ..., M + 2;
[0136] When j = 2, 3, ..., M, the method for obtaining the j-th supplementary target segmentation map includes:
[0137] Obtaining an M-j+1th level decoding feature map obtained in an M-j+1th level decoding unit of the infrared small target detection system when the infrared small target image sample is input into the infrared small target detection system;
[0138] The detection module in the infrared small target detection system is used to process the M-j+1th level decoded feature map, upsample it to the same scale as the corresponding real mask map, and obtain the jth supplementary target segmentation map;
[0139] When j=M+1, the method for obtaining the j-th supplementary target segmentation map includes:
[0140] Obtaining a depth feature map obtained in a feature extraction module of the infrared small target detection system when the infrared small target image sample is input into the infrared small target detection system;
[0141] The detection module in the infrared small target detection system is used to process the depth feature map, and then upsample it to the same scale as the corresponding real mask map to obtain the jth supplementary target segmentation map;
[0142] When j=M+2, the j-th supplementary target segmentation map is obtained by: obtaining a depth feature map obtained in a feature extraction module of the infrared small target detection system when the infrared small target image sample is input into the infrared small target detection system, and M decoding feature maps of different scales obtained in M cascaded decoding units;
[0143] The detection module in the infrared small target detection system is used to process the depth feature map and M decoding feature maps of different scales respectively to obtain M+1 supplementary target segmentation maps of different scales;
[0144] After uniformly transforming the scales of M+1 supplementary target segmentation maps of different scales to the same scale as the corresponding true mask map, they are spliced in the channel dimension to obtain a spliced map;
[0145] The detection module is used to process the spliced image to obtain the j-th supplementary target segmentation map.
[0146] In an optional implementation, the expression of the j-th training objective is:
[0147]
[0148] Among them, A pj A is the pixel set of the target area in the jth supplementary target segmentation map of the infrared small target image sample; gt is the pixel set of the target area in the real mask image of the infrared small target image sample;
[0149]
[0150] In an optional embodiment, the above-mentioned overall training objective also includes: a position training objective; the position training objective is: minimizing the difference loss between the target center point in the target segmentation map of each infrared small target image sample in the training set and the target center point in the corresponding true mask map.
[0151] It should be noted that the above method of measuring difference loss is not limited to the above method, but can also be L2 loss, Huber loss, MAE loss, etc., which are not limited here.
[0152] The related technical solution is the same as the infrared small target detection system provided in the first aspect of the present invention and is not limited here.
[0153] To further illustrate the training method of the infrared small target detection system provided by the present invention, the infrared small target detection system provided by Example 1 is taken as an example and described in detail in conjunction with a specific Example 2 below:
[0154] like Figure 4 As shown, the training method of the infrared small target detection system provided in this embodiment includes an image preprocessing process and a model training process. The specific processing steps are as follows:
[0155] A1. Obtain an infrared small target image dataset with a real mask image, perform image enhancement processing on the infrared small target images in the infrared small target dataset, and obtain enhanced infrared small target images, thereby constructing a training set;
[0156] A2. Extract infrared small target image samples from the training set and input them into the infrared small target detection system to obtain the corresponding target segmentation map and calculate the corresponding loss function. Then, update the parameter values of the infrared small target detection system according to the gradient backpropagation algorithm.
[0157] A3. Determine whether the number of iterations is greater than a preset iteration threshold; if so, go to A2; otherwise, go to A4; in this embodiment, the preset iteration threshold is 400 times.
[0158] A4. Save the parameter value as the final parameter value in the infrared small target detection system.
[0159] In the loss function setting step, the loss function designed by the present invention includes scale-sensitive loss, position-sensitive loss and FocalLoss loss based on relative difference. The specific expression is as follows:
[0160] The scale-sensitive loss based on relative difference includes: the difference loss between the target segmentation map of the infrared small target image sample and the corresponding true mask map and the difference loss between the j-th supplementary target segmentation map of the infrared small target image sample and the corresponding true mask map; j = 2, 3, ..., M + 2;
[0161] In this embodiment, the expression of the difference loss between the target segmentation map of the infrared small target image sample and the corresponding true mask map is:
[0162]
[0163] Among them, A p A is the pixel set of the target area in the target segmentation map of the infrared small target image sample; gt is the pixel set of the target area in the real mask image of the infrared small target image sample;
[0164] When j = 2, 3, ..., M, the method for obtaining the j-th supplementary target segmentation map includes:
[0165] Obtaining an M-j+1th level decoding feature map obtained in an M-j+1th level decoding unit of the infrared small target detection system when the infrared small target image sample is input into the infrared small target detection system;
[0166] The detection module in the infrared small target detection system is used to process the M-j+1th level decoded feature map, upsample it to the same scale as the corresponding real mask map, and obtain the jth supplementary target segmentation map;
[0167] When j=M+1, the method for obtaining the j-th supplementary target segmentation map includes:
[0168] Obtaining a depth feature map obtained in a feature extraction module of the infrared small target detection system when the infrared small target image sample is input into the infrared small target detection system;
[0169] The detection module in the infrared small target detection system is used to process the depth feature map, and then upsample it to the same scale as the corresponding real mask map to obtain the jth supplementary target segmentation map;
[0170] When j=M+2, the j-th supplementary target segmentation map is obtained by: obtaining a depth feature map obtained in a feature extraction module of the infrared small target detection system when the infrared small target image sample is input into the infrared small target detection system, and M decoding feature maps of different scales obtained in M cascaded decoding units;
[0171] The detection module in the infrared small target detection system is used to process the depth feature map and M decoding feature maps of different scales respectively to obtain M+1 supplementary target segmentation maps of different scales;
[0172] After uniformly transforming the scales of M+1 supplementary target segmentation maps of different scales to the same scale as the corresponding true mask map, they are spliced in the channel dimension to obtain a spliced map;
[0173] The detection module is used to process the spliced image to obtain the j-th supplementary target segmentation map.
[0174] In this embodiment, the expression of the j-th training target is:
[0175]
[0176] Among them, A pj A is the pixel set of the target area in the jth supplementary target segmentation map of the infrared small target image sample; gt is the pixel set of the target area in the real mask image of the infrared small target image sample;
[0177]
[0178] The position-sensitive loss is the difference between the target center point in the target segmentation map of the infrared small target image sample and the target center point in the corresponding true mask map. In this embodiment, the expression is:
[0179]
[0180] Among them, d gt and θ gt Represents the target center point c in the real mask image gt Distance and angle in polar coordinate system. First, define the pixel set A of the target area in the target segmentation map of the infrared small target image sample p and the pixel set A of the target area in the real mask image of the infrared small target image sample gt By respectively p and A gt The coordinates of all pixels in the image are averaged to obtain the center points of the predicted target and the real target, which are recorded as c p =(x p ,y p) and c gt =(x gt ,y gt ). Then, the coordinates of the two center points are converted from the rectangular coordinate system to the polar coordinate system. In the polar coordinate system, the predicted center point c p For example, the corresponding distance d p and angle θ p It can be expressed as:
[0181]
[0182] Focal Loss is the loss of whether each pixel in the target segmentation map belongs to the target foreground, denoted as
[0183] In summary, the scale-position-aware joint loss function is defined as follows:
[0184]
[0185] in, α and β are set to 0.3 and 0.7 based on experimental experience.
[0186] During the model training phase, in order to enhance the model's ability to perceive targets of different scales and positions, the present invention designs a scale-position-aware joint loss function (SLJ) to improve the model's target positioning performance.
[0187] To better illustrate the performance of the proposed infrared small target detection system (denoted as HSTNet), the following analysis is performed on the NUAA-SIRST, IRSTD-1k, and IRSTD-Large datasets:
[0188] Figure 5 The experimental results of IRSTD-1k, NUAA-SIRST and IRSTD-Large datasets are given, including: IOU (%), P d (%), F a (10 -6), the HSTNet of the present invention (denoted as Ours) is compared with 14 advanced existing algorithms on the NUAA-SIRST, IRSTD-1K and IRSTD-Large datasets. The HSTNet of the present invention is compared with 14 state-of-the-art algorithms, including CNN-based models such as ACMNet, DNANet, RDIAN, ISDTUNet, MSHNet and UIUNet; CNN-Transformer hybrid methods such as TransUNet, SCTransNet, UCTransNet and MTUNet; and traditional modeling-based methods such as Top-Hat, WSLCM, TLLCM and PSTNN. The results show that the deep learning-based methods significantly outperform the traditional methods in detection performance, highlighting the advantages of deep learning in image processing, especially for small target detection in complex environments. The HSTNet of the present invention performs well in all evaluation indicators, especially in target contour preservation and pixel-level target background recognition. This shows that HSTNet has a strong detail extraction capability and can distinguish subtle differences between the target and the background while maintaining the edge of the target.
[0189] In a third aspect, the present invention provides a method for detecting small infrared targets, comprising:
[0190] The infrared small target to be detected is input into the infrared small target detection system provided by the first aspect of the present invention to obtain a corresponding target segmentation map, thereby realizing infrared small target detection.
[0191] The related technical solution is the same as the infrared small target detection system provided in the first aspect of the present invention and is not limited here.
[0192] To further illustrate the infrared small target detection method provided by the present invention, the infrared small target detection system provided by Example 1 is taken as an example and described in detail in conjunction with a specific Example 3 below:
[0193] This embodiment provides a method for detecting small infrared targets, including the following steps:
[0194] A1. Acquire the image of the small infrared target to be detected;
[0195] A2. Scale the infrared small target image to 256×256 and perform normalization operation;
[0196] A3. Process the infrared small target image using the infrared small target detection system trained in Example 1 to generate a corresponding target segmentation map;
[0197] A4: Apply a Sigmoid activation function to the target segmentation map in A3 to obtain a target probability map. Pixels in the target probability map with a value greater than a preset threshold are considered targets and are reset to 1. Pixels in the target probability map with a value less than or equal to the threshold are considered background pixels and are reset to 0. In this embodiment, the threshold is 0.5.
[0198] In a fourth aspect, the present invention provides an electronic device comprising: a memory and a processor, wherein the memory stores a computer program, and the processor executes the method provided in the second aspect or the third aspect of the present invention when executing the computer program.
[0199] The related technical solutions are the same as the methods provided in the second and third aspects of the present invention and are not limited here.
[0200] In a fifth aspect, the present invention further provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is executed by a processor, the device where the storage medium is located is controlled to execute the method provided in the second aspect or the third aspect of the present invention.
[0201] The related technical solutions are the same as the methods provided in the second and third aspects of the present invention and are not limited here.
[0202] In a sixth aspect, the invention further provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the method provided in the second or third aspect of the invention.
[0203] The related technical solutions are the same as the methods provided in the second and third aspects of the present invention and are not limited here.
[0204] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An infrared small target detection system, characterized in that: include: An encoding module, comprising M cascaded encoding units; the encoding module is used to extract M visual feature maps of different scales of the input infrared small target image; M≥2; M block embedding processing modules, for extracting embedded feature maps of the visual feature maps at M scales in a one-to-one correspondence; the embedding processing modules are used to divide the input visual feature map into multiple local image blocks, linearly project them into the embedding space, and splice the obtained results in the channel dimension to obtain corresponding embedded feature maps; A hybrid self-attention module, comprising: N cascaded hybrid self-attention units; N ≥ 1; each level of hybrid self-attention unit is used to receive the input of M feature maps, and perform 1×1 convolution operation, depth convolution operation and reshaping operation on each feature map in sequence to obtain the corresponding intermediate feature map; each intermediate feature map is sampled at the same sampling rate in the channel dimension and spliced to obtain the corresponding sampled feature map; all sampled feature maps are spliced in the channel dimension to obtain the total sampled feature map; each intermediate feature map is used as the Q matrix, and the total sampled feature map is used as the K matrix and the V matrix to perform attention mechanism calculation to obtain the corresponding spatial channel attention feature map; the M feature maps input to the first level hybrid self-attention unit are the M embedded feature maps obtained by the M block embedding processing modules; when N ≥ 2, the M feature maps input to the second to Nth spatial channel self-attention units are the M spatial channel attention feature maps obtained in the previous level hybrid self-attention unit; A feature extraction module is used to extract features from the visual feature map output by the M-th level encoding unit to obtain a depth feature map; A decoding module comprising M cascaded decoding units; the input end of the first-level decoding unit is connected to the output end of the M-th-level encoding unit through the feature extraction module; the i-th-level decoding unit is further connected to the M-i+1-th-level encoding unit and the M-i+1-th-level hybrid self-attention unit, respectively; i = 1, 2, ..., M; the decoding module is used to decode the deep feature map layer by layer to obtain a decoded feature map; A detection module is used to obtain a target segmentation map based on the decoded feature map.
2. The infrared small target detection system according to claim 1, characterized in that: The i-th level decoding unit is used to receive the visual feature map input by the corresponding decoding unit and the spatial channel attention feature map input by the M-i+1-th level mixed self-attention unit, and fuse the two to obtain a fused feature map; based on the fused feature map, a decoding operation is performed on the feature map input by the previous level to obtain an i-th level decoding feature map; When i=1, the feature map of the previous level input is a depth feature map; When i=2, 3, ..., M, the feature map of the previous level input is the i-1th level decoding feature map; The output of the decoding module is the Mth level decoding feature map.
3. The infrared small target detection system according to claim 1, characterized in that: The system further comprises: M detail enhancement modules for performing feature enhancement on the M visual feature maps of different scales in a one-to-one correspondence to obtain M visual enhancement feature maps of different scales; The detail enhancement module is used to perform a 1×1 convolution operation on the input visual feature map, and then use G cascaded maximum pooling layers to perform downsampling operations to obtain G downsampled feature maps of different scales, and input them one by one into G large-core detail enhancement units to obtain G downsampled enhanced feature maps; after fusing the G downsampled enhanced feature maps, a 1×1 convolution operation is performed to obtain the corresponding visual enhancement feature map; G ≥ 1; The large core detail enhancement unit is used to perform an average pooling operation on the input downsampled feature map to obtain a pooled feature map; subtract the input downsampled feature map from the pooled feature map, perform a 1×1 convolution operation, and add the result to the input downsampled feature map to obtain a downsampled enhanced feature map; The convolution kernel scale of the average pooling operation in the large-core detail enhancement unit corresponding to the downsampled feature map of the g-th scale is larger than the convolution kernel scale of the maximum pooling operation in the large-core detail enhancement unit corresponding to the downsampled feature map of the g+1-th scale, and is larger than the preset scale; g = 1, 2, ..., G-1; The M-i+1th level encoding unit is connected to the i-th level decoding unit through the corresponding detail enhancement module; The i-th level decoding unit is used to receive the visual enhancement feature map input by the corresponding detail enhancement module and the spatial channel attention feature map input by the M-i+1-th level mixed self-attention unit, and fuse the two to obtain a fused feature map; based on the fused feature map, a decoding operation is performed on the feature map input by the previous level to obtain an i-th level decoding feature map; When i=1, the feature map of the previous level input is a depth feature map; When i=2, 3, ..., M, the feature map of the previous level input is the i-1th level decoding feature map; The output of the decoding module is the Mth level decoding feature map.
4. The infrared small target detection system according to claim 1, characterized in that: The detection module is used to perform a 1×1 convolution operation, a Sigmoid operation, and a binarization operation on the decoded feature map in sequence to obtain a target segmentation map.
5. A training method for an infrared small target detection system, characterized in that: include: Input each infrared small target image sample in the training set into the infrared small target detection system to obtain the corresponding target segmentation map; Constructing an overall training objective including a first training objective; wherein the first training objective is to minimize the difference loss between the target segmentation map and the corresponding true mask map of each infrared small target image sample in the training set; Based on the overall training objective, training the infrared small target detection system; Wherein, the infrared small target detection system is the infrared small target detection system according to any one of claims 1 to 4.
6. The training method according to claim 5, characterized in that The expression of the first training objective is: Among them, A p A is the pixel set of the target area in the target segmentation map of the infrared small target image sample; gt is the pixel set of the target area in the real mask image of the infrared small target image sample; 7. The training method according to claim 5 or 6, characterized in that: The total training goal also includes: a j-th training goal; the j-th training goal is to minimize the difference loss between the j-th supplementary target segmentation map and the corresponding true mask map of each infrared small target image sample in the training set; j = 2, 3, ..., M + 2; When j=2, 3, ..., M, the method for obtaining the j-th supplementary target segmentation map includes: obtaining an M-j+1th level decoding feature map obtained in an M-j+1th level decoding unit of the infrared small target detection system when an infrared small target image sample is input into the infrared small target detection system; Using the detection module in the infrared small target detection system, after processing the M-j+1-th level decoding feature map, up-sample it to the same scale as the corresponding real mask map, to obtain the j-th supplementary target segmentation map; When j=M+1, the method for obtaining the j-th supplementary target segmentation map includes: Acquiring a depth feature map obtained in a feature extraction module of the infrared small target detection system when an infrared small target image sample is input into the infrared small target detection system; Using the detection module in the infrared small target detection system, after processing the depth feature map, up-sample it to the same scale as the corresponding real mask map, to obtain the j-th supplementary target segmentation map; When j=M+2, the j-th supplementary target segmentation map is obtained by: obtaining a depth feature map obtained in a feature extraction module of the infrared small target detection system when an infrared small target image sample is input into the infrared small target detection system, and M decoding feature maps of different scales obtained in M cascaded decoding units; Using the detection module in the infrared small target detection system, the depth feature map and the M decoding feature maps of different scales are processed respectively to obtain M+1 supplementary target segmentation maps of different scales; After uniformly transforming the scales of M+1 supplementary target segmentation maps of different scales to the same scale as the corresponding true mask map, they are spliced in the channel dimension to obtain a spliced map; The detection module is used to process the spliced image to obtain the j-th supplementary target segmentation image.
8. The training method according to claim 7, characterized in that: The expression of the j-th training objective is: Among them, A pj A is the pixel set of the target area in the jth supplementary target segmentation map of the infrared small target image sample; gt is the pixel set of the target area in the real mask image of the infrared small target image sample; 9. The training method according to claim 5, characterized in that: The overall training goal also includes: a position training goal; the position training goal is: minimizing the difference loss between the target center point in the target segmentation map of each infrared small target image sample in the training set and the target center point in the corresponding real mask map.
10. A method for detecting small infrared targets, characterized in that: include: The infrared small target to be detected is input into the infrared small target detection system according to any one of claims 1 to 4 to obtain a corresponding target segmentation map, thereby realizing infrared small target detection.
Citation Information
Patent Citations
Infrared small target detection method based on multi-scale self-attention motion background modeling
CN119478349A
Infrared small target detection method and system based on self-adaptive partial self-attention
CN119540539A
Infrared weak and small target detection method based on double-branch attention mechanism
CN119600291A