A Small-Sample SAR Target Recognition Method Based on Hybrid Attention Mechanism
By introducing a hybrid attention mechanism of convolutional blocks and efficient channel attention modules in small sample SAR target recognition, the problems of insufficient feature extraction and overfitting are solved, and the recognition accuracy and model generalization ability are improved.
Patent Information
- Application Number
- CN202310772323.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-28
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-06-28
AI Technical Summary
The existing small sample SAR target recognition method lacks features extracted in the feature extraction network, resulting in overfitting problems and unsatisfactory recognition results.
The convolutional block attention module and the efficient channel attention module are introduced into the deep network with densely connected networks. The semantic information of the high-level feature map is enhanced through the hybrid attention mechanism, and the algorithm's ability to obtain the details and position information of the target object is improved.
It improves the accuracy of small sample target recognition, enhances the generalization ability of the model, and effectively improves the accuracy of image recognition in scenarios with fewer data sets, single categories and complex categories.
Smart Images

Figure CN116824374B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to synthetic aperture radar technology, and particularly to a small-sample SAR target recognition method based on a hybrid attention mechanism. Background Art
[0002] Synthetic aperture radar has the advantages of all-weather and all-day operation and has become an important ground observation device. SAR image recognition has a wide range of applications in target recognition and plays an important role in military and homeland security. With the development of convolutional neural networks, SAR target recognition methods based on convolutional neural networks have achieved remarkable results. However, most existing methods require a large amount of training data, and preparing such high-quality labeled SAR data is labor-intensive and impractical. Therefore, the research on SAR target recognition based on a small amount of labeled data has always been a challenging and active topic.
[0003] In 2019, a semi-supervised recognition method combining a generative adversarial network and a CNN was proposed. GAN was used to generate unlabeled images, which were used as the input of the CNN together with the original labeled images, so as to achieve effective training and recognition with limited training samples. This method has the ability to improve the accuracy and robustness of the neural network system, and has good application prospects in aspects such as SAR image denoising and SAR image super-resolution reconstruction. In addition, the SAR images presented by the same target at different azimuth angles will be very different, and only using the single azimuth angle observation image of the target does not fully utilize the rich information of the synthetic aperture radar image. Recently, researchers have studied the recognition problem of multi-view images and concluded that multi-view image sequences can provide richer classification information than single images. In 2018, a synthetic aperture radar target recognition method based on multi-feature fusion was proposed. After fusing the two feature information of the intensity feature and the gradient magnitude feature, they were input into the classification network, and the effectiveness of the method was verified through experiments. An idea of using multi-azimuth angle SAR images for target recognition was proposed. The SAR image data of three azimuth angles were synthesized into a pseudo-color image for processing, effectively reducing the difference of the target at different azimuth angles and improving the recognition effect. On this basis, the generation method of multi-view SAR data was improved. Without requiring many original synthetic aperture radar images, a large amount of input for network training can be guaranteed. By adopting a topological structure with multi-input signals in parallel, hierarchical learning and multi-layer fusion, superior recognition performance was achieved, and the requirement for the number of original SAR images was reduced.
[0004] In small-sample SAR image recognition, the effective features extracted by the feature extraction network are insufficient, and overfitting problems are likely to occur, resulting in unsatisfactory recognition effects, the same feature weights in different regions and channels, and the inability to extract effective features, and the spatial response effect of the representative object is not good. Summary of the Invention
[0005] The main object of the present invention is to provide a few-shot SAR target recognition method based on a hybrid attention mechanism.
[0006] Aiming at the problem of insufficient features extracted by the feature extraction network when facing the few-shot problem, which leads to overfitting, by introducing a convolutional block attention module and an efficient channel attention module into the deep network of the dense connection network to enhance the semantic information of the high-level feature map, improve the algorithm's ability to obtain the detail and position information of the target object, and improve the accuracy of few-shot target recognition.
[0007] This method can capture the local spatial information of the input feature map channels through network learning, jointly encode the spatial and channel features, and generate a new feature map to improve the performance of the network. First, a convolutional neural network (CNN) model is applied to extract abstract multi-layer feature maps from the original SAR image. Then, a hybrid attention mechanism of spatial weighting and channel weighting is used to enhance the spatial response effect of the representative object and make full use of the features that do not often appear, so that more discriminative features can be extracted. Finally, the aggregated features are output through a fully connected layer.
[0008] The technical solution adopted by the present invention is: a few-shot SAR target recognition method based on a hybrid attention mechanism, including:
[0009] Using a multi-scale feature fusion strategy to fuse multi-layer features;
[0010] Introducing a hybrid attention mechanism into the dense connection network;
[0011] In this network, the backbone extraction network altogether includes 3 DenseBlocks. After the first DenseBlock and the transition layer, a CBAM attention module is added, and after the second DenseBlock and the transition layer, an ECA attention module is added; each DenseBlock contains 4 Inception units;
[0012] The attention mechanism in the network mainly consists of two parts: a CBAM attention module and an ECA attention module;
[0013] The CBAM attention module is used to measure the importance of different features through weight coefficients. CBAM obtains the weight coefficients by measuring the relationship between different feature information, so as to enhance the key information and ignore the irrelevant information;
[0014] The ECA attention module is used to improve the discrimination ability of features. ECA assigns greater weights to the more discriminative channels in the feature map.
[0015] Furthermore, the CBAM attention module includes a channel attention module CAM and a spatial attention module SAM; the channel attention module CAM generates attention features of channels by using the correlation relationship between channels; the spatial attention module SAM is generated through the spatial relationship between features.
[0016] Furthermore, the CBAM attention module is used to extract key information of the picture in the channel domain and the spatial domain, and obtain corresponding channel feature maps and spatial feature maps;
[0017] Given the intermediate feature map F ∈ R C×H×W as the input, CBAM sequentially derives a one-dimensional channel attention map Mc ∈ RC×1×1 and a two-dimensional spatial attention map Ms ∈ R 1×H×W ; The whole attention process is summarized in formulas (1) and (2):
[0018]
[0019]
[0020] where denotes dot product. During dot product, the attention value of the channel propagates along the spatial direction, while the attention in space diffuses along the channel direction; F″ is the final output;
[0021] The channel attention sub-module adopts the maximum downsampling output and the average downsampling output;
[0022] The spatial attention sub-module uses two similar outputs, aggregates them on the axis of the channel, and then passes this information to the convolutional layer;
[0023] First, average pooling and max pooling operations are used to aggregate the spatial information of the feature map, thereby generating two different spatial environment descriptors, which respectively represent the characteristics of average and max pooling; the two descriptors are then sent to a common network to generate a channel attention map M c ∈R C×1×1 ;
[0024] When calculating the spatial attention, first apply average downsampling and max downsampling operations to the channel axis, and then combine them to generate an effective feature description;
[0025] Use downsampling operations on the channel axis to highlight the information area; in the connected feature descriptors, use the convolutional layer to generate a spatial attention map Ms(F) ∈ R H×W ;
[0026] The channel information of the feature map is aggregated by two downsampling operations to obtain two two-dimensional images, which respectively represent the average downsampling characteristics and the maximum downsampling characteristics on the channels;
[0027] Then, a standard convolutional layer is used to combine them to form a two-dimensional attention characteristic curve; the method for calculating the spatial attention is shown in formula (3):
[0028]
[0029] In the formula, σ is the Sigmoid function, and f7×7 is the convolutional operation with a filter size of 7×7.
[0030] Furthermore, the ECA attention module obtains the aggregated features through global average pooling. ECA generates channel weights by performing a fast 1D convolution of size k, where k is adaptively determined by mapping the channel dimension C;
[0031] ECAt avoids dimensionality reduction and effectively captures cross-channel interactions;
[0032] After global channel average pooling without dimensionality reduction, ECA captures local cross-channel interactions by considering each channel and its k adjacent channels. ECA can be effectively implemented by a fast 1D convolution of size k, where the kernel size k represents the coverage of local cross-channel interactions, that is, how many adjacent channels participate in the attention prediction of a channel;
[0033] ECA adopts a banded matrix Wk to obtain the ability of cross-channel information interaction. Wk represents the weight of the learned channel attention to ensure the efficiency and effectiveness of the attention mechanism;
[0034] The banded matrix Wk can be described as formula (4):
[0035]
[0036] The banded matrix Wk involves K×C parameters, which is less than the number of parameters in SE;
[0037] yi represents the weight of the input channel. To calculate the weight of channel yi, only the interaction between it and its k adjacent channels needs to be considered, as shown in formula (5):
[0038]
[0039] where Ω k i represents the set of K adjacent channels of yi and is implemented by a fast 1D convolution with a kernel size of k:
[0040] Ω = σ(C1D k(y))(6)
[0041] Among them, C1D is a one-dimensional convolution; it is called through an effective channel attention module (ECA) that only contains k parameters;
[0042] The coverage of the interaction is proportional to the number of channels C, that is, there is a mapping φ between k and C:
[0043]
[0044] Given the channel dimension C, the kernel size k can be represented by formula (8):
[0045]
[0046] where |t|odd represents the closest odd number to t.
[0047] Advantages of the present invention:
[0048] The present invention enhances the semantic information of the high-level feature map, improves the ability of the algorithm to obtain the detail and position information of the target object, and improves the accuracy of small-sample target recognition;
[0049] It can capture the local spatial information of the input feature map channels through network learning, jointly encode the spatial and channel features, and generate a new feature map, so as to improve the performance of the network;
[0050] In the case of limited training samples, the generalization ability of the model is further improved;
[0051] When performing image recognition in scenarios with fewer data sets, single categories, and complex scenes, the recognition accuracy is effectively improved;
[0052] Using a hybrid attention mechanism of spatial weighting and channel weighting, it enhances the spatial response effect of the representative object and makes full use of the features that do not often appear, so that more discriminative features can be extracted.
[0053] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The present invention will be further described in detail below with reference to the drawings. Brief Description of the Drawings
[0054] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0055] Figure 1 is the network structure diagram of the present invention;
[0056] Figure 2 is the CBAM structure diagram of the present invention;
[0057] Figure 3 This is the ECA module diagram of the present invention. Detailed implementation manners
[0058] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0059] (1) The features in the deeper layers of the network have a larger receptive field and stronger feature expression ability compared to the features in the shallower layers of the network, and are not easily affected by interference such as occlusion, deformation, and rapid movement. However, due to the downsampling operation, the features in the deeper layers of the network will lose the spatial detail information of the target. These spatial detail information are crucial for the spatial positioning of the target. In a convolutional neural network, the features in the shallower layers preserve the spatial detail information of the target, and the features in the deeper layers have rich semantic information. In order to obtain better recognition results, a multi-scale feature fusion strategy is used to fuse multi-layer features. And in order to improve the feature expression ability in the network, a hybrid attention mechanism is introduced into the densely connected network, as Figure 1 shown.
[0060] (2) In this network, the backbone extraction network altogether includes 3 DenseBlocks, and a CBAM attention module is added after the first DenseBlock and the transition layer, and an ECA attention module is added after the second DenseBlock and the transition layer. Each DenseBlock contains 4 Inception units.
[0061] (3) The attention mechanism in the network mainly consists of two parts: CBAM and ECA. Using CBAM can measure the importance of different features through weight coefficients. CBAM obtains the weight coefficients by measuring the relationship between different feature information, so as to enhance the key information and ignore the irrelevant information; using ECA can improve the discriminative ability of the features, and ECA assigns greater weights to the more discriminative channels in the feature map. By using the hybrid attention mechanism to increase the performance ability, focus on the important features, and suppress the unnecessary features.
[0062] (4) The CBAM module contains two independent sequential sub-modules, a channel attention module (ChannelAttention Module, CAM) and a spatial attention module (Spartial Attention Module, SAM).
[0063] The Channel Attention Module (CAM) generates channel attention features by leveraging the correlation relationships between channels. Since each channel is a feature detector, the focus of the channels is on the input image content. Then, spatial dimension compression is performed on the input feature map, effectively improving the utilization rate of channel attention. During the process of aggregating spatial information, max pooling and average pooling are adopted, and both of these operations can well improve the expression performance of the network. The Spatial Attention Module (SAM) is generated based on the spatial relationships between features. Quite different from channel attention, spatial attention focuses on which regions are regarded as an information area, which can make up for the deficiencies of channel attention. When calculating spatial attention, average pooling and max pooling operations are first applied to the channel axis and then combined to generate an effective feature description. Using the downsampling operation on the channel axis can highlight the information area.
[0064] As Figure 2 shown, the key information extraction of the picture in the channel domain and the spatial domain is realized, and the corresponding channel feature map and spatial feature map are obtained. This not only saves parameters and computing power but also ensures that it can be integrated into the existing network architecture as a plug-and-play module.
[0065] Given the intermediate feature map F ∈ R C×H×W as the input, CBAM sequentially derives the one-dimensional channel attention map Mc ∈ RC×1×1 and the two-dimensional spatial attention map Ms ∈ R 1×H×W . The entire attention process can be summarized as equations (1) and (2):
[0066]
[0067]
[0068] where denotes element-wise multiplication. During element-wise multiplication, the attention values of the channels propagate along the spatial direction, while the spatial attention spreads along the channel direction. F″ is the final output.
[0069] The channel attention sub-module adopts the max pooling output and the average pooling output; the spatial attention sub-module uses similar two outputs, aggregates on the channel axis, and then passes this information to the convolutional layer.
[0070] First, average pooling and max pooling operations are used to aggregate the spatial information of the feature map, thereby generating two different spatial context descriptors, which respectively represent the characteristics of average pooling and max pooling. The two descriptors are then fed into a common network to generate the channel attention map M c ∈ R C×1×1 .
[0071] The spatial attention model is generated through the spatial relationship between features. Different from channel attention, spatial attention focuses on which regions are regarded as an information area, which can make up for the deficiencies of channel attention. When calculating spatial attention, first apply average downsampling and maximum downsampling operations to the channel axis, and then combine them to generate an effective feature description. Using the downsampling operation on the channel axis can highlight the information area. In the connected feature descriptor, a convolutional layer is used to generate a spatial attention map Ms(F) ∈ R that encodes the key and inhibitory positions. H×W 。
[0072] The channel information of the feature map is aggregated using two downsampling operations, obtaining two two-dimensional images. They represent the average downsampling characteristics and the maximum downsampling characteristics on the channel respectively. Then a standard convolutional layer is used to combine them to form a two-dimensional attention characteristic curve. Simply put, the method for calculating spatial attention is shown in Equation (3):
[0073]
[0074] In the formula, σ is the Sigmoid function, and f7×7 is the convolutional operation with a filter size of 7×7.
[0075] (5) As Figure 3 shown, it is the ECA structure diagram. For the aggregated features obtained through global average pooling (GAP), ECA generates channel weights by performing a fast 1D convolution of size k, where k is adaptively determined through the mapping of the channel dimension C. ECAt avoids dimensionality reduction and effectively captures cross-channel interactions. After global channel average pooling without dimensionality reduction, ECA captures local cross-channel interactions by considering each channel and its k adjacent channels, and ECA can be effectively implemented through a fast 1D convolution of size k, where the kernel size k represents the coverage of local cross-channel interactions, that is, how many adjacent channels participate in the attention prediction of one channel.
[0076] ECA adopts a banded matrix Wk to obtain the ability of cross-channel information interaction. Wk represents the weight of the learned channel attention to ensure the efficiency and effectiveness of the attention mechanism. The banded matrix Wk can be described as Equation (4):
[0077]
[0078] The banded matrix Wk involves K×C parameters, which are usually less than the number of parameters in SE. At the same time, it also avoids the mutual independence between different channels. yi represents the weight of the input channel. To calculate the weight of channel yi, only the interaction between it and its adjacent k channels needs to be considered. As shown in formula (5):
[0079]
[0080] In Ω k i represents the set of K adjacent channels of yi. This method can be implemented by a fast 1D convolution with a kernel size of k:
[0081] Ω = σ(C1D k (y))(6)
[0082] where C1D is a one-dimensional convolution. The above method is called through an effective channel attention module (ECA) that only contains k parameters. This scheme can effectively reduce the complexity of the model and can effectively capture local cross-channel interactions.
[0083] The purpose of the ECA model is to correctly capture local cross-channel interaction behaviors, so it is necessary to determine the coverage area of the interaction (that is, the kernel size k of the one-dimensional convolution). For different CNN structures, it can be optimized by manually adjusting the convolution blocks with different numbers of channels. However, manual adjustment requires the use of the cross-validation method, which consumes a large amount of computing resources.
[0084] The coverage range of the interaction (that is, the kernel size k of the one-dimensional convolution) is proportional to the number of channels C, that is, there is a mapping φ between k and C:
[0085]
[0086] Given the channel dimension C, the kernel size k can be represented by formula (8):
[0087]
[0088] where |t|odd represents the closest odd number to t.
[0089] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A small - sample SAR target recognition method based on a hybrid attention mechanism, characterized in that, it includes: Using a multi - scale feature fusion strategy to fuse multi - layer features; Introducing a hybrid attention mechanism into a dense connection network; In this network, the backbone extraction network altogether contains 3 DenseBlocks. After the first DenseBlock and the transition layer, a CBAM attention module is added. After the second DenseBlock and the transition layer, an ECA attention module is added; each DenseBlock contains 4 Inception units; The attention mechanism in the network mainly consists of two parts: a CBAM attention module and an ECA attention module; The CBAM attention module is used to measure the importance of different features through weight coefficients. CBAM obtains weight coefficients by measuring the relationship between different feature information, so as to enhance key information and ignore irrelevant information; The ECA attention module is used to improve the discriminative ability of features. ECA assigns greater weights to the more discriminative channels in the feature map.
2. The small - sample SAR target recognition method based on the hybrid attention mechanism according to claim 1, characterized in that, the CBAM attention module includes a channel attention module CAM and a spatial attention module SAM; the channel attention module CAM generates channel attention features by using the correlation relationship between channels; The spatial attention module SAM is generated by the spatial relationship between features.
3. The small - sample SAR target recognition method based on the hybrid attention mechanism according to claim 1, characterized in that, the CBAM attention module is used to extract key information of the picture in the channel domain and the spatial domain, and obtain corresponding channel feature maps and spatial feature maps; Given the intermediate feature map F ∈ R C×H×W as input, CBAM sequentially infers the one-dimensional channel attention map Mc ∈ RC×1×1 and the two-dimensional spatial attention map Ms ∈ R 1×H×W ; the entire attention process is summarized as equations (1) and (2): Among them represents dot product. During dot product, the attention values of channels propagate along the spatial direction, while the spatial attention diffuses along the channel direction; F″ is the final output; The channel attention sub - module adopts the maximum down - sampling output and the average down - sampling output; The spatial attention sub - module uses two similar outputs, aggregates them on the channel axis, and then transfers this information to the convolutional layer; First, average pooling and max pooling operations are used to aggregate the spatial information of the feature map, thereby generating two different spatial context descriptors that represent the characteristics of average and max pooling respectively; the two descriptors are then fed into a common network to generate the channel attention map M c ∈R C×1×1 ; When calculating the spatial attention, first apply the average down - sampling and maximum down - sampling operations to the channel axis, and then combine them to generate an effective feature description; Use downsampling operations on the channel axis to highlight the information area; in the connected feature descriptors, a convolutional layer is used to generate a spatial attention map Ms(F) ∈ R that encodes the key and inhibitory positions H×W ; The channel information of the feature map is aggregated by using two down - sampling operations to obtain two two - dimensional images, which respectively represent the average down - sampling characteristics and the maximum down - sampling characteristics on the channel; Then use a standard convolutional layer to combine them to form a two - dimensional attention characteristic curve; the method for calculating the spatial attention is shown in formula (3): In the formula, σ is the Sigmoid function, and f 7×7 is the convolutional operation with a filter size of 7×7.
4. The small - sample SAR target recognition method based on the hybrid attention mechanism according to claim 1, characterized in that, the ECA attention module obtains aggregated features through global average pooling. ECA generates channel weights by performing a fast 1D convolution of size k, where k is adaptively determined by the mapping of the channel dimension C; ECAt avoids dimensionality reduction and effectively captures cross - channel interactions; After global channel average pooling without dimensionality reduction, ECA captures local cross-channel interactions by considering each channel and its k adjacent channels, and ECA can be effectively implemented by a fast 1D convolution of size k, where the kernel size k represents the coverage of local cross-channel interactions, that is, how many adjacent channels participate in the attention prediction of one channel; ECA adopts a banded matrix Wk to obtain the ability of cross-channel information interaction. Wk represents the weight of the learned channel attention to ensure the efficiency and effectiveness of the attention mechanism; The banded matrix Wk can be described by formula (4): There are K×C parameters involved in the banded matrix Wk, which is less than the number of parameters in SE; yi represents the weight of the input channel. To calculate the weight of channel yi, only the interaction between it and its adjacent k channels needs to be considered, as shown in formula (5): where Ω k i represents the set of K adjacent channels of yi, implemented by a fast 1D convolution with a kernel size of k: Ω=σ(C1D k (y))(6) where C1D is a one-dimensional convolution; it is called by an effective channel attention module (ECA) that only contains k parameters; The coverage of the interaction is proportional to the number of channels C, that is, there is a mapping φ between k and C: Given the channel dimension C, the kernel size k can be represented by formula (8): where |t|odd represents the closest odd number to t.
Citation Information
Patent Citations
Synthetic aperture radar (SAR) image target detection method
US20230169623A1
Contextual visual-based SAR target detection method and apparatus, and storage medium
US20230184927A1