SAR Image Aircraft Target Detection Method Based on Joint Attention Mechanism
The joint attention mechanism in SAR image detection enhances target-specific features and contextual awareness, improving detection accuracy by integrating local and global attention features in a pyramid network, addressing the limitations of existing methods that neglect background information.
Patent Information
- Application Number
- CN202211065572.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-08-31
Smart Images

Figure CN115410102B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of radar target detection, and in particular to a method for detecting aircraft targets in SAR images based on a joint attention mechanism. Background Art
[0002] Synthetic Aperture Radar (SAR) technology is a pulsed radar technology that uses a mobile radar carried on a satellite or an aircraft to obtain high-precision radar target images of geographical areas. It has the working capabilities of all-weather and all-time and certain penetration capabilities. Given these advantages, it is widely used in fields such as mineral exploration, marine environment monitoring, and military defense. In particular, the research on the detection of aircraft targets is of great significance both in the military field and the civilian field. Therefore, the research on the detection of aircraft in SAR images has received extensive attention from scholars at home and abroad.
[0003] Traditional aircraft target detection mainly focuses on the structural features and scattering features of aircraft targets. For the structural features, the detection of aircraft targets is mainly completed through the unique structures of aircraft, such as structures like "Y" and "T". The scattering features are due to the special imaging mechanism of SAR images, and the target is usually composed of a series of strong scattering points. The scattering features are specifically realized by the geometric features, gray-scale statistical features, and texture features of the target to detect aircraft targets.
[0004] In recent years, with the continuous development and popularization of deep learning theories and methods, good results have been achieved in many fields. As an important part of image interpretation, target detection is one of the core issues in SAR image understanding. Deep features have strong descriptive capabilities and show good effects in both detection and classification. However, most of the existing deep learning methods mainly focus on the feature information of aircraft targets themselves in SAR images through convolution and local attention, without overly focusing on background information and clutter. Although this excludes some interference information, it also ignores the information of the surrounding positions of aircraft targets, making it impossible to compare the differences between aircraft targets and surrounding information, resulting in poor accuracy of SAR image target detection. Therefore, how to design a SAR image target detection method that can take into account both the feature information of aircraft targets themselves and the surrounding position information is a technical problem that needs to be solved urgently. Summary of the Invention
[0005] Aiming at the deficiencies of the above-mentioned existing technologies, the technical problem to be solved by the present invention is: how to provide a method for detecting aircraft targets in SAR images based on a joint attention mechanism, so as to effectively fuse the local attention features and global attention features of SAR images, and further be able to take into account both the feature information of aircraft targets themselves and the surrounding position information, thereby improving the accuracy of SAR image target detection.
[0006] To solve the above technical problems, the present invention adopts the following technical solutions:
[0007] A method for detecting aircraft targets in SAR images based on a joint attention mechanism, comprising:
[0008] S1: Obtain the SAR image to be detected;
[0009] S2: Input the SAR image to be detected into the trained target detection model, and output the corresponding target detection prediction value;
[0010] When training the target detection model, first input a training set containing several SAR images into the target detection model; secondly, extract deep feature maps of different levels of the SAR image through a deep neural network; then input the corresponding deep feature maps of each level into the corresponding joint attention layer of the pyramid network for extracting local and global joint attention features, and at the same time splice the output of the upper joint attention layer in the pyramid network with the deep feature map input to the adjacent lower joint attention layer as the input of the adjacent lower joint attention layer; then respectively make predictions based on the joint attention feature maps output by each joint attention layer of the pyramid network to obtain the corresponding prediction boxes and classification prediction probabilities; finally, generate the target detection prediction value through each prediction box and classification prediction probability, and perform model training based on the target detection prediction value;
[0011] S3: Based on the target detection prediction value output by the target detection model, implement the target detection of the SAR image to be detected.
[0012] Preferably, in step S2, ResNet50 is used as the backbone network of the deep neural network for extracting deep feature maps.
[0013] Preferably, in step S2, the deep feature maps of different levels refer to deep feature maps with different scales and numbers of channels.
[0014] Preferably, each joint attention layer of the pyramid network includes a local attention module for extracting local attention feature maps and a global attention module for extracting global attention feature maps; add the local attention feature map and the global attention feature map to obtain the joint attention feature map of the corresponding joint attention layer.
[0015] Preferably, the local attention module includes two independent attention branches of channel attention and spatial attention;
[0016] The input feature maps are respectively input into two attention branches to extract the channel attention feature map and the spatial attention feature map; then the channel attention feature map and the spatial attention feature map are added and after passing through the Sigmoid activation operation, they are multiplied by the input feature map and a residual connection is performed to obtain the local attention feature map.
[0017] Channel attention: First, global average pooling is performed on the input feature map to obtain the channel vector Fc ∈ C×1×1; then the cross-channel attention is estimated from the channel vector Fc through a multi-layer perceptron with one hidden layer; finally, the scale of the output of the spatial branch is adjusted through a batch normalization layer to obtain the channel attention feature map;
[0018] Spatial attention: First, the input feature map is projected from C×H×W to the reduced-dimensional C / r×H×W through a 1×1 convolution; then two 3×3 dilated convolutions are used to utilize the context information; finally, the feature map is simplified again to a 1×H×W spatial attention map through a 1×1 convolution, and the scale of the output of the spatial branch is adjusted by applying a batch normalization layer to obtain the spatial attention feature map.
[0019] Preferably, the global attention module first uses three 1×1 convolutions on the input feature map to obtain three feature maps Q, K, and V, and the number of channels of the feature maps Q and K is less than that of the input feature map; then an Affinity operation is performed on the feature maps Q and K: the vectors on each channel dimension in the feature map Q are matrix-multiplied with all the vectors in the horizontal and vertical directions at the corresponding positions in the feature map K, and then the softmax function is used to perform weighted averaging on the channel dimension to obtain the feature map A; finally, the feature map A and the feature map V are fused, and the fused feature map is subjected to a residual connection with the input feature map to obtain the global attention feature map.
[0020] Preferably, in step S2, the joint attention feature map p output by the (i + 1)-th layer joint attention layer in the pyramid network i+1 is upsampled by a factor of two and concatenated with the feature map obtained by passing the depth feature map input to the i-th layer joint attention layer through a 1×1 convolution to obtain the corresponding feature map as the input of the i-th layer joint attention layer.
[0021] Preferably, in step S2, the joint attention feature maps output by each network layer of the pyramid network are respectively input into the corresponding region proposal network and region of interest pooling layer: the region proposal network performs a sliding window operation on the joint attention feature map, and two CNNs are respectively used as feature extractors within the sliding window to extract regression box features and class features, obtaining proposal boxes of the target; then, the region of interest pooling layer performs pooling processing on the proposal boxes of the target to adjust the sizes of the proposal boxes, and finally, a feature map with proposal boxes is obtained as the input of the fully connected layer, thereby outputting the regression parameters and classification parameters of the prediction boxes.
[0022] Preferably, the coordinates of the prediction boxes are calculated based on the regression parameters of the prediction boxes, and the classification parameters of the prediction boxes are processed by the softmax function to obtain the classification prediction probabilities of each category; then, based on the prediction box coordinates, the prediction boxes and their classification prediction probabilities are mapped onto the SAR image, and the prediction boxes are cropped to adjust the coordinates of the out-of-bounds prediction boxes to the SAR image boundary; finally, the target categories with low probabilities are removed, and non-maximum suppression processing is performed to suppress the redundant prediction boxes, obtaining the SAR image with prediction boxes and classification prediction probabilities as the target detection prediction values.
[0023] Preferably, in step S2, the prediction boxes and the classification prediction probabilities are jointly trained according to the number of training iterations and the stochastic gradient descent algorithm in combination with the cross-entropy loss function and the SmoothL1 loss function to complete the training of the target detection model.
[0024] The method for detecting aircraft targets in SAR images based on the joint attention mechanism in the present invention has the following beneficial effects:
[0025] In the present invention, the pyramid network for extracting local and global joint attention features respectively extracts the local attention feature and the global attention feature of the depth feature map to obtain the joint attention feature. On the one hand, the local attention feature guides the network to pay more attention to the feature information of the aircraft target itself in the SAR image, but not overly focus on the background and clutter; on the other hand, the global attention feature makes up for the neglect of the surrounding position information of the aircraft target caused by convolution and local attention, and better realizes the detection of the aircraft target by comparing the differences with the surrounding position information. Therefore, by fusing the local attention feature and the global attention feature of the SAR image, the feature information of the aircraft target itself and the surrounding position information can be taken into account, thereby improving the accuracy of SAR image target detection.
[0026] Secondly, in the pyramid network of the present invention, the output of the upper joint attention layer is fused with the depth feature map input by the adjacent lower joint attention layer as the input of the adjacent lower joint attention layer, enabling effective fusion of low-dimensional features that retain texture features and high-dimensional features that retain semantic information. Among them, the high-dimensional features are highly correlated with the targets in the SAR image, contain rich target information, which is beneficial to improving the correct detection rate of the targets. However, the target positions are relatively rough. The low-dimensional features can provide discriminative target information and meet the requirements of gray-scale and rotation invariance, having the advantage of accurate target positions, but containing less semantic information of the features. Therefore, fusing high-dimensional features with low-dimensional features can not only provide richer discriminative target information for the target detection model while ensuring the relevance of aircraft targets, but also provide more accurate target positions, thereby further improving the accuracy of SAR image target detection.
[0027] Finally, the present invention extracts depth feature maps of different levels (different scales and numbers of channels) of the SAR image, which can enrich the feature information of aircraft targets through multi-scale feature fusion and solve the problem of different sizes of aircraft targets, thereby further improving the accuracy of SAR image target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to make the objectives, technical solutions, and advantages of the invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings, where:
[0029] Figure 1 FIG. is the logic block diagram of the SAR image aircraft target detection method based on the joint attention mechanism;
[0030] Figure 2 FIG. is the network structure diagram of the target detection model;
[0031] Figure 3 FIG. is the schematic diagram of the network structure of the pyramid network;
[0032] Figure 4 FIG. is the schematic diagram of the framework of the local attention module;
[0033] Figure 5 FIG. is the schematic diagram of the framework of the global attention module;
[0034] Figure 6 FIG. is a partial image of the constructed aircraft detection dataset;
[0035] Figure 7 FIG. is the detection result diagram of the existing deep learning target detection model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the figures herein can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the figures is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0037] It should be noted that like reference numerals and letters denote like items in the following figures. Therefore, once an item is defined in one figure, it does not require further definition and explanation in subsequent figures. In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the inventive product is customarily placed during use. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and should not be construed as indicating or implying relative importance. In addition, terms such as "horizontal" and "vertical" do not mean that the components are required to be absolutely horizontal or hanging, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined. In the description of the present invention, it should also be noted that unless otherwise clearly defined and limited, the terms "set", "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0038] The following is a more detailed description through specific embodiments:
[0039] Embodiment:
[0040] In this embodiment, a method for detecting aircraft targets in SAR images based on a joint attention mechanism is disclosed.
[0041] Such asFigure 1 As shown in Figure 1 , a method for detecting aircraft targets in SAR images based on a joint attention mechanism includes:
[0042] S1: Obtain the SAR image to be detected;
[0043] S2: Input the SAR image to be detected into the trained target detection model, and output the corresponding target detection prediction value;
[0044] When training the target detection model, first input a training set containing several SAR images into the target detection model; secondly, extract the depth feature maps C of different levels of the SAR image through a deep neural network i (i = 1, 2, 3, 4); then input the depth feature maps of each level into the corresponding joint attention layer of the pyramid network for extracting local and global joint attention features, and at the same time splice the output of the upper joint attention layer in the pyramid network with the depth feature map input to the adjacent lower joint attention layer as the input of the adjacent lower joint attention layer; then respectively make predictions based on the joint attention feature maps output by each joint attention layer of the pyramid network to obtain the corresponding prediction boxes and classification prediction probabilities; finally, generate the target detection prediction value through each prediction box and classification prediction probability, and perform model training based on the target detection prediction value;
[0045] S3: Implement the target detection of the SAR image to be detected based on the target detection prediction value output by the target detection model.
[0046] In this embodiment, the target detection prediction value refers to the SAR image with prediction boxes and classification prediction probabilities. According to the SAR image with prediction boxes and classification prediction probabilities, the position and type of the aircraft target in the image can be detected.
[0047] The present invention extracts the local attention feature and the global attention feature of the depth feature map through the pyramid network for extracting local and global joint attention features to obtain the joint attention feature. On the one hand, the local attention feature guides the network to pay more attention to the feature information of the aircraft target itself in the SAR image, but does not overly focus on the background and clutter; on the other hand, the global attention feature makes up for the neglect of the surrounding position information of the aircraft target caused by convolution and local attention, and better realizes the detection of the aircraft target by comparing the differences with the surrounding position information. Therefore, by fusing the local attention feature and the global attention feature of the SAR image, the feature information of the aircraft target itself and the surrounding position information can be taken into account, thereby improving the accuracy of SAR image target detection.
[0048] Secondly, in the pyramid network of the present invention, the output of the upper-level joint attention layer is fused with the depth feature map input by the adjacent lower-level joint attention layer as the input of the adjacent lower-level joint attention layer, enabling effective fusion of low-dimensional features that retain texture features and high-dimensional features that retain semantic information. Among them, the high-dimensional features are highly correlated with the targets in the SAR image, contain rich target information, and are beneficial to improving the correct detection rate of the targets. However, the target positions are relatively rough. The low-dimensional features can provide discriminative target information and meet the requirements of gray-scale and rotation invariance, with the advantage of accurate target positions, but contain less feature semantic information. Therefore, fusing high-dimensional features with low-dimensional features can not only provide richer discriminative target information for the target detection model while ensuring the relevance of aircraft targets, but also provide more accurate target positions, thereby further improving the accuracy of SAR image target detection.
[0049] Finally, the present invention extracts depth feature maps of different levels (different scales and numbers of channels) of the SAR image, which can enrich the feature information of aircraft targets through multi-scale feature fusion and solve the problem of different sizes of aircraft targets, thereby further improving the accuracy of SAR image target detection.
[0050] In the specific implementation process, ResNet50 is used as the backbone network of the deep neural network for extracting depth feature maps.
[0051] In the specific implementation process, depth feature maps of different levels refer to depth feature maps with different scales and numbers of channels.
[0052] The present invention extracts depth feature maps of different levels (different scales and numbers of channels) of the SAR image, which can enrich the feature information of aircraft targets through multi-scale feature fusion and solve the problem of different sizes of aircraft targets, thereby further improving the accuracy of SAR image target detection.
[0053] In the specific implementation process, in combination with Figure 3 as shown, the joint attention feature map P output by the (i + 1)-th layer joint attention layer in the pyramid network i+1 is upsampled by a factor of two and concatenated with the feature map obtained by 1×1 convolution of the depth feature map input by the i-th layer joint attention layer to obtain the corresponding feature map as the input of the i-th layer joint attention layer. In combination with Figure 2 as shown, the feature map C5 is directly downsampled by a factor of two from the feature map C4, and then the feature map C5 directly passes through the joint attention layer to obtain the joint attention feature map P5.
[0054] In the pyramid network of the present invention, the output of the upper joint attention layer is fused with the depth feature map input by the adjacent lower joint attention layer as the input of the adjacent lower joint attention layer, so that it is possible to effectively fuse the low-dimensional features that retain texture features and the high-dimensional features that retain semantic information. Among them, the high-dimensional features are highly correlated with the targets in the SAR image, contain rich target information, and are beneficial to improving the correct detection rate of targets. However, the target positions are relatively rough. The low-dimensional features can provide discriminative target information and meet the requirements of gray-scale and rotation invariance, and have the advantage of accurate target positions, but contain less feature semantic information. Therefore, fusing high-dimensional features with low-dimensional features can not only provide richer discriminative target information for the target detection model while ensuring the relevance of aircraft targets, but also provide more accurate target positions, thereby further improving the accuracy of SAR image target detection.
[0055] Each joint attention layer of the pyramid network includes a local attention module (Bottleneck Attention Module, BAM) for extracting local attention feature maps and a global attention module (Criss-Cross Attention, CCA) for extracting global attention feature maps; the local attention feature map and the global attention feature map are added together to obtain the joint attention feature map of the corresponding joint attention layer.
[0056] As Figure 4 shown, the local attention module includes two independent attention branches: channel attention and spatial attention;
[0057] The input feature map is respectively input into the two attention branches to extract the channel attention feature map and the spatial attention feature map; then the channel attention feature map and the spatial attention feature map are added together, and after passing through the Sigmoid activation operation, they are multiplied by the input feature map and a residual connection is performed to obtain the local attention feature map.
[0058] Channel attention: First, global average pooling is performed on the input feature map to obtain a channel vector Fc ∈ C×1×1; then, a multi-layer perceptron (MLP) with a hidden layer is used to estimate cross-channel attention from the channel vector Fc; finally, a batch normalization (BN) layer is used to adjust the ratio of the output of the spatial branch to obtain the channel attention feature map; where the hidden activation size of the multi-layer perceptron (MLP) is set to C / r×1×1, and r is the reduction rate;
[0059] Spatial attention: First, project the input feature map from C×H×W to the reduced-dimensional C / r×H×W through a 1×1 convolution, using the same reduction ratio r as in the channel attention branch; then utilize context information through two 3×3 dilated convolutions; finally, simplify the feature map back to a 1×H×W spatial attention map through a 1×1 convolution, and apply a batch normalization layer to adjust the scale of the output of the spatial branch to obtain the spatial attention feature map.
[0060] As Figure 5 shown, the global attention module first uses three 1×1 convolutions on the input feature map to obtain three feature maps Q, K, and V. Among them, the number of channels of feature maps Q and K is less than that of the input feature map; then perform an Affinity operation on feature maps Q and K: multiply the vectors on each channel dimension in feature map Q with all the vectors in the horizontal and vertical directions at the corresponding positions in feature map K, and then use the softmax function to perform weighted averaging on the channel dimension to obtain feature map A; finally, fuse feature map A and feature map V, and perform a residual connection between the fused feature map and the input feature map to obtain the global attention feature map.
[0061] In the present invention, the pyramid network for extracting local and global joint attention features extracts the local attention feature and the global attention feature of the depth feature map respectively to obtain the joint attention feature. On the one hand, the network is guided by the local attention feature to pay more attention to the feature information of the aircraft target itself in the SAR image, but not overly focus on the background and clutter; on the other hand, the global attention feature makes up for the neglect of the position information around the aircraft target caused by convolution and local attention, and better realizes the detection of the aircraft target by comparing the differences with the surrounding position information. Therefore, by fusing the local attention feature and the global attention feature of the SAR image, the feature information of the aircraft target itself and the surrounding position information can be taken into account, thereby improving the accuracy of SAR image target detection.
[0062] In the specific implementation process, input the joint attention feature maps output by each network layer of the pyramid network into the corresponding Region Proposal Network (RPN) and Region of Interest (ROI) pooling layer respectively: perform a sliding window operation on the joint attention feature map through the region proposal network, and extract the regression box feature and the class feature respectively through two CNNs as the feature extractors within the sliding window to obtain the proposal box of the target (the working principle is similar to the existing Fast R-CNN model); then perform pooling processing on the proposal box of the target through the ROI pooling layer to adjust the size of the proposal box, and finally obtain the feature map with the proposal box as the input of the fully connected layer, and further output the regression parameters and classification parameters of the prediction box.
[0063] In this embodiment, the RPN, ROI pooling layer, and fully connected layer are all existing mature models. The present invention does not improve the structure and working logic of the models, but only applies the models to process the joint attention feature map of the present invention, and then obtains the regression parameters and classification parameters of the prediction boxes. Among them, the regression parameters and classification parameters are obtained by training the fully connected layer through coordinate information and class probabilities, and then the coordinate information and class probabilities are output through the fully connected layer. The class refers to whether the object detected by the RPN is a target.
[0064] The coordinates of the prediction boxes are calculated according to the regression parameters of the prediction boxes, and the classification parameters of the prediction boxes are processed by the softmax function to obtain the classification prediction probabilities of each class. Then, according to the coordinates of the prediction boxes, the prediction boxes and their classification prediction probabilities are mapped to the SAR image, and the prediction boxes are cropped to adjust the coordinates of the out-of-bounds prediction boxes to the boundaries of the SAR image. Finally, the target classes with low probabilities are removed, and non-maximum suppression processing is performed to suppress the redundant prediction boxes, and the SAR image with prediction boxes and classification prediction probabilities is obtained as the target detection prediction value.
[0065] In this embodiment, calculating the coordinates of the prediction boxes, calculating the classification prediction probabilities, mapping the prediction boxes and their classification prediction probabilities to the SAR image, cropping the prediction boxes, and performing non-maximum suppression processing are all completed by existing mature means, and the present invention only needs to obtain the SAR image with prediction boxes and classification prediction probabilities through existing means.
[0066] In the specific implementation process, the prediction boxes and classification prediction probabilities in the target detection prediction value are jointly trained according to the number of training iterations (Epochs) and the stochastic gradient descent algorithm combined with the cross-entropy loss function and the SmoothL1 loss function to complete the training of the target detection model.
[0067] In this embodiment, the target detection model is trained by existing mature means. Among them, the number of iterations (Epochs) and the stochastic gradient descent algorithm are both existing mature technologies, and the present invention does not improve them. The joint training of the cross-entropy loss function and the SmoothL1 loss function means that the sum of the cross-entropy loss and the SmoothL1 loss is used as the training loss of the model, and no changes are made to the cross-entropy loss function and the SmoothL1 loss function themselves, but only the true labels and predicted labels in their formulas are replaced with the true classes and classification prediction probabilities in the present invention.
[0068] To better illustrate the advantages of the technical solution of the present invention, the following experiment is disclosed in this embodiment.
[0069] 1. Evaluation Metrics
[0070] 1) Average Precision, which adopts six average precision metrics of Microsoft COCO, including AP, AP50, AP75, APs, APm, and APl. Among them, AP evaluates the average precision score through ten Intersection of Union (IoU) thresholds of 0.50:0.05:0.95 between the prediction results and the ground truth. AP50 and AP75 are the average precision scores evaluated at IoU of 0.5 and 0.75 respectively. APs, APm, and APl refer to the average precision scores of small, medium, and large aircraft detection methods at ten IOU thresholds. The specific calculation is as follows:
[0071]
[0072] Among them, p represents precision, r represents recall rate, and p is a function with r as a parameter.
[0073] 2) Precision refers to the proportion of positive samples that are correctly detected as positive samples among all positive samples.
[0074] 3) Recall rate is the proportion of predicted samples that are correctly detected as positive samples. The calculation methods of the two evaluation metrics are as follows:
[0075]
[0076]
[0077] 2. Experimental Data
[0078] Construct a SAR image aircraft target detection dataset, obtaining 1872 images of 256×256. Table 1 gives the specific information of the dataset division. Some examples of the dataset are shown in Figure 6 .
[0079] Table 1 Aircraft Target Detection Dataset
[0080]
[0081] 3. Model Settings
[0082] For Faster R-CNN that selects ResNet50 as the backbone network and integrates the Feature Pyramid Network (FPN), the backbone network is initialized with Imagenet pre-trained weights. The dataset is randomly divided into a training set and a test set at a ratio of 8:2. The model is trained using the Stochastic Gradient Descent (SGD) algorithm, with the learning rate set to 0.005, the weight decay set to 0.0005, and the momentum set to 0.9. The batch size of gradient descent is 2. The cross-entropy loss function is used as the classification and regression loss, and the Smooth L1 loss is used as the bounding box regression loss. The total number of training iterations is set to 15 epochs. In BAM, the hyperparameters r = 16 and d = 4.
[0083] 4. Model Performance Evaluation
[0084] To verify the performance of the object detection model proposed in the present invention, this experiment respectively detects the original Faster R-CNN and the methods combined with BAM, CCA attention, and local and global joint attention. The results are shown in Table 2.
[0085] Table 2 Model Performance Evaluation Table
[0086]
[0087] As can be seen from Table 2, after adding local attention BAM and global attention CCA, AP50 is increased by 0.6% and 0.5% respectively, indicating the effectiveness of BAM and CCA for the detection network. For the case of using both BAM and CCA, AP50 is increased by 1.0% compared with the original Faster R-CNN, indicating that local attention and global attention promote each other.
[0088] 5. Performance Comparison
[0089] To verify the performance of the model proposed in the present invention, it is compared with a variety of CNN-based object detection networks.
[0090] Table 3 presents the experimental results of different network detection performances. It can be found from Table 3 that the performance of the model proposed by the present invention is better than that of other detection networks, and the detection accuracy for aircraft target detection in the experimental dataset reaches 90.2%. The detection accuracy of the model proposed by the present invention is 1.1% higher than that of the basic Faster R-CNN. In addition, the detection performance of the basic Faster R-CNN is also better than that of other typical target detection networks in the table. This is because single-stage networks such as RetinaNet, SSD-300, and YOLOv3 do not have an RPN network similar to that in Faster R-CNN, and do not achieve the ability to pre-perceive target regions. They directly detect aircraft targets from the entire input image, so the detection effect is not as accurate as that of Faster R-CNN.
[0091] Table 3 Comparison Table of Model Performances
[0092]
[0093] Figure 7 is the detection result diagram of different network models, where (a) is the ground truth of different scenarios, (b) is the detection result of Faster R-CNN, (c) is the detection result of RetinaNet network, (d) is the detection result of SSD-300, and (e) is the detection result of the model proposed by the present invention. In addition, each column represents a different scenario, and (I), (II), (III), and (IV) are used to represent the four scenarios respectively.
[0094] As can be seen from Figure 7 , for scenarios (I) and (II), both Faster R-CNN and RetinaNet have obvious false alarms or repeated detections due to strong background clutter interference; SSD-300 has a certain number of missed detections and ignores incomplete aircraft; while the model proposed by the present invention can better solve the problems of false alarms and missed detections because it adopts a feature pyramid with local and global joint attention. At the same time, in scenarios (III) and (IV), the aircraft targets are small and densely arranged. Both the RetinaNet and SSD-300 models have missed detections, and the detection box positioning is inaccurate. The performance of the model proposed by the present invention is significantly better than these two networks.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and do not limit the technical solutions. Those of ordinary skill in the art should understand that any modifications or equivalent replacements to the technical solutions of the present invention without departing from the purpose and scope of the technical solutions should be covered by the scope of the claims of the present invention.
Claims
1. A method for detecting aircraft targets in SAR images based on a joint attention mechanism, characterized in that, Including: S1: Obtain the SAR image to be detected; S2: Input the SAR image to be detected into the trained object detection model, and output the corresponding object detection prediction value; When training the object detection model, first input the training set containing several SAR images into the object detection model; secondly, extract the depth feature maps of different levels of the SAR image through the deep neural network; then input the depth feature maps of each level into the corresponding joint attention layer of the pyramid network for extracting local and global joint attention features, and at the same time splice the output of the upper joint attention layer in the pyramid network with the depth feature map input to the adjacent lower joint attention layer as the input of the adjacent lower joint attention layer; then respectively perform predictions based on the joint attention feature maps output by each joint attention layer of the pyramid network to obtain the corresponding prediction boxes and classification prediction probabilities; finally, generate the object detection prediction value through each prediction box and classification prediction probability, and perform model training based on the object detection prediction value; In step S2, each joint attention layer of the pyramid network includes a local attention module for extracting local attention feature maps and a global attention module for extracting global attention feature maps; add the local attention feature map and the global attention feature map to obtain the joint attention feature map of the corresponding joint attention layer; The local attention module includes two independent attention branches: channel attention and spatial attention; Input the input feature map into the two attention branches respectively to extract the channel attention feature map and the spatial attention feature map; Then add the channel attention feature map and the spatial attention feature map, and after passing through the Sigmoid activation operation, multiply it with the input feature map and perform residual connection to obtain the local attention feature map; Upsample the joint attention feature map p output by the joint attention layer of the (i + 1)-th layer in the pyramid network by a factor of two and concatenate it with the feature map obtained by performing 1×1 convolution on the depth feature map input to the joint attention layer of the i-th layer to obtain the corresponding feature map as the input to the joint attention layer of the i-th layer; i+1 S3: Implement the object detection of the SAR image to be detected based on the object detection prediction value output by the object detection model.
2. The method for detecting aircraft targets in SAR images based on the joint attention mechanism according to claim 1, wherein: In step S2, ResNet50 is used as the backbone network of the deep neural network for extracting depth feature maps.
3. The method for detecting aircraft targets in SAR images based on the joint attention mechanism according to claim 1, wherein: In step S2, the depth feature maps of different levels refer to the depth feature maps with different scales and numbers of channels.
4. The method for detecting aircraft targets in SAR images based on the joint attention mechanism according to claim 1, wherein: Channel attention: First, perform global average pooling on the input feature map to obtain the channel vector Fc ∈ C×1×1; then estimate the cross-channel attention from the channel vector Fc through a multi-layer perceptron with one hidden layer; finally, adjust the ratio of the output of the spatial branch through the batch normalization layer to obtain the channel attention feature map; Spatial attention: First, project the input feature map from C×H×W to the reduced-dimensional C / r×H×W through a 1×1 convolution; then use two 3×3 dilated convolutions to utilize the context information; finally, simplify the feature map to a 1×H×W spatial attention map through a 1×1 convolution, and apply the batch normalization layer to adjust the ratio of the output of the spatial branch to obtain the spatial attention feature map.
5. The method for detecting aircraft targets in SAR images based on the joint attention mechanism according to claim 1, wherein: The global attention module first uses three 1×1 convolutions on the input feature map to obtain three feature maps Q, K, and V. The number of channels of feature maps Q and K is less than that of the input feature map. Then, an Affinity operation is performed on feature maps Q and K: the vectors in each channel dimension of feature map Q are matrix-multiplied with all the vectors in the horizontal and vertical directions at the corresponding positions in feature map K, and then the softmax function is used to perform weighted averaging on the channel dimension to obtain feature map A. Finally, feature map A and feature map V are fused, and the fused feature map is connected to the input feature map through a residual connection to obtain the global attention feature map.
6. The method for detecting aircraft targets in SAR images based on the joint attention mechanism according to claim 1, wherein: In step S2, the joint attention feature maps output by each network layer of the pyramid network are respectively input into the corresponding region proposal network and region of interest pooling layer: the region proposal network performs a sliding window operation on the joint attention feature map, and two CNNs are respectively used as feature extractors within the sliding window to extract regression box features and class features to obtain the proposal boxes of the targets. Then, the region of interest pooling layer performs pooling processing on the proposal boxes of the targets to adjust the sizes of the proposal boxes, and finally obtains the feature map with proposal boxes as the input of the fully connected layer, and further outputs the regression parameters and classification parameters of the prediction boxes.
7. The method for detecting aircraft targets in SAR images based on the joint attention mechanism according to claim 6, characterized in that: The coordinates of the prediction boxes are calculated according to the regression parameters of the prediction boxes, and the classification parameters of the prediction boxes are processed by the softmax function to obtain the classification prediction probabilities of each category. Then, according to the prediction box coordinates, the prediction boxes and their classification prediction probabilities are mapped onto the SAR image, and the prediction boxes are cropped to adjust the coordinates of the out-of-bounds prediction boxes to the boundaries of the SAR image. Finally, the target categories with low probabilities are removed, and non-maximum suppression processing is performed to suppress the redundant prediction boxes, and the SAR image with prediction boxes and classification prediction probabilities is obtained as the target detection prediction value.
8. The method for detecting aircraft targets in SAR images based on the joint attention mechanism according to claim 7, wherein: In step S2, according to the number of training iterations and the stochastic gradient descent algorithm, combined with the cross-entropy loss function and the SmoothL1 loss function, the prediction boxes and the classification prediction probabilities are jointly trained to complete the training of the target detection model.