Aircraft target detection method based on improved YOLOv5
By introducing the CA attention module, multi-stage jump feature fusion module and Swin Transformer module into the YOLOv5 network, the problem of the YOLO algorithm having a high missed error detection rate due to large differences in target scales in aircraft detection is solved, and higher detection accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202311491141.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-05-16
AI Technical Summary
The existing YOLO algorithm has a high missed detection and error detection rate due to large differences in target scales in aircraft small target detection tasks.
The CA attention module is introduced into the YOLOv5 network to enhance the receptive field and target positioning capabilities of the network model, and improve the neck network through the multi-level jump feature fusion module and the Swin Transformer module to improve the multi-scale characterization capability and detection accuracy of dense scenes.
It effectively reduces the missed detection and missed detection rates in aircraft target detection, improves detection accuracy and generalization capabilities, and meets the needs of real-time applications.
Smart Images

Figure CN120014484A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image target detection, and in particular to an aircraft target detection method based on improved YOLOv5. Background Art
[0002] With the rapid development of remote sensing technology, remote sensing images have shown the characteristics of high spectral resolution, high spatial resolution and high temporal resolution. High-resolution visible light remote sensing images have clear geometric structure, texture information and rich spatial information, which can detect and identify some important remote sensing targets such as aircraft, bridges, vehicles, ships, etc.
[0003] The target detection algorithm based on deep learning has strong feature abstraction and feature representation capabilities, and can use image data to automatically learn the features of the target. The target detection algorithm based on deep learning is mainly divided into two research directions: candidate region-based and regression-based methods. Due to the difference in design ideas, their network structures and training methods are very different. The two research methods have their own advantages in detection accuracy and detection speed. Algorithms based on candidate regions such as FPN, Fast R-CNN, Faster R-CNN, and Mask R-CNN have achieved very high detection accuracy in mainstream data test sets. The detection process includes two steps: candidate box extraction and candidate box classification and regression. Although the candidate region-based algorithm shares parameters, the training of these two parts is performed separately. Therefore, the algorithm model is large, the training takes a long time, and the hardware system requirements are high, which cannot meet the needs of real-time applications. The regression-based algorithm is represented by the SSD, RetinaNet, and YOLO series. Among them, SSD has a lower detection accuracy, and YOLO has achieved a good balance between accuracy and detection speed. Its main idea is to integrate the entire detection process into one network, generate fewer candidate boxes, and use a single network to estimate all parameters of the candidate boxes at the same time. Therefore, this method has a faster detection speed and can reach real-time level in some video applications, which can achieve more efficient detection efficiency than Faster R-CNN.
[0004] The quality of high-resolution optical images is affected by lighting conditions, flight attitude, etc., and aircraft targets are small and dense. The scales of different types of aircraft vary greatly, and their backgrounds are also relatively complex. These characteristics seriously affect the accuracy of aircraft target detection. At present, many researchers have made some improvements to the YOLO target detection algorithm. Dai et al. improved the network structure and multi-scale detection module of the YOLOv3 algorithm, but the model generalization ability is not strong and needs to rely on more data support. Xu, Hou et al. improved the backbone network of YOLO and used DenseNet to enhance the feature extraction ability of aircraft targets, but did not solve the problem of a large number of missed detections of small targets in the presence of a large number of occlusions. Wu Jie et al. proposed a lightweight multi-scale detection network, optimized the loss function and the global pooling layer, but the accuracy of aircraft target detection was limited. Summary of the invention
[0005] In order to address the problem that the existing YOLO algorithm has a high missed detection and false detection rate in the task of small aircraft target detection due to large target scale differences, the present invention provides an aircraft target detection method based on an improved YOLOv5.
[0006] The present invention provides an aircraft target detection method based on improved YOLOv5, comprising:
[0007] Step 1: Build an improved YOLOv5 network, specifically including: connecting a CA module after each C3 structure in the backbone network; or, connecting a CA module after the first C3 structure and before the second C3 structure in the bottom-up fusion process in the neck network; or, connecting a CA module before each detection head in the prediction network;
[0008] Step 2: training the improved YOLOv5 network to obtain an aircraft target detection model;
[0009] Step 3: Input the remote sensing image to be tested into the aircraft target detection model to obtain the detection result.
[0010] Furthermore, in step 1, it specifically includes: directly inputting the shallow feature map M2 output by the backbone network into the neck network to perform feature fusion with the remaining three feature maps.
[0011] Furthermore, a multi-level jump feature fusion network is used as the neck network; wherein, formula (4) is used to represent the fusion process of the multi-level jump feature fusion network on features of different levels:
[0012]
[0013] Among them, Q represents the feature map after feature fusion, X is the feature map directly downsampled from the original image, Y is the shallow high-resolution feature map in the feature pyramid, and Z is the feature map after upsampling the deep high-semantic feature map. It represents the attention weight matrix E calculated by the CA module after the X and Y feature maps are element-wise added.
[0014] Furthermore, the BottleNeck module in each C3 structure located between the neck network and the prediction network is replaced with a Swin Transformer module.
[0015] Furthermore, in step 1, the YOLOv5s model is selected as the basic network to construct an improved YOLOv5 network.
[0016] Beneficial effects of the present invention:
[0017] (1) To address the problem of large scale variations of aircraft targets, the CA attention module is integrated into the backbone network, neck network or prediction network to enhance the receptive field of the network model and the ability to accurately locate the target. Among them, the best effect is achieved by integrating CA attention into the backbone network.
[0018] (2) Aiming at the problem that the aircraft target is small in scale and the multi-scale representation ability is reduced during feature fusion, the original neck network is improved in two aspects: on the one hand, the shallow feature map M2 output by the backbone network is directly input into the neck network for feature fusion with the other three feature maps, which retains more feature information of small targets and provides more sufficient target information for feature fusion, which is conducive to the detection of small targets. Without adding an additional target detection layer, the target feature information can be enriched and the multi-scale expression ability of the network can be enhanced. On the other hand, a multi-level jump feature fusion module is proposed, which uses far-jump links to directly transmit the original bottom-level features to the deep semantic nodes, and the M2 from the backbone network is directly input into the neck network for feature fusion. 2 、M 3 The shallow feature map is connected to the convolution layer connected to the large-scale feature layer and the medium-scale feature layer. At the end of the network, the underlying features and semantic information are fused again, and the deep and shallow feature maps are connected with dynamic attention weights. The weights of feature layers of different scales are reasonably allocated, so that the model can better learn the underlying features and semantic information, and further improve the model's multi-scale learning capabilities.
[0019] (3) In order to address the common problem of dense distribution in aircraft detection tasks, the Swin Transformer module is integrated to allow the network to focus more on useful information in dense areas, enhance the model's ability to obtain global feature information of the target, and improve the network's generalization ability and detection accuracy in dense scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A schematic diagram of a flow chart of an aircraft target detection method based on an improved YOLOv5 provided in an embodiment of the present invention;
[0021] Figure 2 A distribution diagram of sample detection frames provided by an embodiment of the present invention;
[0022] Figure 3 Schematic diagram of the structure after the CA modules are integrated at different positions of the YOLOv5 network provided in an embodiment of the present invention: (a) CA_Backbone; (b) CA_Neck; (c) CA_Prediction;
[0023] Figure 4 An improved schematic diagram of shallow feature enhancement of a neck network provided by an embodiment of the present invention;
[0024] Figure 5 A schematic diagram of the structure of a multi-level jump feature fusion module provided in an embodiment of the present invention;
[0025] Figure 6 A schematic diagram of the structure of the Swin Transformer module provided in an embodiment of the present invention;
[0026] Figure 7 A schematic diagram of the structure of a C3STR module provided in an embodiment of the present invention;
[0027] Figure 8 Part of the image data provided by the embodiment of the present invention;
[0028] Fig. 9 A schematic diagram of the structure of the AT-YOLOv5s model provided in an embodiment of the present invention;
[0029] Fig.10 Comparison of the detection results of the YOLOv5 network provided in the embodiment of the present invention and the AT-YOLOv5s model of the present invention on the DOTA data set: the first row shows the detection results of the YOLOv5s network in four scenarios of dense, truncated, similar and repeated detection; the second row shows the detection results of the AT-YOLOv5s in four scenarios of dense, truncated, similar and repeated detection; wherein red is a false detection, green is a missed detection, and blue is a repeated detection;
[0030] Fig.11The detection results of the YOLOv5s network provided in an embodiment of the present invention and the AT-YOLOv5s model of the present invention on the RSOD dataset are compared: the first row represents the detection results of the YOLOv5 network in four scenarios: complex background, dense distribution, large scale difference, and blurred image; the second row represents the detection results of the AT-YOLOv5s network in four scenarios: complex background, dense distribution, large scale difference, and blurred image. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0032] Example 1
[0033] The YOLO series of target detection algorithms has now developed to YOLOv5, which combines the advantages of previous algorithms and is superior in detection accuracy and speed. It is currently a target detection model with strong versatility and high performance. The YOLOv5 network structure is divided into four components: input, Backbone, Neck, and Prediction. YOLOv5 uses Mosaic data enhancement in the input part; Backbone mainly includes SPPF structure and C3 structure; Neck uses PANet (Path Aggregation Network) structure connection; the loss function of the bounding box in Prediction is GIoU (Generalized IoU) loss.
[0034] The YOLOv5 version has four basic network models of different sizes: YOLOv5s, YOLOv5m, YOLOv51 and YOLOv5x. Among them, the YOLOv5s network model is the network with the shallowest depth and the smallest feature map width. The three models of YOLOv5m, YOLOv51 and YOLOv5x are the products of continuous deepening and widening based on YOLOv5s. As an implementable method, considering the real-time performance of target detection, this embodiment selects the YOLOv5s model as the basic network for target detection. It is understandable that in scenarios where real-time detection is not required, the other three models can also be selected as the basic network for target detection.
[0035] like Figure 1As shown, an aircraft target detection method based on improved YOLOv5 provided by an embodiment of the present invention comprises the following steps:
[0036] S101: Build an improved YOLOv5 network;
[0037] Specifically, in the aircraft target detection task, the scales of different aircraft targets vary greatly, which requires the constructed aircraft target detection model to have a large receptive field, but at the same time have the ability to accurately locate the target, so that it can detect large-scale aircraft targets while not missing or misdetecting small-scale aircraft targets. Based on this situation, this embodiment integrates the CA module into the YOLOv5 network.
[0038] Placing the CA module at different positions in the network model will have different effects on the detection results. In this embodiment, three attention fusion methods at different positions are designed, that is, CA attention is fused in BackBone, Neck and Prediction modules respectively. The reason why the Input module is not considered is that it only processes data and does not involve target feature extraction and fusion operations. The three model structures generated are recorded as CA_Backbone, CA_Neck and CA_Prediction respectively; Figure 3 As shown in the figure, CA_Backbone connects a CA module after each C3 structure in the backbone network; CA_Neck connects a CA module after the first C3 structure and before the second C3 structure in the bottom-up fusion process in the neck network; CA_Prediction connects a CA module before each detection head in the prediction network.
[0039] The CA (CoordAttention) module introduced in this embodiment is a novel lightweight network attention mechanism proposed by Qibin Hou et al. in CVPR2021, which embeds position information into channel attention. Different from the use of global pooling in channel attention to compress spatial information into channel descriptors, which makes it difficult to retain position information, CA decomposes global pooling into two one-dimensional feature encoding processes in different directions in order to capture long-range dependencies and retain accurate position information at the same time, so that the generated attention weight map contains spatial coordinate information. The specific implementation is divided into two steps: coordinate information embedding and coordinate attention generation. In the coordinate information embedding part, the input feature Figure XUse the pooling kernels of (H, 1) and (1, W) respectively, and perform adaptive average pooling in the horizontal and vertical directions at the same time, so as to generate two feature maps with independent direction perception. When the channel is c, they are expressed as shown in the following equations (1) and (2). In the coordinate attention generation part, a feature map with both position and channel information is generated. The two feature maps in the horizontal and vertical directions are spliced in the third dimension, and then a convolution operation is performed with a weight-sharing 1×1 convolution kernel. The process feature map is generated through regularization and nonlinear activation. The generated feature map is then split into two feature vectors, and transformed using two 1×1 convolution kernels to obtain a feature vector of the same size as the input dimension. Finally, the attention weight map in two directions is obtained through the Sigmoid (x) activation function, which can capture the long-range dependency of the input feature map along a specific spatial direction. Finally, the input feature Figure X Multiplying with the two attention weight maps respectively, the output of the CA module is expressed as follows (3).
[0040]
[0041]
[0042] y c (i,j)=x c (i,j)×g c h (i)×g c w (j)(3)
[0043] The CA module simultaneously retains the information of different channels and the spatial position encoding information, enhances the representation ability of the feature map, and helps the network model accurately locate the position of the target of interest.
[0044] It should be noted that, unlike other network structures that introduce attention mechanisms into the backbone network, this embodiment connects CA after the C3 structure (rather than after other structures in the backbone network) to further enhance the receptive field of the network model and its ability to accurately locate the target. Figure 2 is a distribution diagram of sample detection frames after the method of this embodiment is used on the DOTA dataset, as shown in Figure 2 As shown in the figure, it can be seen that the scale of the sample detection box is unevenly distributed. Therefore, the introduction of the CA module can enhance the receptive field of the network model and its ability to accurately locate the target, thereby solving the problem of false detection and missed detection caused by large changes in the scale of the aircraft target to a certain extent.
[0045] S102: training the improved YOLOv5 network to obtain an aircraft target detection model;
[0046] S103: Inputting the remote sensing image to be detected into the aircraft target detection model to obtain a detection result.
[0047] Example 2
[0048] It is understandable that for the target detection algorithm based on the Convolutional Neural Network (CNN), the feature map after deep convolution has a small scale and a large receptive field, and has richer feature semantic information, but the feature information of the target will be lost in multiple convolution operations, which is not conducive to the detection of small targets; while the feature map of shallow convolution has a large scale and a small receptive field, and although it lacks global semantic information, it can obtain more feature information of small targets. The original YOLOv5 network uses 8x, 16x, and 32x downsampled feature maps P 3 , P 4 , P 5, To predict small, medium and large-sized targets respectively. However, the aircraft targets in remote sensing images are small in scale. After multiple downsampling operations, the target feature information is lost too much, and the original three-scale feature maps are difficult to detect such small targets. Figure 4 As shown in the figure, {M2, M3, M4, M5} are feature maps of four scales, corresponding to downsampling of 4, 8, 16, and 32 times, respectively. If the input image size is 416×416 pixels, the sizes of the four scale feature maps correspond to {104×104, 52×52, 26×26, 13×13} pixels, and most of the aircraft targets in the image occupy a small proportion of the entire image, so their size in the deep feature map may be only a few pixels, which may cause missed detection or wrong detection of aircraft targets.
[0049] Therefore, in order to further reduce the missed detection and false detection rate of extremely small aircraft targets, based on the above embodiment, the neck network of the YOLOv5 network is further improved in the embodiment of the present invention to enhance the model's feature extraction capability for small targets. In this embodiment, the neck network is mainly improved in the following two aspects.
[0050] The first aspect of improvement: Figure 4 As shown in the red line part, the feature map M is obtained by downsampling by 4 times 2 It is directly input into the neck network for feature fusion with the other three scale feature maps M3, M4, and M5.
[0051] Specifically, M 2 The feature map has a higher resolution and can retain more feature information of small targets, providing more sufficient target information for feature fusion, which is conducive to the detection of small targets. Without adding an additional target detection layer, the target feature information can be enriched and the multi-scale expression ability of the network can be enhanced.
[0052] The second improvement: a multi-level jump feature fusion module is proposed. Figure 5 As shown in the red line part in the figure, the original bottom-level features are directly transmitted to the deep semantic nodes using far-jump links. 2 、M 3 The shallow feature map is connected to the convolution layer connected to the large-scale feature layer and the medium-scale feature layer, and the bottom-level features and semantic information are fused again at the end of the network. At the same time, a weighted method is used to reflect the contribution of features of different scales to the fused features. Attention weights are assigned to feature maps of different levels at the spatial and channel levels. Different from the original fixed feature weights, the parameters are continuously updated and optimized during the feature learning process. The fusion process of features of different levels by the multi-level jump feature fusion module can be expressed by formula (4):
[0053]
[0054] Among them, Q represents the feature map after feature fusion, X is the feature map directly downsampled from the original image, Y is the shallow high-resolution feature map in the feature pyramid, and Z is the feature map after upsampling the deep high-semantic feature map. It represents the attention weight matrix E obtained by CA attention calculation after element-wise add of X and Y feature layers. The weight E is processed by Sigmoid function and its value range is 0 to 1.
[0055] Specifically, the YOLOv5 network obtains feature maps of different scales through a multi-scale feature extraction module, in which the shallow feature map has a high resolution and accurate target location information, while the deep feature map has a low resolution and rich semantic information. Therefore, fusing information from feature maps of different scales will be more conducive to target detection and recognition. However, the original YOLOv5 uses the PANet feature fusion network as the neck network, which adopts a top-down and bottom-up fusion method to connect different feature layers with an equal relationship, ignoring shallow features, and small-scale information is not fully utilized. To address this problem, the multi-level jump feature fusion module in this embodiment, when the input image data contains a large number of small target samples, the attention weight parameter will give higher weights to the shallow features. Figure X , Y, while the deep feature map Z is given a lower weight, so that the network can reasonably allocate weights between shallow features and deep features. Compared with the original feature fusion network of YOLOv5, the multi-level feature jump module uses far jump links to increase the underlying features of the original image, and uses dynamic attention weights to fuse more small target features, which can improve the detection accuracy of aircraft targets and achieve more accurate recognition.
[0056] The aircraft target detection method provided by the embodiment of the present invention combines deep semantic information with shallow features, and reasonably distributes the weights of feature layers of different scales, so that the model can better learn the underlying features and semantic information, increase the network's learning ability for small targets, and improve the network model's ability to locate and regress targets.
[0057] Example 3
[0058] It is understandable that CNN enhances the ability to extract image features and reduce the number of network parameters through local connections and parameter sharing methods. It has great advantages in extracting low-level image features and visual structures, but it has limitations in acquiring global feature information of images. Therefore, in order to enhance the model's ability to acquire global feature information of the target; at the same time, in order to improve the generalization ability and detection accuracy of the network in dense scenes, on the basis of the above embodiments, the neck network of the YOLOv5 network is further improved in the embodiments of the present invention, specifically: the original BottleNeck module in each C3 structure between the neck network and the prediction network is replaced with a Swin Transformer module.
[0059] Specifically, because the Prediction module of the YOLOv5s network predicts the target through three feature maps of different scales, it mainly predicts large targets on the deep feature map and predicts small targets on the shallow feature map, so adding the Swin Transformer module before prediction can enhance the model's ability to focus on feature information at different local locations. Figure 7 As shown, at the position directly connected to the three detection heads in the network, a Swin Transformer module is used to replace the original BottleNeck module in the C3 module with a Swin Transformer Block structure, and the formed new module is recorded as C3STR.
[0060] The SwinTransformer model is an improved model proposed by Microsoft in 2021 based on the Transformer. It has two great advantages. First, it uses the network layering method to select different sampling values to reduce the network calculation amount. Second, it uses the sliding window method to realize the cross-window connection of local features, so that the network model can obtain the information features of adjacent windows and increase the network receptive field. The SwinTransformer module is composed of residual connections of units such as the normalization layer (LayerNorm, LN), the multi-layer perceptron (Multilayer Perceptron, MLP), and the sliding window multi-head self-attention mechanism (Multi-head SelfAttention, MSA). The specific structure is as follows Figure 6As shown. The SwinTransformer module has two Transformers. The first module W-MSA part uses a regular window division strategy from the upper left corner pixel to evenly divide the 8×8 feature map into 2×2 windows of size 4×4 (M=4). The second module SW-MSA part uses a different window mechanism from the previous layer, namely the sliding window method. Compared with the regular division of the window position, the window position is shifted by (M / 2, M / 2) pixels, and then 3×3 non-overlapping windows are obtained. The sliding window division method introduces connections between adjacent non-overlapping windows in the previous layer, which greatly increases the receptive field range of the network.
[0061] The aircraft target detection method provided by the embodiment of the present invention uses the SwinTransformer model at the position connected to the three detection heads in the YOLOv5 basic network, so that the network model can better detect aircraft targets.
[0062] In order to verify the effectiveness of the target detection method provided by the embodiment of the present invention, the present invention also provides the following experiments.
[0063] 1. Dataset analysis and processing
[0064] The DOTA dataset is a commonly used dataset for remote sensing image target detection. The images in the dataset come from different sensors and platforms, such as Google Earth images, my country's high-resolution satellite series, and the Resource Satellite Data Application Center. In addition, the dataset has diverse resolutions, with image sizes ranging from 800*800 to 4000*4000. The DOTA dataset contains 2806 images in 15 categories, including 202 aircraft targets. However, due to the different sizes of images in the dataset, excessive image resolution affects the detection accuracy of small aircraft targets. Therefore, the images in the dataset are cut, the image size is set to 1024*1024, the overlap of adjacent slices is 200, and images that do not contain aircraft targets are removed. Finally, the training set contains 2053 images and the test set contains 1478 images. Some images in the dataset are as follows Figure 8 shown.
[0065] (II) Experimental environment and settings
[0066] A deep learning framework based on Windows 10 operating system, Python 3.8 and PyTorch 1.7.0 (such as Fig. 9As shown, referred to as AT-YOLOv5s model; wherein, the red dotted line represents the far jump link, and the black dotted line represents the improved fusion process of the feature enhancement module of the present invention); the experimental hardware environment is Intel(R) Xeon(R), 32GB memory, NVIDIA RTX3080 image processor. The experimental training parameters in this paper are: batch size (training samples per batch) is set to 8, epoch (training iteration) is set to 150, learning rate is set to 0.0001, and momentum parameter is set to 0.5; Adam gradient descent is used, and the learning rate is preheated using the Warmup strategy training when the model starts training. 3 enpoches are set for preheating. After the end, the cosine annealing learning algorithm is used to update the learning rate. This paper adopts AP 50 ,AP 75 ,AP 50-95 (average precision), P (Precision), R (recall), Params (Parameters), FPS (frames per second) and other indicators are used as evaluation indicators of model performance. 50 、AP 75 They represent the average detection accuracy of the target when the IoU value is 0.5 and 0.75 respectively. AP50:95 represents the average detection accuracy of the target under all thresholds with IoU values ranging from 0.5 to 0.95 (with a step size of 0.05). Params is the parameter quantity of the model, which is used to measure the consumption of computer memory resources. FPS refers to how many images the network model can detect per second, which is used to measure the real-time performance of the model.
[0067] (III) Analysis of experimental results
[0068] In order to verify the effectiveness of the AT-YOLOv5 model in the aircraft target detection task, this paper set up 5 comparative experiments. The first experiment is to verify the improvement effect of the CA module on the network model, and compare the fusion effect of CA attention at different positions of YOLOv5. The second experiment is to verify the effectiveness of the feature enhancement module improved in this paper, and compare the ordinary four-detection head structure with the improved three-detection head structure proposed in this paper. The third experiment is to test the effectiveness of the embedded Swin Transformer module on the network model. The fourth experiment is an ablation experiment to test the effectiveness of the superposition of various improved modules on the performance of the network model. The fifth experiment compares the performance of the algorithm in this paper with the mainstream algorithm in the aircraft detection task to prove the performance advantage of the AT-YOLOv5s algorithm. Train each group of models in the 5 experiments respectively, and use the precision P, recall R, and average precision AP 50 ,AP 50-95 and AP 75As a measurement indicator, the YOLOv5s network accuracy is used as the performance measurement benchmark.
[0069] (1) Comparative Experiment on Attention Fusion Design
[0070] This paper integrates the design of CA attention to enhance the receptive field of the network model. From the results in Table 1, it can be seen that after adding the CA module, the detection precision P, recall R and AP value of the network model have been improved. This is because after adding CA attention, the target position information and semantic information obtained by the network feature map have been increased. Among them, the CA_Backbone model with improved attention at the Backbone position has the highest detection accuracy and AP 50 Compared with the YOLOv5s model, it has been improved by 2.6% to 89.3%; while the accuracy of the attention fusion module at the Neck and Prediction positions has been slightly improved. Because in the Backbone network module, although the extracted features include less semantic information, there is more position and shape contour information. Therefore, after adding CA attention, the feature information acquired by the target can be further enriched through the feature fusion of channel and spatial dimensions. At the Neck intermediate layer and the Prediction output end, the feature map has undergone multiple up- and down-sampling operations, and the feature information of the target is lost. Although the CA attention mechanism can also enhance the feature representation capability, it has limited improvement on the detection results.
[0071] Table 1 Comparative experimental results of attention fusion module
[0072]
[0073]
[0074] (2) Experimental verification of feature enhancement module
[0075] The feature enhancement module includes two parts: feature extraction and feature fusion. First, in the feature extraction module, through the comparative experiments in Table 2, we can see that compared with the YOLOv5s basic network, the accuracy of the two improved multi-scale feature extraction modules has been improved, but the shallow feature enhancement method has obviously better accuracy and shorter detection time. Compared with the four-detection head model, the accuracy is improved by 0.9%, and the inference time of each picture is also improved by 0.5ms. Because the shallow features are added to participate in the fusion, only the convolution layer of this level is added, and the network model only adds very few parameters. The additional detection head method requires 4 times the sampling of the feature map P 2 The shallow prediction layer may also cause language ambiguity. This shows that shallow feature enhancement is helpful for improving the detection effect of the network.
[0076] Table 2 Comparative experimental results of multi-scale feature extraction models
[0077]
[0078] In order to further enhance the multi-scale representation capability of the network, the feature fusion module is improved on the basis of shallow feature enhancement. The downsampled feature map of the input image is directly connected to the optimized large-scale and medium-scale feature layers by long jumps, and the attention model is used to dynamically assign weights. In this way, not only the feature information of the image itself can be fused, but also the deep and shallow feature maps can be weighted fused according to the contribution of the feature map to the fused features. As can be seen from Table 3, through feature enhancement processing, the model detection accuracy AP 50 Reaching 89.4%, an improvement of 2.7% over the YOLOv5s baseline.
[0079] Table 3 Verification experimental results of feature enhancement module
[0080]
[0081] (3) Swin Transformer verification experiment
[0082] This paper integrates the Swin Transformer module before target prediction, which enhances the network model's ability to model global feature information, obtains rich contextual information, and improves the network detection accuracy. From Table 4, we can see that after the introduction of Swin Transformer, the model can detect aircraft targets well, with an accuracy of 93.5% and a recall rate of 84.8%. The AP50 is 2.0% higher than the YOLOv5s baseline network, which verifies the effectiveness of the Swin Transformer module in the network.
[0083] Table 4 Experimental results of Swin Transformer module
[0084]
[0085] (4) Ablation experiment
[0086] Ablation experiments were conducted to verify the impact of the mixed use of improved modules on the performance of the network model. As shown in Table 5, the following four groups of models were trained respectively, and the accuracy of the original YOLOv5s was used as the performance benchmark. √ indicates the use of improved modules, and - indicates no use. From the results in the table, it can be seen that after each improved module is superimposed, each indicator will be improved to a certain extent. Moreover, the effect of improving all three together is the best, AP 50 ,AP 50-95Both achieved the optimal results, which were 90.6% and 65.6% respectively, which were 3.9% and 1.9% higher than the benchmark, further verifying the feasibility of the improved algorithm.
[0087] Table 5 Ablation test results
[0088]
[0089] (5) Comparative experimental analysis
[0090] In order to further verify the detection performance of the AT-YOLOv5s model, three target detection models were selected for comparative experiments on the DOTA aircraft dataset. The detection results are shown in Table 6. As can be seen from the table, the improved model has a lower detection speed due to the increase in the number of detected targets and the increase in model parameters. It transmits 18 frames less per second than the YOLOv5s model, but is still significantly faster than algorithms such as YOLOv3 and YOLOv4. 50 ,AP 75 ,AP 50-95 The detection accuracy indicators are significantly better than other algorithms. The YOLOv3 model uses Darknet53 as the backbone network and uses the FPN network for connection. The detection accuracy AP 50 The detection rate is above 80%, but due to the multiple downsampling operations, the feature information of small targets is lost a lot, so the detection effect is not as good as that of YOLOv4 and YOLOv5s models. The AT-YOLOv5 model proposed in this paper introduces the attention mechanism in the backbone and the SwinTransformer mechanism in the prediction, focusing more attention on the dense small target area, enhancing the small target feature extraction ability, and adding shallow feature maps to the neck feature fusion network. The attention mechanism is used for multi-scale weighted jump connections to improve the feature map fusion ability of small targets. Therefore, the detection performance of the improved algorithm in this paper is significantly better than the above mainstream models, and effectively improves the recognition accuracy of small aircraft targets.
[0091] Table 6 Detection results of different small target detection models on the DOTA dataset
[0092]
[0093] (IV) Visualization of experimental results
[0094] This paper uses CA attention fusion to build a new backbone network, proposes a feature enhancement module, and applies the Swin Transformer module. The effectiveness of the algorithm is verified in experiments. Fig.10The detection effect of the algorithm is shown in Figure 2. Some detection results in the DOTA aircraft dataset are visualized, and the results in different scenarios are compared. It can be seen that the original YOLOv5s network missed some dense aircraft, while the improved network AT-YOLOv5s proposed in this paper can detect these targets, including smaller-scale aircraft targets and truncated targets in dense scenes; at the same time, it reduces the false detection phenomenon. For example, in similar scenes, the original network mistakenly marked 4 T-apron signs as aircraft targets, and under the influence of complex wing and tail shadows, the original network identified one aircraft target as two, while the improved network model proposed in this paper solves the problems of false detection and duplication.
[0095] (V) Migration experiment
[0096] In order to fully verify the effectiveness of the AT-YOLOv5s algorithm, this paper conducted a comparative verification on the RSOD remote sensing dataset. The RSOD dataset was released by Wuhan University in 2015. The images in the dataset are mainly from Google Earth and Tiandi Map, and are acquired under different weather, seasons and imaging conditions. Among them, the aircraft category targets have large scale differences, diverse within the class, and complex backgrounds, which are suitable for migration experiments to verify the effectiveness of this algorithm. There are 446 images of aircraft types in the dataset, a total of 4993 aircraft targets, 334 training sets, and 112 test sets. The training parameter settings are the same as other comparative experiments. On the RSOD remote sensing dataset, compared with the YOLOv5s baseline model AP 50 The accuracy is 95.9%, and the detection accuracy AP of the improved model 50 It increased by 0.8% to 96.7%. Fig.11 From some detection examples in , we can see that this model outperforms the YOLOv5s model in terms of accuracy, false alarm rate, repeated detection, and dense target detection. Experiments show that the algorithm in this paper has strong generalization ability and can achieve good robustness for datasets with small target size, large scale difference, and high target overlap.
[0097] In view of the problem that the scale of aircraft targets in the DOTA data set varies greatly and is densely distributed, the present invention introduces an attention mechanism to enhance the receptive field and improve the network model's ability to accurately locate aircraft targets. In order to improve the network model's multi-feature representation capability and generalization capability in dense scenes, a shallow feature enhancement and multi-level jump feature fusion module are proposed. At the same time, the SwinTransformer module is used to enhance the model's ability to capture local information and improve network detection accuracy. Although the FPS of the AT-YOLOv5s algorithm of the embodiment of the present invention is lower than that of the YOLOv5s algorithm, it can still meet the real-time requirements of the aircraft target detection task. Subsequent research will be conducted from the perspectives of network pruning and distillation to improve the detection rate of the network model.
[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An aircraft target detection method based on improved YOLOv5, characterized in that: include: Step 1: Build an improved YOLOv5 network, specifically including: connecting a CA module after each C3 structure in the backbone network; or, connecting a CA module after the first C3 structure and before the second C3 structure in the bottom-up fusion process in the neck network; or, connecting a CA module before each detection head in the prediction network; Step 2: training the improved YOLOv5 network to obtain an aircraft target detection model; Step 3: Input the remote sensing image to be tested into the aircraft target detection model to obtain the detection result.
2. The aircraft target detection method based on improved YOLOv5 according to claim 1, characterized in that: In step 1, specifically, it includes: directly inputting the shallow feature map M2 output by the backbone network into the neck network to perform feature fusion with the other three feature maps.
3. The aircraft target detection method based on improved YOLOv5 according to claim 2, characterized in that: A multi-level jump feature fusion network is used as the neck network; wherein, formula (4) is used to represent the fusion process of the multi-level jump feature fusion network on features of different levels: Among them, Q represents the feature map after feature fusion, X is the feature map directly downsampled from the original image, Y is the shallow high-resolution feature map in the feature pyramid, and Z is the feature map after upsampling the deep high-semantic feature map. It represents the attention weight matrix E calculated by the CA module after the X and Y feature maps are element-wise added.
4. The aircraft target detection method based on improved YOLOv5 according to claim 1, characterized in that: In step 1, the BottleNeck module in each C3 structure between the neck network and the prediction network is replaced with a SwinTransformer module.
5. The aircraft target detection method based on improved YOLOv5 according to claim 1, characterized in that: In step 1, the YOLOv5s model is selected as the basic network to build an improved YOLOv5 network.