Unmanned aerial vehicle aerial image small target detection method based on improved YOLOv8
By introducing a MAFPN structure into the Neck part of the YOLOv8 network, adding a small target detection head, and replacing the detection head with a Dyhead detection head, the problem of low target detection accuracy in aerial images of the drone is solved, and higher detection accuracy and fewer missed and missed detection are achieved.
Patent Information
- Application Number
- CN202510189745.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-06
AI Technical Summary
In the aerial images of drones, the target scale changes greatly, the background is complex, and it is susceptible to light factors, resulting in low target detection accuracy and many missed and missed detections.
The MAFPN structure was introduced in the Neck part of the YOLOv8 network, a small object detection head with a size of 160×160 was added, and the detection head was replaced with a Dyhead detection head with a variety of attention mechanisms.
It effectively improves the detection accuracy of small target objects in aerial images of drones, reduces missed and misdetections, and improves the detection performance of the model on the VisDrone2019 dataset.
Smart Images

Figure CN120107830A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and deep learning target detection, and specifically relates to a method for detecting small targets in unmanned aerial vehicle aerial images based on improved YOLOv8. Background Art
[0002] As an important part of the drone industry, drone aerial photography has played a significant role in many fields in recent years, such as smart agriculture, power inspection, disaster assessment, military field, and intelligent transportation. However, compared with conventional target detection tasks, drone aerial images have problems such as large target scale changes, complex backgrounds, and susceptibility to lighting factors, which makes it easy for drone aerial images to miss detection and misdetection, and the detection is difficult. Traditional target detection algorithms often do not work well in drone aerial photography scenarios. Therefore, studying drone aerial small target detection algorithms is of great significance to improving the target detection accuracy of drone aerial images and improving the processing capabilities of drone aerial images.
[0003] The development of deep learning detection algorithms can be mainly divided into two-stage algorithms represented by Faster-R-CNN and Cascade r-cnn, and single-stage algorithms represented by RetinaNet, SSD, TOOD and YOLO series. Two-stage detection algorithms such as the classic Faster-R-CNN network use the RPN structure to generate candidate boxes, and then project the candidate boxes onto the feature map it generates. This results in a large number of model parameters and a slow detection speed. Therefore, it is difficult to deploy on terminal devices that require real-time performance and a small number of model parameters. As a representative of single-stage algorithms, the YOLO series algorithm has undergone multiple iterations and is widely used in drone aerial photography for small target detection due to its fast detection speed, high accuracy, small number of model parameters, and easy terminal deployment. Summary of the invention
[0004] In order to solve the problem of low detection accuracy of UAV aerial images, the present invention proposes a small target detection method for UAV aerial images based on improved YOLOv8, which effectively improves the detection accuracy of small target objects in UAV aerial images and reduces the cases of missed detection and false detection.
[0005] This paper takes YOLOv8s as the benchmark model and makes three related improvements, mainly including: introducing the MAFPN structure in MAF-YOLO in the Neck part; adding a small target detection head of 160×160 size to the network; and introducing the Dyhead detection head with multiple attention mechanisms.
[0006] The specific scheme is as follows: the MAFPN structure in MAF-YOLO is introduced in the Neck part to integrate and fuse the multi-scale information of different resolution layers: because the PAFPN structure cannot simultaneously and efficiently and adaptively fuse high-level semantic information and low-level spatial information, lacks information exchange at different depth levels, and for small target objects with characteristics such as low resolution, contrast and blur, the network structure is often difficult to effectively fuse and extract the features of shallow target objects. Based on this, the present invention introduces the MAFPN structure in MAF-YOLO in the Neck part of the network. The module mainly uses two paths to effectively extract and fuse the effective information of features between different depth levels of the network, mainly including: in the first bottom-up path, the network extracts multi-scale features from Backbone and performs preliminary auxiliary fusion in the shallow layer of Neck; then, the network collects information of each layer through dense connections in the second top-down path, and guides the detection head Head to obtain diversified output information at different resolutions, thereby realizing rich feature interaction and fusion of the network, thereby enhancing the network's representation ability for multi-scale objects.
[0007] Add a 160×160-sized small target detection head to the network detection head: After the feature fusion part, the YOLOv8 network mainly outputs three feature layers of different scales, namely: 80×80 small target detection layer, 40×40 medium target detection layer and 20×20 large target detection layer. The smaller the downsampling multiple of the original input image, the larger the resolution of the corresponding feature map obtained, and the more detailed features of the corresponding target object retained. In order to improve the network's detection ability for small target objects, the present invention downsamples and upsamples the outputs of the P1 layer and the P3 layer on the basis of the original YOLOv8 network, and then concats the obtained upsampling and downsampling results with the P2 layer to finally obtain a 160×160-sized small target detection head, so that the network contains a total of 160×160-sized small target detection layer, 80×80 small target detection layer, 40×40 medium target detection layer and 20×20 large target detection layer. Four detection layers, so more shallow target feature information can be retained.
[0008] Introducing the Dyhead detection head with multiple attention mechanisms: The original YOLOv8 network detection head has problems such as lack of contextual information connection, limited expression ability, and weak processing ability for multi-scale target objects. When faced with small target objects that are easily interfered or occluded by the background, have low resolution, low contrast, and small size, it is usually difficult to accurately detect and locate. In order to improve the problem of difficult positioning and detection of small target objects, the original YOLOv8 detection head is replaced with the Dynamic Head framework with attention mechanism proposed by Microsoft. This detection head embeds scale perception, spatial position perception, and task perception into a target detection head at the same time, and uses the attention mechanism to unify different target detection heads, which can significantly improve the network detection head's expression ability for objects of different scales without significantly increasing the amount of calculation.
[0009] Furthermore, the Dyhead detection head can be expressed as a combination of three types of sequence attention, and each type of attention focuses on only one dimension. Its calculation formula is: W(F) = π C (π S (π L (F)·F)·F)·F, where F is the input three-dimensional tensor of L×S×C, S=H×W is the size of the feature map, and π L (F), π s (F) and π C (F) represent scale-aware attention, spatial position-aware attention, and task-aware attention, respectively.
[0010] The improved YOLOv8 network is used as the target detection model, and the training set in the VisDrone2019 dataset is used as the network input to train the network. It is iterated for 100 rounds until the model converges, thus obtaining the final improved YOLOv8 drone aerial image target detection model.
[0011] Furthermore, after obtaining the trained network weights, the test set images in the VisDrone2019 dataset are input, and the network will output the target detection box and confidence score of the image to be detected, and obtain the actual detection effect of the drone aerial image.
[0012] In summary, the present invention first introduces the MAFPN structure in MAF-YOLO in the Neck part of the network to improve the PAFPN structure of the benchmark model YOLOv8s (i.e., the original YOLOv8 network). This module can integrate and fuse multi-scale information of different resolution layers. Secondly, in order to retain more shallow network feature information, a small target detection head that can retain more high-resolution feature information is designed. Then, in the detection head part, the original YOLOv8 detection head is replaced with the Dyhead dynamic detection head with scale perception, spatial position perception, and task perception. The above improvements effectively improve the detection accuracy of the model on the VisDrone2019 dataset. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 A flowchart of a method for detecting small targets in drone aerial photography based on improved YOLOv8 in the present invention;
[0014] Figure 2 It is a structural diagram of the improved YOLOv8 network based on the baseline model YOLOv8s in the present invention;
[0015] Figure 3 It is a module structure diagram of MAFPN in the present invention;
[0016] Figure 4 It is the module structure diagram of Dyhead in the present invention;
[0017] Figure 5 This is a diagram of the structure of multiple DyHead modules in series in the present invention.
[0018] Figure 6 This is a comparison chart of model indicators before and after improvement in the present invention;
[0019] Figure 7 This is a visual comparison chart of the model detection effects before and after the improvement in the present invention. DETAILED DESCRIPTION
[0020] In order to more clearly explain the purpose, technical solution and innovation of the present invention, the present invention will be described in detail below with reference to the accompanying drawings.
[0021] The present invention discloses a method for detecting small targets in drone aerial photography based on an improved YOLOv8 model. The process is as follows: Figure 1 As shown, the specific implementation steps are as follows:
[0022] Step S1: Get the VisDrone2019 dataset, divide it into training set, validation set and test set, and convert it into YOLO data format;
[0023] Step S2: Using the original YOLOv8 network as the benchmark model, an improved YOLOv8 UAV aerial photography small target detection network model is constructed. The improvement is mainly to make relevant improvements to the Neck part and detection head of the original YOLOv8 network;
[0024] Step S3: Input the images of the training set and the validation set into the improved YOLOv8 UAV aerial photography small target detection network model to train and verify the image data respectively, and obtain the final target detection model;
[0025] Step S4: Based on the improved target detection model, the drone aerial image is detected to obtain the detection result.
[0026] In this embodiment, the above step S1 is specifically as follows:
[0027] Download the image data from the official website of the VisDrone2019 dataset. The dataset has a total of 8599 images, of which 6471, 548, and 1610 are training, validation, and test images, respectively. It includes ten common traffic scene targets such as pedestrian, bicycle, people, motor, and bus, and covers different scenes such as dense, occluded, and sparse. Then use the Python script to convert it into the txt format of the YOLO sequence for the next step of training.
[0028] In this embodiment, the above step S2 is specifically as follows:
[0029] Taking the original YOLOv8 network as the benchmark model, the MAFPN structure in MAF-YOLO is introduced in the Neck part to integrate and fuse the multi-scale information of different resolution layers of the network, add a 160×160 small target detection head to the network, and introduce a Dyhead detection head with multiple attention mechanisms. The improved YOLOv8 model is as follows: Figure 2 shown.
[0030] Specific improvements include:
[0031] The MAFPN structure is introduced to effectively extract and fuse the effective information of features between different depth levels of the network, so as to achieve richer feature interaction and fusion of the network, such as Figure 3 shown.
[0032] A small target detection head of 160×160 size is added to the network, which effectively retains the high-resolution information of the shallow network and improves the network's ability to detect small target objects.
[0033] The detection head of the original network is replaced with the Dyhead dynamic detection head that can integrate multiple attention mechanisms. This detection head can significantly improve the network's ability to express objects of different scales without significantly increasing the amount of computation.
[0034] The above step S3 is specifically as follows:
[0035] In the specific implementation of the present invention, the hardware configuration of the platform for the experiment is: the operating system is Ubuntu 20.04, the CPU uses an Intel (R) Core (TM) i7-13700KF processor, the GPU is NVIDIA GeForce RTX 4080Super, 16GB video memory, the deep learning framework is Pytorch 2.0.0, the Cuda version is 11.7, and the Python version is 3.10.15. In order to ensure the accuracy and fairness of the experiment, the present invention did not load the pre-training weights during the experiment, the bachsize of the experiment was set to 4, the iteration round was 100 epochs, the initial learning rate was 0.01, and the AdamW optimizer was used.
[0036] The main index parameters in target detection are precision, recall, average AP of a single category, mAP, Parameters, GFLOPs and FPS, etc. The mAP50 and mAP50:90 parameters represent the average results of mAP when the IoU between the predicted box and the true box is 50% and within the range of 50%-95% IoU threshold, respectively. Parameters and GFLOPS parameters are indicators for measuring the size of model parameters and the amount of model calculation, respectively. The present invention uses parameters such as mAP50, Parameters and GFLOPs to evaluate model performance.
[0037] Specifically, the DyHead module structure and the multi-DyHead module series structure diagram are as follows: Figure 4 , Figure 5 As shown in Table 1, in order to study the effect of adding different Dyhead modules on the VisDrone2019 dataset, the present invention conducts a network performance comparison experiment with different numbers of Dyhead modules in series, as shown in Table 1. The results show that on the VisDrone2019 dataset, with the increase in the number of Dyhead modules in series, the mAP50 index of the network first increases and then decreases, but the increase is not large, and the corresponding calculation amount and parameter amount parameters of each Dyhead module network in series increase by about 0.5GFLOPS and 0.5M respectively. Considering factors such as the calculation amount and accuracy of different Dyhead modules in series, the present invention only uses one Dyhead detection head.
[0038] Table 1 Performance comparison of different Dyhead modules in series
[0039]
[0040] In order to verify the effect of adding different modules on the experiment, the present invention sets up 8 groups of experiments on the VisDrone2019 dataset to conduct ablation experiments between modules. Given the network with the same experimental parameters such as batchsize and epoch, the small target detection layer, MAFPN multi-scale feature fusion module and Dyhead dynamic detection head are added to the YOLOv8s benchmark model to verify the effectiveness of different modules for improving the model through different combinations. The "√" in the table represents the use of the module. The results are summarized as shown in Table 2.
[0041] Table 2 Ablation experiments between different modules
[0042]
[0043] As can be seen from Table 1, in Experiment 2, after the Neck part of the PAN-FPN structure of the original YOLOv8s network was replaced with the MAFPN structure that can effectively fuse different depth features, the network's mAP50 and mAP50:90 increased by 1.1 and 0.9 points respectively while the number of parameters and the amount of calculation were slightly reduced, indicating that the MAFPN structure can effectively fuse the feature information of different levels of the network without increasing the number of parameters and the amount of calculation. In Experiment 3, after the detection head of the original network was replaced with the Dyhead dynamic detection head, the network's mAP50 and mAP50:90 indicators increased by 1.0 and 0.7 points respectively, indicating that the Dyhead detection head can effectively unify the three attention mechanisms and thus improve the model's detection accuracy for small target objects. In Experiment 5, after adding a 160×160 small target head to the network, the network's computational workload increased, but compared with the baseline model, mAP50 and mAP50:90 increased by 5.0 and 3.4 points respectively, and the network's model parameters did not change much, which shows that adding a small target detection head, that is, retaining shallow network information, is effective for small target detection. Finally, in Experiment 8, when the three improved modules were added to the baseline network model at the same time, the network increased the computational workload due to the effect of the superposition module, but the parameters did not change much, and the mAP50 and mAP50:90 indicators increased by 8.1 and 5.6 points respectively compared with the baseline YOLOv8s network, and the network achieved a large increase. The results of the above ablation experiments show that the improved model in this paper can effectively improve the detection accuracy of drone small target image detection without a significant increase in parameters.
[0044] Step 3 also includes a visual comparison of regression loss box loss, cls loss, and dfl loss classification loss, as well as precision, recall, mAP50, and mAP50:90 during training. Figure 6 As shown, from Figure 6 It can be seen from the curve graph that the improved model of the present invention fluctuates more gently during the training process than the original model, and eventually tends to converge. The parameter indicators such as mAP50, mAP50:90, Precision and Recall of the improved model all increase.
[0045] At the same time, the present invention selected target detection networks such as YOLOv5, YOLOv7 and GOLD-YOLO to conduct comparative experiments between different algorithms on the VisDrone2019 dataset. The experimental results are shown in Table 3.
[0046] Table 3 Comparative experiments between different algorithms
[0047]
[0048] It can be seen from Table 3 that the relevant parameter indicators of the improved model of the present invention have certain advantages over mainstream target detection algorithms such as YOLOv5s, YOLOv7s and GOLD-YOLO.
[0049] The above step S4 is specifically as follows:
[0050] The VisDrone2019 dataset selects three scenes: high altitude perspective, dim and dense, and conducts visual comparative analysis of the original network model and the improved network model. Figure 7 As shown, Figure 7 (a) Figure 7 (b) and Figure 7 (c) The selected original image, the YOLOv8s detection effect diagram and the detection effect diagram of the improved model respectively.
[0051] From the first set of experiments with high-altitude perspectives, it can be seen that the original model produced false detections at high-rise locations and the confidence of false detections was higher than that of the improved model, while the confidence of correctly detected targets was lower than that of the improved model. From the second set of visualization comparisons of dim scenes, it can be seen that the original YOLOv8s model had significantly higher missed detections than the improved model. In the comparison of the third set of dense scenes, the improved model still had a higher detection rate than the original model in areas that were difficult to distinguish in the image. From the three sets of visualization comparisons, it can be seen that the improved model can detect more small target objects than the original model in different scenes, such as high-altitude perspectives, dim and dense, which shows that the improved model in this paper has a higher detection rate for small target objects than the original model. Through the visualization comparison analysis of the detection effects of the two models, it can be seen that the improved model in this paper effectively improves the detection accuracy of the original model for small target objects on the VisDrone2019 dataset.
[0052] In summary, in order to solve the problem that the target size of the VisDrone2019 dataset is small and it is difficult to accurately detect and locate, resulting in low detection accuracy of small target objects, the present invention has made relevant improvements based on the YOLOv8s benchmark model. The MAFPN structure in MAF-YOLO is introduced in the Neck part of the network to improve the PAFPN structure of the original YOLOv8, and a small target detection head that can retain more high-resolution feature information is designed. The original YOLOv8 detection head is replaced with the Dyhead dynamic detection head with scale perception, spatial position perception and task perception. Finally, by conducting ablation experiments between different modules, the results of the improved model of the present invention increased by 8.1 and 5.6 points in mAP50 and mAP50:95 respectively on the VisDrone2019 dataset compared with the benchmark YOLOv8s model.
Claims
1. A method for detecting small targets in drone aerial photography based on improved YOLOv8, characterized in that: S1. Obtain the VisDrone2019 dataset and divide it into training set, validation set and test set, and convert them into YOLO data format respectively; S2, using the original YOLOv8 network as the benchmark model to build an improved YOLOv8 drone aerial photography small target detection network model; Improvements to the original YOLOv8 network include: introducing the MAFPN structure in MAF-YOLO in the Neck part to integrate and fuse multi-scale information at different resolution layers; adding a 160×160 small target detection head to retain more shallow network feature information so that the network can more effectively extract the features of small target objects; introducing a Dyhead detection head with multiple attention mechanisms to enable the network to effectively extract the features of small target objects; S3, respectively input the images of the training set and the validation set into the improved YOLOv8 UAV aerial photography small target detection network model to train and verify the image data, and obtain the final target detection model; S4. Based on the improved target detection model, the drone aerial images are detected to obtain the detection results.
2. The method for detecting small targets in drone aerial photography based on improved YOLOv8 as claimed in claim 1, characterized in that: The MAFPN structure is used for multi-scale feature fusion. The MAFPN structure in MAF-YOLO is introduced in the Neck part to integrate and fuse multi-scale information of different resolution layers, including: First, in the first bottom-up path, the network extracts multi-scale features from Backbone and performs preliminary auxiliary fusion in the shallow layer of Neck; Then, the network collects information from each layer through dense connections in the second top-down path, and obtains diverse output information at different resolutions through the detection head; Finally, the detection head predicts the object bounding box and its corresponding category based on the feature map of each scale to calculate its loss.
3. The method for detecting small targets in drone aerial photography based on improved YOLOv8 as claimed in claim 1, characterized in that: The addition of a 160×160 small target detection head to retain more shallow network feature information specifically includes: First, based on the baseline model YOLOv8s, the output results of the P1 layer and the P3 layer are downsampled and upsampled respectively. Then, the up-sampling and down-sampling results are concatenated with the P2 layer to finally obtain a small object detection head of 160x160 size. Finally, the network contains four detection layers of different scales: a 160×160 tiny target detection layer, an 80×80 small target detection layer, a 40×40 medium target detection layer, and a 20×20 large target detection layer.
4. The method for detecting small targets in drone aerial photography based on improved YOLOv8 as claimed in claim 1, characterized in that: The Dyhead detection head is given a detection layer and an input tensor F∈R L×S×C It can be expressed as a combination of three sequential attentions, each of which focuses on only one dimension. The calculation formula is: W(F) = π C (π S (π L (F)·F)·F)·F, where F represents the input L×S×C three-dimensional tensor, S=H×W is the size of the feature map, and π L (F), π s (F) and π C (F) represent scale-aware attention, spatial position-aware attention, and task-aware attention, respectively.