Unmanned aerial vehicle aerial photography small target detection method based on improved YOL0v8
By improving the YOLOv8 model and introducing a variety of optimization technologies, the problem of low detection accuracy of small and medium-sized drone aerial photography is solved, and higher detection accuracy and lower computing resource consumption are achieved.
Patent Information
- Application Number
- CN202510282082.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has low accuracy in detection of small and medium-sized drone aerial photography, especially in complex backgrounds and low-resolution scenarios, feature extraction capabilities are insufficient and special optimization mechanisms are lacking.
By improving the YOLOv8 model, WIoU v3 loss function, adaptive partial convolution module, SAMA attention mechanism, ADWN downsampling module and P2 detection layer are introduced to improve model performance from multiple aspects: feature extraction, feature fusion and loss optimization.
It significantly improves the detection accuracy and generalization ability of small targets in complex scenarios, and effectively reduces the consumption of computing resources.
Smart Images

Figure CN120219993A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and object detection, and particularly to a method for detecting small targets in UAV aerial photography based on improved YOLOv8. Background Art
[0002] UAV aerial photography is widely used in fields such as disaster monitoring, crop assessment, and urban planning due to its flexible perspective and wide coverage. However, the detection accuracy of small targets still faces many challenges in complex scenarios. Existing technologies mainly rely on deep learning models (such as the YOLO series) for object detection, but these methods perform poorly when dealing with small targets with low resolution, complex backgrounds, and close target spacing. The reason is that the feature extraction ability of existing models is insufficient, feature fusion is not sufficient, and there is a lack of a dedicated optimization mechanism for small targets, resulting in a significant decline in detection accuracy and generalization performance. Summary of the Invention
[0003] In view of the many problems existing in the above-mentioned prior art, the present invention provides a method for detecting small targets in UAV aerial photography based on improved YOLOv8. By improving the YOLOv8 model, the present invention introduces a variety of optimization techniques, including the WIoU v3 loss function, an adaptive partial convolution module, a SAMA attention mechanism, an ADWN downsampling module, and a P2 detection layer, to improve the model performance in terms of feature extraction, feature fusion, and loss optimization. Through these improvements, the detection accuracy and generalization ability of small targets in complex scenarios are significantly improved, and at the same time, the consumption of computing resources is effectively reduced.
[0004] A method for detecting small targets in UAV aerial photography based on improved YOLOv8 includes the following steps:
[0005] Preprocess the UAV aerial photography image data to generate standardized input data adapted to the YOLOv8 model;
[0006] Input the standardized input data into the improved YOLOv8 model, and the improvements include the following features:
[0007] Use WIoU v3 as the loss function for bounding box regression, and optimize the gradient gain allocation through a dynamic non-monotonic mechanism to suppress the excessive or harmful gradients generated by extreme samples;
[0008] Replace the C2f module in the YOLOv8 model with a lightweight adaptive partial convolution module to reduce the computational amount and improve the ability to extract spatial features;
[0009] Adopt the SAMA parameter-free self-attention mechanism to replace the downsampling convolution module, and generate three-dimensional attention weights by optimizing the energy function;
[0010] An adaptive downsampling module is introduced. Through the design of multi-branch feature fusion, the final feature map contains both the original feature information and the features of the additional processing path;
[0011] A P2 detection layer is newly added to the Neck part of the network. By fusing multi-scale features, the detection efficiency is improved, and the P5 detection branch is retained at the same time;
[0012] Based on the improved YOLOv8 model, multi-object detection is carried out, and the detection results of the target category and the bounding box position are output.
[0013] Preferably, the preprocessing steps of the image data include the following specific operations: normalizing the UAV aerial image data, scaling the pixel values to between zero and one to reduce the difference in the distribution of the input data; randomly cropping the image to generate different-sized fields of view; enhancing the orientation of the image through horizontal flipping; generating standardized input data adapted to the YOLOv8 model.
[0014] Preferably, during the bounding box regression process, the WIoU v3 loss function calculates the weighted intersection over union of the predicted box and the target box, and combines the distance quantization of the center point offset and the width-height difference to achieve the gradient gain optimization of medium-quality samples.
[0015] Preferably, the lightweight adaptive partial convolution module divides the input feature map into multiple consecutive channel groups, selects some channels according to the set ratio for regular convolution operations, and outputs the unselected channels through direct mapping.
[0016] Preferably, the regular convolution operation uses a 3×3 convolution kernel and a unit stride to ensure that the spatial resolution of the feature map remains the same before and after convolution.
[0017] Preferably, the SAMA parameter-free self-attention mechanism calculates the energy distribution of the surrounding area for each position point of the feature map to generate a three-dimensional attention weight matrix for weighting, and the weight matrix adjusts the intensity of the feature map through pixel-by-pixel multiplication.
[0018] Preferably, the adaptive downsampling module includes two parallel feature paths. The first path performs average pooling operations to extract low-frequency information; the second path performs max pooling operations to capture high-frequency edge features; finally, the output features of these two paths are linearly weighted and fused to generate an enhanced downsampled feature map.
[0019] Preferably, the newly added P2 detection layer in the Neck part upsamples the low-level feature map through bilinear interpolation, performs point-by-point weighted fusion with the high-level feature map, and after a set of convolution operations on the fusion result, generates a high-resolution feature map for optimizing the detection performance of small targets.
[0020] Preferably, the feature fusion of the P2 detection layer combines the spatial detail information of the low-level feature map and the semantic information of the high-level feature map through a weighted average strategy, enabling the model to effectively detect targets at different scales.
[0021] Preferably, the multi-object detection results are weighted and calculated by the intersection-over-union ratio of the class prediction score and the bounding box position prediction for each object to generate a comprehensive confidence score, and the final detection results are screened by setting a confidence threshold.
[0022] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:
[0023] By introducing the WIoU v3 loss function, the present invention realizes the optimized allocation of the bounding box gradient gain, enhancing the generalization ability of the model for medium-quality samples;
[0024] By means of the lightweight adaptive partial convolution module, the present invention reduces the computational complexity while improving the ability to extract spatial features;
[0025] By adopting the SAMA parameter-free self-attention mechanism, the present invention realizes the precise attention to the features of key regions and suppresses non-key information;
[0026] By introducing the ADWN downsampling module, the present invention retains more detail information and optimizes the feature representation;
[0027] By adding a new P2 detection layer, the present invention significantly improves the multi-scale feature fusion and the detection performance of small targets;
[0028] The present invention effectively solves the problems of insufficient feature expression ability and low detection accuracy in small target detection in the prior art, and at the same time improves the detection efficiency and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a schematic diagram of the overall structure of the improved YOLOv8 model of the present invention;
[0030] Figure 2 It is a schematic diagram of the PCNV adaptive partial convolution structure in the embodiment of the present invention;
[0031] Figure 3 It is a schematic diagram for comparing three attention steps in the embodiment of the present invention;
[0032] Figure 4 It is a schematic diagram of the ADWN module structure in the embodiment of the present invention;
[0033] Figures 5a - 5i It is a schematic diagram of the images collected by various types of drones performing diverse tasks from multiple angles and different scenarios in the embodiment of the present invention;
[0034] Figures 6a - 6c These are the comparison curves of amAP0.5, Precision(c), and Recall in the embodiments of the present invention;
[0035] Figure 7 This is the comparison diagram of the original image, the heat map of YOLOv8s, and the heat map of the improved model in the embodiments of the present invention;
[0036] Figure 8 This is the comparison schematic diagram of the inference results of the three models in the embodiments of the present invention. Detailed implementation manners
[0037] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.
[0038] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0039] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0040] The YOLO model has achieved remarkable success in the field of computer vision, especially in object detection tasks. Researchers have improved this method and added new modules, proposing many classic models. YOLOv8 adopted an improved version of CSPDarknet53 as its backbone network, and through five downsampling processes, five feature levels of different scales were refined, namely B1 to B5. In the design of this backbone network, the traditional cross-stage partial (CSP) module was replaced by the C2f module, which, through gradient shunting connections, enhanced the richness of feature extraction while maintaining the network's lightweight nature. The C2f module first processes the input information through the CBS module (convolution, batch normalization, and SiLU activation function), and then the backbone network uses the Spatial Pyramid Pooling Fast (SPPF) module to pool the feature map to adapt to different-sized outputs. Compared with the traditional Spatial Pyramid Pooling (SPP), SPPF reduces the computational amount and latency through three consecutive max pooling layers. Inspired by PANet, YOLOv8 adopted the PAN-FPN structure in its neck design. This design maintains the model's lightweight while, by eliminating the convolutional step after upsampling, retains the original performance. In YOLOv8, P4 - P5 and N4 - N5 represent two different-scale features in the PAN structure and the FPN structure, respectively. FPN enhances the semantic information of features by fusing B4 - P4 and B3 - P3 in a top-down manner but may lose some localization information. To address this issue, PAN-FPN strengthens the learning of position information by fusing P4 - N4 and P5 - N5, achieving path enhancement. This top-down and bottom-up network structure effectively combines shallow position information and deep semantic information through feature fusion, enhancing the diversity and integrity of features. In the detection part, YOLOv8 adopted a decoupled head structure that uses two independent branches to perform object classification and bounding box regression prediction respectively, and different loss functions are applied to these two types of tasks. The binary cross-entropy loss (BCE loss) is used for the classification task, while the distribution focal loss (DFL) and CIoU loss are used for the bounding box regression task. This detection structure not only improves the detection accuracy but also speeds up the model's convergence rate. As an anchor-free detection model, YOLOv8 can simply identify positive and negative samples and dynamically allocate samples through a task-aligned allocator, further improving the model's detection accuracy and robustness.
[0041] The present invention provides a method for detecting small targets in UAV aerial photography based on the improved YOL0v8: PAS-YOLO (i.e., the improved YOLOv8 model), as Figure 1As shown below. First, WIoU v3 is adopted as the loss function for bounding box regression. It implements a reasonable allocation of gradient gain through a dynamic non-monotonic mechanism, effectively suppressing the excessive or harmful gradients generated by extreme samples. WIoU v3 focuses on medium-quality samples, enhancing the generalization ability of the model and improving the overall performance. Second, it is proposed to replace the backbone C2f module in YOLOv8 with the lightweight adaptive partial convolution PCNV, which reduces the computational amount while enhancing the ability to extract spatial features. Then, it is proposed to use SAMA attention to replace the downsampling convolution, which can effectively identify key features and highlight them. And the adaptive downsampling ADWN module is introduced. Through the branch design, the final feature map can contain both the original feature information and the additional information obtained through different processing paths, enhancing the feature expression ability of the model. Finally, a new P2 detection layer is added to enhance feature fusion, solve the problems of multi-scale and low resolution of network features, and thus optimize the overall efficiency of the model.
[0042] The implementation process of the present invention will be described through specific embodiments as follows:
[0043] In the field of UAV aerial photography, there are a large number of small-size targets in the object detection task. Reasonable design of the loss function can significantly improve the detection performance of the model. WIoU v1 introduces distance as a metric for attention. When the target box and the prediction box overlap within a certain range, reducing the penalty of the geometric metric enables the model to obtain better generalization ability. The formula for calculating WIoU v1 is shown in formulas (1)-(3):
[0044] L WIoUV1 =R WIoU ×L IoU (1)
[0045]
[0046] L IoU =1 - IoU (3)
[0047] The present invention uses WIoU v3 to evaluate the quality of anchor boxes by defining the outlier β, and constructs a non-monotonic focusing factor γ based on β, which is applied to WIoU v1. A smaller β indicates a higher quality of the anchor box, so assigning a smaller γ value will reduce its weight in the loss function. On the contrary, a larger β indicates a lower quality of the anchor box, and a smaller gradient gain is assigned to reduce the adverse gradients generated by them. By dynamically adjusting the loss weights of high-quality and low-quality anchor boxes, the model can focus on samples of average quality and improve performance. The formula for WIoU v3 is shown in equations (4)-(6), where δ and α in equation (5) are adjustable hyperparameters for adapting different models.
[0048] L WIoUv3 =γ×LWIoUv1 (4)
[0049]
[0050] Among them, The monotonic focus coefficient. WIoU v3 combines the advantages of EIoU
[29] and SioU
[30] . It evaluates the quality of anchor boxes through a dynamic non-monotonic mechanism, enabling the model to pay more attention to medium-quality anchor boxes, thereby improving the accuracy of object localization. And by dynamically adjusting the loss weights of small targets, the ability of the model to detect small targets is enhanced. The meanings of relevant symbols in Formulas (1)-(6) are shown in Table 1:
[0051] Table 1. Explanation of symbols in the loss function
[0052]
[0053] In the field of UAV aerial photography, lightweight networks such as MobileNet
[31] , ShuffleNet
[32] , and GhostNet
[33] use depth convolution or group convolution to extract features and reduce parameters. Although these models reduce FLOP, they rarely consider FLOPs, and parameter reduction does not necessarily improve the computational speed. For an input feature of size h×w×c, the FLOP required for using a regular convolution of size k×k is shown in Equation (7). c in Equation (7) represents the number of channels of the input data.
[0054] FLOPs Conv = h×w×k 2 ×c 2 (7)
[0055] The depth convolution kernel performs a sliding operation in the input channel space to derive the output channel features, and the FLOPs of the depth convolution are calculated as shown in Equation (8).
[0056] FLOPs DWConv = h×w×k 2 ×c (8)
[0057] Therefore, lightweight network blocks designed with depth or group convolution sometimes fail to accelerate the operation and may instead increase the latency. The present invention uses adaptive partial convolution (PCNV) to replace the C2f module in the backbone network. As Figure 2 shown, the PCNV module processes the selected continuous features in the input channels using regular convolution, while retaining other features through identity mapping, thereby maintaining the number of channels unchanged. This method not only simplifies the calculation process but also effectively reduces the storage requirements, providing an efficient processing strategy for small target detection.
[0058] For the input feature with channel number cp Perform convolution calculations on the first continuous feature and derive the formula for calculating the FLOP of PCNV, as shown in formula (9):
[0059]
[0060] If the segmentation ratio r = 1 / 4, the FLOPs of PCNV are only 1 / 16 of those of conventional Conv. The memory access volume of PCNV is as shown in formula (10):
[0061]
[0062] Since only c p channels are involved in feature extraction, other (c - cp) channels cannot be removed to prevent PCNV from degenerating into a general convolution with fewer channels, which goes against the original intention of reducing computational redundancy. For continuous or regular memory access, select the first or last continuous channel as the representative of the feature map for calculation, which not only reduces the number of memory accesses and the number of parameters but also ensures the effective extraction of spatial information. The meanings of the relevant symbols in formulas (7)-(10) are shown in Table 2:
[0063] Table 2. Explanation of PCNV convolution symbols
[0064]
[0065] Many existing attention modules only generate 1D or 2D weights, which may lead to insufficient feature detection of small targets. The present invention proposes to adopt the SAMA parameter-free self-attention mechanism to generate 3D attention weights by optimizing the energy function, so as to better focus on key regions while maintaining computational efficiency and model lightweight. An energy function solution that can converge quickly is defined and derived based on the principles of neuroscience. SAMA is superior to the SE and CBAM attention modules in multiple scenarios, as Figure 3 shown.
[0066] The present invention finds important neurons and defines an energy function, which uses binary labels and adds conventional terms. Finally, the minimum energy can be calculated by formula (11):
[0067]
[0068] where μ t and are defined as shown in formula (12) below:
[0069]
[0070] Here, μ t represents the average value of neurons, represent their variances. The difference between the target neuron and other neurons x in the input feature channels is correlated through an energy function. The total number of neurons M on the channel can be obtained by H×W. The importance of each neuron can be measured by calculating i and this value optimizes the features through a scaling operation. The entire optimization stage of the SAMA module is: where in formula (13), X is the input feature,
[0071]
[0072] is the output feature enhanced by sigmoid. And E groups all the channels and spatial dimensions and adds sigmoid to limit the excessive values in E. An energy function is designed following the spatial inhibition theory to evaluate the attention weights of neurons. Through this function, an innovative attention module is constructed, aiming to highlight key features while suppressing non-key information, thus significantly improving the efficiency of feature extraction. The meanings of the relevant symbols in formulas (11)-(13) are shown in Table 3:
[0073] Table 3. Explanation of SAMA attention symbols
[0074]
[0075] In YOLOv8, the traditional 3×3 convolutional layer is responsible for downsampling. Although it can capture local features, some detailed information may be lost. In the present invention, by introducing the ADWN downsampling module, it can adaptively adjust the sampling rate, which not only reduces the computational amount but also retains more detailed information, thus improving the accuracy of target detection in the field of drone aerial photography. As Figure 4 shown.
[0076] By performing multi-path processing on the input feature map X to make up for the information loss caused by traditional single-path convolution. First, an average pooling operation is performed on the input feature map X, and X' is obtained through step (14). This step helps to retain more global information while reducing the spatial dimension of the feature map:
[0077] X' = Avgpool2d(X) (14)
[0078] Information compression is achieved by calculating the mean of adjacent elements through average pooling, which can effectively reduce the loss of details. Then, the obtained feature map X' is divided into two sub-feature maps X1 and X2, as shown in step (15):
[0079] X' = [X1, X2] (15)
[0080] Then, apply a convolution operation with a kernel size of 3, a stride of 2, and a padding of 1 to X1, and obtain X1' through step (16):
[0081] X1' = Conv(X1; k = 3, s = 2, p = 1) (16)
[0082] This operation can capture local features in X1 and achieve downsampling through convolution with a stride of 2. For another sub-feature map X2, first perform a max pooling operation to obtain X2_pool through step (17), and this step helps to retain significant features in the image:
[0083] X2_pool = MaxPool2d(X2) (17)
[0084] Then, perform a 1×1 convolution operation on the max-pooled X2_pool to obtain X2' through step (18):
[0085] X2' = Conv(X2_pool; k = 1, s = 1, p = 0) (18)
[0086] Finally, concatenate the outputs X1' and X2' of the two paths in the channel dimension through step (19) to obtain the final output feature map:
[0087] Output = Concat(X1', X2') (19)
[0088] By retaining multi-information through the dual path, the sampling rate of the feature map can be dynamically adjusted, and combined with pooling to capture global and local features, the quality of the feature map is improved, the features for object detection are optimized, and richer and more accurate feature information is provided for subsequent object detection. The meanings of the relevant symbols in formulas (14)-(19) are shown in Table 4:
[0089] Table 4. Explanation of symbols for ADWN downsampling
[0090] Symbol Explanation X, X1, X2, X1′, X2′ Are all feature maps
[0091] In traditional YOLO series models, due to the limited number of detection layers, the feature resolution is insufficient and the receptive field is limited, making it difficult to better capture the features of small targets. Therefore, in the present invention, a P2 layer is newly added on the basis of three detection scales. Assuming the original image size is F, the receptive field R can be calculated by the following formula:
[0092] R(F) = R(F - 1) + (k - 1) × s#(20)
[0093] Let \(R(F)\) be the receptive field of the \(F\)-th layer, \(R(F - 1)\) be the receptive field of the previous layer, \(k\) be the convolution kernel size, and \(s\) be the current stride. The size of the feature map of the P2 layer is \(160\times160\). Assume that the convolution kernel size and the stride are the same (a \(3\times3\) convolution kernel with a stride of 2). From the above formula, the receptive field of the first downsampling (from 640→320) is: \(R = 1+(3 - 1)\times2 = 5\); the receptive field of the second downsampling (from 320→160) is: \(R = 5+(3 - 1)\times2 = 9\). Therefore, the receptive field of the P2 layer is a \(9\times9\) region. The feature map of the P2 layer has a higher resolution and a smaller receptive field. Therefore, finer details can be detected within the same original image region, making it suitable for detecting small-scale targets. The meanings of the relevant symbols in formula (20) are shown in Table 5 as follows:
[0094] Table 5. Symbol Explanation of the P2 Detection Layer
[0095] Symbol Explanation F Size of the feature map R Size of the receptive field k, s Size and stride of the convolutional kernel R(F) Represents the receptive field of the F-th layer
[0096] Experimental Verification in Practical Applications:
[0097] The VisDrone2019 dataset is an important UAV aerial photography dataset jointly collected and created by Tianjin University and the AISKYEYE data mining team. It covers a dataset of 10 different types of targets, is shot in many cities in China, and uses various UAVs to perform diverse tasks from multiple angles and different scenarios, capturing rich information, involving the diversity of target categories (from single to complex), the variability of the number of targets (from scarce to numerous), the extensiveness of target distribution (from sparse to dense), and images under different lighting conditions.
[0098] As shown in Figures 5(a)-(i) respectively: (a) Sparse object distribution; (b) Dense object distribution; (c) Few objects; (d) Many objects; (e) Many object types; (f) Very small objects; (g) Morning; (h) Evening; (i) Night.
[0099] The hardware platform used in the experimental training stage is shown in Table 6
[0100]
[0101] Some key parameter settings during model training are shown in Table 7
[0102]
[0103] To verify the improvement effect of the improved model on the detection performance, a comparative experiment was conducted between the improved model and the baseline YOLO series models. The YOLO algorithm family used in this invention is as follows: YOLOv 3[7], which first used multi-scale detection, and its lightweight version YOLOv 3-tiny; YOLOv 4[8] used the idea of CSPNet
[40] to construct a new backbone network structure called CSPDarknet 53. YOLOv 5 used Mosaic data augmentation and adopted the Focus structure. YOLOv 7 used the efficient network architecture ELAN
[10] . YOLOV9 introduced programmable gradient information and an efficient aggregation network. YOLOV10 adopted the latest network, trained without NMS and optimized the architecture. YOLOv11 first introduced the Transformer backbone network, adopted a dynamic head design, and used dual label assignment. The results of the comparative experiment are shown in Table 8.
[0104] Table 8. Detection Results of Some YOLO Series Models and the Prediction Model
[0105]
[0106] From the experimental results in Table 8, the PAS-YOLO model (i.e., the improved YOLOv8 model) outperformed these models in terms of average detection accuracy compared with the YOLO series models, showing the best overall performance.
[0107] The present invention also conducted a comparative experiment between the PAS-YOLO algorithm model and other mainstream models. Faster R-CNN[3] optimized the high computational and structural complexity problems of R-CNN; Cascade R-CNN
[35] proposed a multi-stage detection architecture based on R-CNN; RetinaNet
[34] proposed Focal loss; CenterNet
[36] proposed anchor-free detection; the FSAF
[37] algorithm solved the disadvantages of heuristic feature selection; ATSS
[38] proposed an adaptive training sample selection mechanism.
[0108] Table 9. Detection Results of Classical Models and the Proposed Model
[0109]
[0110] According to the comparative experiment results in Table 9, the PAS-YOLO model (i.e., the improved YOLOv8 model) has the best performance compared with the performance of other mainstream models.
[0111] To verify the effectiveness of each improvement strategy proposed in the present invention, ablation experiments were conducted on the baseline model using the VisDrone2019 dataset, and the experimental results are shown in Table 10.
[0112] Table 10. Detection results after introducing different improvement strategies
[0113]
[0114] The experimental results in Table 10 show that when applied to the baseline model, each improvement strategy has improved the detection performance to varying degrees. Introducing the PCNV convolution into the backbone to replace the C2f module reduces redundant calculations, enhances the context awareness ability, and increases the mAP50 by 0.9%. Replacing the neck convolution with the SAMA attention increases the mAP50 by 1.5%. By dynamically adjusting the attention weights, the missed detection rate of small targets can be effectively reduced. Using the efficient and fast ADWN module reduces the computational and storage burdens of feature fusion and the parameters of the model, thereby increasing the mAP50 by 1.0%. Adding a new P2 detection layer enables full fusion of shallow and deep information, increasing the mAP50 by 2.9%. WIoU v3 is integrated into the regression loss of the prediction box. By adopting a more refined sample allocation strategy, WIoU v3 enhances the localization accuracy of the model, thereby increasing the mAP50 index by 0.4%. The improved model further reduces the missed detection rate of small targets, and the average detection accuracy is increased by 6.7%. It is proved that the small target detection algorithm proposed in the present invention effectively improves the detection performance of the model.
[0115] To prove the scalability of this algorithm, relevant comparative experiments were carried out on the remote sensing dataset RSOD, and the experimental results are shown in Table 11.
[0116] Table 11. Detection results of classical models and the model of the present invention
[0117]
[0118] According to the comparative experimental results in Table 11, the algorithm proposed in the present invention has increased the mAP0.5 by 2.3% on the dataset RSOD, proving that the algorithm is scalable and applicable to small target detection scenarios.
[0119] To visually and conveniently draw the detection effect diagram of the model, a comparative experimental analysis of the detection performance of the model was carried out from three aspects: curve graph, heat map, and model inference results.
[0120] First, Figure 6 shows the change curves of some important evaluation metrics of the PAS-YOLO model (i.e., the improved YOLOv8 model), YOLOv8, and YOLOv5 during the training process.
[0121] According to Figure 6, the improved model has the best results in three detection metrics: precision, recall, and mAP0.5. Compared with YOLOv5 and YOLOv8s, this method has a faster training speed and better detection effect.
[0122] Secondly, the Gradient-weighted Class Activation Mapping (Grad-CAM
[39] ) technique is used to generate heatmaps for the YOLOv8 and PAS-YOLO models (i.e., the improved YOLOv8 model). From the experimental results as Figure 7 shown, it can be inferred that YOLOv8s has drawbacks in focusing on small objects and is not sensitive enough to recognize distant objects. In contrast, the model proposed in the present invention performs better in suppressing background noise and pays more attention to small objects. In addition, the attention of the model is more concentrated on the central region of the object, which helps to improve the accuracy of bounding box prediction and thus improve the overall detection performance of the model.
[0123] Finally, inference experiments are carried out using three algorithm models of PAS-YOLO (i.e., the improved YOLOv8 model), YOLOv8s, and YOLOv5s. The present invention selects three scenarios of urban roads, public facility sites, and market intersections as experimental data to intuitively verify the detection effect of the improved method. The detection results of the three models are as Figure 8 shown.
[0124] Figure 8 Respectively: the inference result of YOLOv5, the inference result of YOLOv8s, and the inference result of the PAS-YOLO model (i.e., the improved YOLOv8 model). Compared with YOLOv5s and YOLOv8s, the method proposed in the present invention has the best detection accuracy for objects in the far field of view, while improving the leakage detection rate of the model for occluded and dense targets, effectively improving the detection performance.
[0125] The present invention proposes a target detection algorithm model of PAS-YOLO (i.e., an improved YOLOv8 model), which is designed specifically for small target detection and can cope with challenges such as target size variation, occlusion, and unstable illumination. First, the WIoU v3 loss function is introduced. Through a dynamic sample allocation strategy, it effectively reduces the model's attention to extreme samples and improves the overall performance. Second, aiming at the problems of computational redundancy and information loss in small target detection in aerial images, the lightweight adaptive partial convolution PCNV is integrated into the backbone network of YOLOv8, replacing the C2f module to reduce computational redundancy and enhance context awareness ability, and reduce information loss. In addition, in the neck structure of the model, the parameter-free self-attention SAMA is used to replace the downsampling convolution to highlight key features and suppress secondary information, enhancing the feature extraction ability. And by introducing the adaptive downsampling ADWN module, the model can adaptively adjust the sampling rate, reducing the computational burden while retaining key information. Finally, by adding a P2 detection layer, the multi-scale feature fusion ability of the network is enhanced, improving the performance of small target detection. Experiments are carried out on the VisDrone and RSOD datasets, including ablation experiments, comparative experiments, and scalability experiments, verifying the effectiveness of the proposed method from multiple perspectives. The experimental results show that compared with the baseline model, the proposed method improves the mAP performance by 6.7% and 2.3% on the VisDrone and RSOD datasets respectively. Generally speaking, this method is suitable for deployment in complex environments, showing good generality and robustness. Future research will focus on using unsupervised learning theory to reduce the distribution difference between public datasets and self-built datasets and reduce the dependence of deep learning on labeled data.
[0126] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.
[0127] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for detecting small targets in drone aerial photography based on improved YOL0v8, characterized in that: The following steps are involved: Preprocess the drone aerial image data to generate standardized input data suitable for the YOLOv8 model; The standardized input data is input into an improved YOLOv8 model, wherein the improvement includes the following features: WIoU v3 is used as the loss function for bounding box regression, and the gradient gain distribution is optimized through a dynamic non-monotonic mechanism to suppress excessive or harmful gradients generated by extreme samples; Use lightweight adaptive partial convolution modules to replace the C2f module in the YOLOv8 model to reduce the amount of calculation and improve the ability to extract spatial features; The SAMA parameter-free self-attention mechanism is used to replace the downsampling convolution module, and the three-dimensional attention weights are generated by optimizing the energy function; An adaptive downsampling module is introduced, and through a multi-branch feature fusion design, the final feature map contains both the original feature information and the features of the additional processing path; A new P2 detection layer is added to the Neck part of the network to improve detection efficiency through multi-scale feature fusion, while retaining the P5 detection branch; Multi-target detection is performed based on the improved YOLOv8 model, and the detection results of target categories and bounding box positions are output.
2. The method according to claim 1, characterized in that The preprocessing steps of the image data include the following specific operations: normalizing the drone aerial image data and scaling the pixel values to between zero and one to reduce the difference in input data distribution; randomly cropping the image to generate fields of view of different sizes; enhancing the direction of the image by horizontally flipping it; and generating standardized input data suitable for the YOLOv8 model.
3. The method according to claim 1, characterized in that The WIoU v3 loss function optimizes the gradient gain for medium-quality samples by calculating the weighted intersection-over-union of the predicted box and the target box during bounding box regression, and combining the distance quantization of the center point offset and the width-height difference.
4. The method according to claim 1, characterized in that The lightweight adaptive partial convolution module divides the input feature map into multiple continuous channel groups, selects some channels according to a set ratio to perform regular convolution operations, and outputs the unselected channels by direct mapping.
5. The method according to claim 4, characterized in that The regular convolution operation uses a three-by-three convolution kernel and a unit step size to ensure that the spatial resolution of the feature map remains consistent before and after the convolution.
6. The method according to claim 1, characterized in that The SAMA parameter-free self-attention mechanism generates a three-dimensional attention weight matrix for weighting by calculating the energy distribution of the surrounding area for each position point of the feature map. The weight matrix adjusts the intensity of the feature map by pixel-by-pixel multiplication.
7. The method according to claim 1, characterized in that The adaptive downsampling module includes two parallel feature paths. The first path performs an average pooling operation to extract low-frequency information; the second path performs a maximum pooling operation to capture high-frequency edge features; and finally, the output features of the two paths are linearly weighted to generate an enhanced downsampling feature map.
8. The method according to claim 1, characterized in that The newly added P2 detection layer in the Neck part performs bilinear interpolation upsampling on the low-level feature map, and performs point-by-point weighted fusion with the high-level feature map. The fusion result is subjected to a set of convolution operations to generate a high-resolution feature map for optimizing the detection performance of small targets.
9. The method according to claim 8, characterized in that The feature fusion of the P2 detection layer combines the spatial detail information of the low-level feature map and the semantic information of the high-level feature map through a weighted average strategy, so that the model can effectively detect targets at different scales.
10. The method according to claim 1, characterized in that The multi-target detection results generate a comprehensive confidence score by performing a weighted operation on the intersection-and-union ratio of the category prediction score and the bounding box position prediction of each target, and the final detection result is screened by setting a confidence threshold.
Citation Information
Patent Citations
Small target detection method for remote sensing image
CN118115893A
Unmanned aerial vehicle aerial photography target detection method based on PCRS-YOLO network
CN118799766A
Detection method for tiny complex target in aerial image of unmanned aerial vehicle
CN118865170A
Method for detecting infrared ship target based on improved yolov7
US20250078541A1
Cited By
Target detection method and device based on infrared image, equipment and storage medium
CN120953593A
Ultrasonic real-time focus positioning and benign and malignant auxiliary diagnosis platform based on deep learning
CN121329943A
Steel structure weld defect identification method and system
CN121482059A
Low-altitude unmanned aerial vehicle small target detection method and system based on improved YOLOv8
CN122090330A