Unmanned aerial vehicle aerial photography small target detection method based on improved YOLOv11s algorithm

By improving the YOLOv11s algorithm and optimizing the network structure and loss function, the problems of information loss and low accuracy in the detection of small aerial targets of drone photography are solved, and higher detection accuracy and robustness are achieved.

CN120014244AActive Publication Date: 2025-05-16NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510175425.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing drone aerial target detection algorithm has the problems of risk of information loss, missed detection, false detection and low detection accuracy in small target detection.

Method used

By improving the YOLOv11s algorithm, the backbone feature extraction network, the neck fusion network and the head detection network are optimized, the PPA_C3k2 module and the TriFPN structure are introduced, and the loss function is adjusted to improve the detection and positioning capabilities of small targets.

Benefits of technology

It improves the detection accuracy and positioning accuracy of small targets of aerial photography by drones, reduces the risk of information loss, and enhances the robustness and real-timeness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014244A_ABST
    Figure CN120014244A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision and target detection, and discloses an unmanned aerial vehicle aerial photography small target detection method based on an improved YOLOv11s algorithm, and the method comprises the steps: constructing an unmanned aerial vehicle aerial photography image data set, constructing an unmanned aerial vehicle aerial photography small target detection model, and optimizing the unmanned aerial vehicle aerial photography small target detection model. Training parameters of the unmanned aerial vehicle aerial photography small target detection model are set, the improved unmanned aerial vehicle aerial photography small target detection model is trained by using the training set, training weights are gradually saved in the training process, after training is finished, the model weight with the minimum verification set loss function is selected and loaded into the unmanned aerial vehicle aerial photography small target detection model, and the unmanned aerial vehicle aerial photography small target detection model is obtained. And the test set is used for evaluation. According to the invention, the detection and positioning capability of a small target is improved, and the improvement of the average precision mAP of the model is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and target detection, and specifically relates to a method for detecting small targets in drone aerial photography based on an improved YOLOv11s algorithm. Background Art

[0002] In recent years, drone technology has developed rapidly and has been widely used in many fields, such as agricultural monitoring, disaster relief, security patrols, and environmental monitoring. With its high maneuverability, low cost, and real-time data collection capabilities, drones have shown unique advantages in aerial photography detection tasks. Driven by computer vision and deep learning, drone aerial object detection has become a hot research direction. Drones can be equipped with high-resolution cameras to perform high-altitude photography in complex environments, and combined with target detection algorithms to achieve accurate recognition of ground objects. However, due to the complex environment of drone aerial photography, such as lighting changes, occlusion, scale changes, etc., target detection faces many challenges.

[0003] At present, mainstream target detection methods are mainly divided into two categories: methods based on traditional computer vision and methods based on deep learning.

[0004] (1) Traditional methods: Early target detection mainly relied on manual feature extraction and classifiers, such as SIFT, HOG, SURF and other features combined with SVM and Adaboost for target recognition. However, these methods have limited effect on target detection in complex backgrounds and are difficult to adapt to the changing environment of drone aerial photography;

[0005] (2) Deep learning methods: Deep learning methods are mainly divided into two categories: Two-stage algorithms: such as Faster R-CNN, this method first generates candidate regions through the region proposal network (RPN), and then performs fine classification and bounding box regression on these regions. Two-stage methods usually have higher detection accuracy and are suitable for fine detection tasks, but because they require multiple forward propagations, the reasoning speed is slow and it is difficult to meet the real-time requirements of drone aerial photography. One-stage algorithms: such as YOLO and SSD, these methods directly extract features from the input image and perform target detection. Compared with the two-stage methods, they have higher detection speed. The YOLO series of algorithms have been widely used in drone target detection tasks due to their end-to-end training, high-speed reasoning and good accuracy. For example, YOLOv3 improves the small target detection capability through multi-scale prediction, YOLOv5 further optimizes the detection accuracy and reasoning speed, and YOLOv7 improves the detection performance through efficient model structure optimization.

[0006] Although existing research has made some progress in drone aerial target detection, the following challenges still exist: (1) Limited small target detection capability: Due to the high shooting altitude of drones, the target accounts for a small proportion of the image, resulting in weak feature information of small targets, and existing models still have difficulty in accurately identifying small targets; (2) Insufficient feature fusion: Although existing studies have introduced structures such as FPN and PAN for feature fusion, in the small target detection task, deep features are prone to lose local information, and shallow features are difficult to express high-level semantics, resulting in the small target detection effect still needs to be improved; (3) The 5-fold downsampling of the YOLO algorithm has certain advantages for large target detection. Downsampling can reduce the amount of calculation and increase the speed, while helping to capture the overall features of large targets and improve detection accuracy; however, for small targets, excessive downsampling brings significant disadvantages: multiple downsampling causes the details of small targets to be gradually blurred or even disappear, making it difficult to accurately identify; after multiple downsampling, the position of small targets becomes inaccurate, and the positioning accuracy of the detection results decreases; the low-resolution image after downsampling may cause the background to be confused with the small target, thus affecting the detection accuracy. Summary of the invention

[0007] In order to solve the above technical problems, the present invention provides a method for detecting small targets in drone aerial photography based on an improved YOLOv11s algorithm. The method improves the detection and positioning capabilities of small targets by optimizing the YOLOv8s model framework, achieves an improvement in the average precision mAP of the model, and solves the technical problems of the YOLOv11s algorithm having a large risk of information loss, missed detection, false detection, and low detection accuracy when detecting images with complex backgrounds and small pixels such as small targets in drone aerial photography, thereby improving the accuracy, robustness and real-time performance of drone aerial photography target detection.

[0008] In order to achieve the above object, the present invention is achieved through the following technical solutions:

[0009] The present invention is a method for detecting small targets in drone aerial photography based on an improved YOLOv11s algorithm, which specifically includes the following steps:

[0010] Step 1: Build a drone aerial image dataset and randomly divide the drone aerial image dataset into a training set, a validation set, and a test set according to a preset ratio;

[0011] Step 2: construct a UAV aerial photography small target detection model based on the YOLOv11s algorithm, wherein the UAV aerial photography small target detection model is an improved YOLOv11s algorithm including a backbone feature extraction network, a neck fusion network and a head detection network;

[0012] Step 3, optimizing the drone aerial photography small target detection model in step 2, including optimizing the backbone feature extraction network, optimizing the neck fusion network, and optimizing the loss function of the head detection network;

[0013] Step 4: Set the training parameters of the UAV aerial photography small target detection model, use the training set in step 1 to train the improved UAV aerial photography small target detection model, and gradually save the training weights during the training process;

[0014] Step 5. After the training is completed, select the model weight with the smallest loss function of the validation set and load it into the UAV aerial photography small target detection model, and evaluate it with the test set.

[0015] A further improvement of the present invention is that in step 1, the format of the drone aerial image data set is YOLO format.

[0016] A further improvement of the present invention is that in step 1, the drone aerial image data set is divided into a training set, a validation set and a test set in a ratio of 8:1:1.

[0017] A further improvement of the present invention is that in step 1, the backbone feature extraction network adopts four downsampling operations, the step size of the first downsampling layer is 1, and four feature maps of 320*320, 160*160, 80*80, and 40*40 are obtained to reduce the risk of information loss. The neck fusion network receives three feature maps of 160*160, 80*80, and 40*40. Through a larger resolution, the network can better detect and locate small targets on a larger feature map.

[0018] A further improvement of the present invention is that the number of output channels of the first Conv module in the backbone feature extraction network is 16, the number of output channels of the second Conv module is 32, the number of output channels of the first PPA_C3k2 module is 64, the number of output channels of the third Conv module is 64, the number of output channels of the second PPA_C3k2 module is 128, the number of output channels of the fourth Conv module is 128, the number of output channels of the third PPA_C3k2 module is 128, the number of output channels of the fifth Conv module is 256, the number of output channels of the fourth PPA_C3k2 module is 256, the number of output channels of the SPPF is 256, and the number of output channels of the C2PSA is 256.

[0019] A further improvement of the present invention is that: in step 3, the backbone feature extraction network in the UAV aerial photography small target detection model is optimized, specifically: the PPA_C3k2 module is introduced into the backbone feature extraction network of the YOLOv11s algorithm, and the PPA_C3k2 module includes a local branch, a global branch and a C3k2 branch. The local branch is responsible for capturing detail information, and the global branch is responsible for helping to obtain the context information of the global branch. The C3k2 branch further enhances the expression ability of the feature by introducing a convolutional jump connection; specifically, the PPA_C3k2 module adopts a parallel multi-branch method to realize the multi-branch extraction process, which specifically includes the following steps:

[0020] Step 3.1.1, given the input feature tensor F′∈R H′×W′×C , adjust the input feature tensor through point-by-point convolution to obtain the point-by-point convolution output tensor F∈R H′×W′×C′ ;

[0021] Step 3.1.2: Calculate the local branch output tensor F through the local branch, global branch and C3k2 branch respectively local ∈R H′×W′×C′ , global branch output tensor F global ∈R H′×W′×C′ and the C3k2 branch output tensor F C3K2 ∈R H′×W′×C′ ;

[0022] Step 3.1.3: Output the local branch tensor F local , global branch output tensor F global and the C3k2 branch output tensor F C3K2 Add together to get the multi-branch fusion output tensor F ~ ∈R H′×W′×C′ ;

[0023] Step 3.1.4: Multi-branch fusion output tensor F obtained by multi-branch feature extraction ~ ∈R H′×W′×C′ , use the attention module CBAM to perform adaptive feature enhancement, and get the PPA_C3k2 module output F PPA_C3K2 ∈R H′×W′×C′ ;

[0024] A further improvement of the present invention is that in step 3.1.2, the local branch outputs the tensor F local and the global branch output tensor F global The specific calculation method is:

[0025] Step 3.1.2.1. Divide the input feature tensor F′ into a set of spatially continuous blocks using computationally efficient operations including unfolding and reshaping.

[0026] Step 3.1.2.2: Perform channel averaging to obtain the channel average output tensor

[0027] Step 3.1.2.3, use the feedforward neural network FFN to perform linear calculations to obtain linear calculation features, then apply the activation function to obtain the probability distribution of the linear calculation features in the spatial dimension, multiply the obtained probability distribution by the linear calculation features, adjust the weights of the linear calculation features, obtain weighted features, and use the feature selection module (FeatureSelection) to select task-related features from tokens and channels for the weighted features. Specifically, set The weighted result is expressed as where t i ∈R d represents the i-th output token. The feature selection module selects each token to operate on, and the output is where s∈R c′ and P∈R c′×C′ is a task-specific parameter, sim(·,·) is the cosine similarity function in the range [0,1];

[0028] Step 3.1.2.4: Apply linear transformation of p to each token for channel selection, and then perform reshaping and interpolation operations to finally generate the local branch output tensor F. local ∈R H′×W′×C′ and the global branch output tensor F global ∈R H′×W′×C′ .

[0029] The further improvement of the present invention is that in step 3, the neck fusion network of the drone aerial photography small target detection model is optimized specifically as follows: the TriFPN structure is introduced into the YOLOv11s algorithm, and the TriFPN structure includes three enhancement paths, namely, two top-down path fusions and one bottom-up path enhancement. The TriFPN structure introduces an adaptive fusion module, which is AFConcat, and adds an additional weight to each input feature to allow the network to learn the importance of each feature. The formula of the fusion mechanism of the adaptive fusion module is as follows:

[0030]

[0031] Among them, O is the output value, which represents the result after the adaptive weight fusion mechanism; w i represents the weight of the i-th input feature map, i represents the index number of the i-th input feature map; I i is the input value, representing the i-th input feature map; in the above formula, w iIt is a learnable weight that determines the contribution of each input feature to the final output feature. Through the above adaptive weight fusion method, the TriFPN structure can improve accuracy while minimizing the computational cost.

[0032] A further improvement of the present invention is that the optimization of the loss function of the head detection network in step 3 refers to using a combination of CIoU and dot distance Dot Distance (DotD) as a positioning loss function to enhance the small target positioning capability.

[0033] The beneficial effects of the present invention are:

[0034] The present invention optimizes the network module and introduces the PPA_C3k2 module to maintain and enhance the representation of small targets;

[0035] The present invention reduces the risk of information loss and improves the detection and positioning capabilities of small targets by optimizing the network structure and using a larger resolution feature map;

[0036] The present invention optimizes the network structure and adds a top-down path fusion, so that the high-resolution feature layer can fuse more high-level semantic features;

[0037] The present invention optimizes the network structure and adjusts the width of the backbone network to half of YOLOv11s, thereby balancing the network parameters while improving the accuracy.

[0038] The present invention improves the detection accuracy and stability of small targets by improving the loss function. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is the overall flow chart of the present invention.

[0040] Figure 2 It is a network structure diagram of the present invention.

[0041] Figure 3 It is a diagram of the PPA_C3k2 module in the backbone network of the present invention.

[0042] Figure 4 It is a Patch-Aware module diagram in the PPA_C3k2 module of the present invention.

[0043] Figure 5 It is the CBAM module diagram in the PPA_C3k2 module of the present invention.

[0044] Figure 6 It is the MAP curve of YOLOv11s.

[0045] Figure 7 It is a MAP curve diagram of the method of the present invention.

[0046] Figure 8 It is a training result diagram of the method of the present invention.

[0047] Fig. 9 This is a picture of the detection effect of YOLOv11s in a dense scene.

[0048] Fig.10 This is a diagram showing the detection effect of the method of the present invention in a dense scene. DETAILED DESCRIPTION

[0049] The following will disclose the embodiments of the present invention with drawings. For the purpose of clear description, many practical details will be described together in the following description. However, it should be understood that these practical details should not be used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.

[0050] like Figure 1-2 As shown, the present invention is a method for detecting small targets in drone aerial photography based on an improved YOLOv11s algorithm, which specifically includes the following steps:

[0051] Step 1: Build a drone aerial image dataset and divide it into training set, validation set and test set in a ratio of 8:1:1.

[0052] Step 2: construct a UAV aerial photography small target detection model based on the YOLOv11s algorithm, wherein the UAV aerial photography small target detection model is an improved YOLOv11s algorithm including a backbone feature extraction network, a neck fusion network and a head detection network;

[0053] Step 3, optimizing the drone aerial photography small target detection model in step 2, including optimizing the backbone feature extraction network, optimizing the neck fusion network, and optimizing the loss function of the head detection network;

[0054] The PPA_C3k2 module is introduced into the backbone feature extraction network of the original YOLOv11s algorithm. The advantage of PPA_C3k2 lies in the multi-branch feature extraction strategy. It adopts a parallel multi-branch method. Each branch is responsible for extracting features of different scales and levels, which helps to capture the multi-scale features of objects, thereby improving the accuracy of small object detection.

[0055] like Figure 3 As shown, the PPA_C3k2 module includes a local branch, a global branch and a C3k2 branch. The local branch is responsible for capturing detail information, the global branch is responsible for helping to obtain context information of the global branch, and the C3k2 branch further enhances the expressiveness of features by introducing convolutional jump connections.

[0056] The process of implementing multi-branch extraction using the parallel multi-branch method in the PPA_C3k2 module specifically includes the following steps:

[0057] Step 3.1.1, given the input feature tensor F′∈R H′×W′×C , adjust the input feature tensor through point-by-point convolution to obtain the point-by-point convolution output tensor F∈R H′×W′×C′ ;

[0058] Step 3.1.2: Calculate the local branch output tensor F through the local branch, global branch and C3k2 branch respectively local ∈R H′×W′×C′ , global branch output tensor F global ∈R H′×W′×C′ and the C3k2 branch output tensor F C3K2 ∈R H′×W′×C′ Specifically, the local and global branches are distinguished by controlling the patch size parameter p, and the patch size is determined by the aggregation and displacement of non-overlapping patches in the spatial dimension; then the attention matrix between non-overlapping patches is calculated to achieve the extraction and interaction of local and global features. The specific local branch output tensor F local and the global branch output tensor F global The specific calculation method is:

[0059] First, the input feature tensor F′ is partitioned into a set of spatially contiguous blocks using computationally efficient operations including unrolling and reshaping.

[0060] Then, channel-wise averaging is performed to produce a channel-wise averaged output tensor

[0061] Use the feedforward neural network FFN to perform linear calculations to obtain linear calculation features, then apply the activation function to obtain the probability distribution of the linear calculation features in the spatial dimension, multiply the obtained probability distribution by the linear calculation features, adjust the weights of the linear calculation features, and obtain weighted features. Use the feature selection module (Feature Selection) to select task-related features from tokens and channels for the weighted features. Specifically, set The weighted result is expressed as where t i ∈R d represents the i-th output token. The feature selection module selects each token to operate on, and the output is where s∈R c′ and P∈R c′×C′are task-specific parameters, sim(·,·) is the cosine similarity function in the range [0,1]; here, s acts as a task embedding, specifying which tokens are relevant to the task. Each token t i They are reweighted according to their relevance to the task embedding (measured by cosine similarity), effectively modeling token selection.

[0062] Then, a linear transformation of p is applied to each token for channel selection, followed by reshaping and interpolation operations, and finally a local branch output tensor F is generated. local ∈R H′×W′×C′ and the global branch output tensor F global ∈R H′×W′×C′ ,like Figure 4 shown.

[0063] Step 3.1.3: Output the local branch tensor F local , global branch output tensor F global and the C3k2 branch output tensor F C3K2 Add together to get the multi-branch fusion output tensor F ~ ∈R H′×W′×C′ ;

[0064] Step 3.1.4: Multi-branch fusion output tensor F obtained by multi-branch feature extraction ~ ∈R H′×W′×C′ , use the attention module CBAM to perform adaptive feature enhancement, and get the PPA_C3k2 module output F PPA_C3K2 ∈R H′×W′×C′ ,like Figure 5 shown.

[0065] The step size of the first downsampling layer in YOLOv11s is changed to 1, and a total of 4 downsamplings are performed to obtain four feature maps of 320*320, 160*160, 80*80, and 40*40 to reduce the risk of information loss. The neck network receives three feature maps of 160*160, 80*80, and 40*40. With a larger resolution, the network can better detect and locate small targets on larger feature maps.

[0066] Adjust the width of the backbone network to half of YOLOv11s, specifically: the number of output channels of the first Conv module in the backbone network is 16, the number of output channels of the second Conv module is 32, the number of output channels of the first PPA_C3k2 module is 64, the number of output channels of the third Conv module is 64, the number of output channels of the second PPA_C3k2 module is 128, the number of output channels of the fourth Conv module is 128, the number of output channels of the third PPA_C3k2 module is 128, the number of output channels of the fifth Conv module is 256, the number of output channels of the fourth PPA_C3k2 module is 256, the number of output channels of SPPF is 256, and the number of output channels of C2PSA is 256.

[0067] In step 3, the neck fusion network of the drone aerial photography small target detection model is optimized as follows: the TriFPN structure is introduced into the YOLOv11s algorithm. The TriFPN structure includes three enhancement paths, namely two top-down path fusions and one bottom-up path enhancement, a total of three enhancement paths, which are used to construct high-level semantic feature maps of all scales, and enhance the entire feature hierarchy with precise positioning signals at lower levels. The adaptive fusion module introduced in the TriFPN structure is AFConcat, which adds an additional weight to each input feature to allow the network to learn the importance of each feature. The formula of the fusion mechanism of the adaptive fusion module is as follows:

[0068]

[0069] Among them, O is the output value, which represents the result after the adaptive weight fusion mechanism; w i represents the weight of the i-th input feature map, i represents the index number of the i-th input feature map; I i is the input value, representing the i-th input feature map; in the above formula, w i It is a learnable weight that determines the contribution of each input feature to the final output feature. Through the above adaptive weight fusion method, the TriFPN structure can improve accuracy while minimizing the computational cost.

[0070] Optimizing the loss function of the head detection network refers to using a combination of CIoU and dot distance (DotD) as a positioning loss function to enhance the small target positioning capability.

[0071] Dot Distance (DotD) metric: Since a slight change in the positional relationship may make the CIOU value of a very small target very large, the use of the CIOU indicator will reduce the detection of very small targets. Therefore, a new metric Dot Distance (DotD) is introduced for small targets.

[0072] DotD is defined as the normalized Euclidean distance between the center points of two bounding boxes.

[0073] A small target can be regarded as a point, and the width and height are much less important than the center point. Therefore, DotD is defined as:

[0074]

[0075] Where D represents the Euclidean distance between two center points, and S represents the average size of all objects in a specific data set. In order to limit the range to [0,1], an exponential form is used for normalization. D and S are expressed as:

[0076]

[0077] Among them, (x A -x B ) and (y A -y B ) represent the coordinates of the two center points, M represents the number of images, and N i Indicates the number of bounding boxes.

[0078] Dot Distance is used to reduce the sensitivity of the original CIoU metric, thereby helping the network to effectively locate the area where small targets are located. Most drone data sets contain not only small targets but also medium and large targets. A combination of CIoU and Dot Distance is proposed as a positioning loss function:

[0079] Loss = 1-0.5CIoU-0.5DotD

[0080] The above loss function takes into account both medium and large targets and small targets. On the one hand, it enables the model to accurately predict the position and size of the target, and on the other hand, it accelerates the convergence of the network.

[0081] Step 4: Set the training parameters and use the training set to train the improved YOLOv11s model. The specific parameters are set as follows: the input image size is 640×640, the optimizer is Adam, the initial learning rate is 0.01, the momentum is 0.92, the number of iterations is set to 300, and the batch normalization size is 8. During the training process, the training weights are gradually saved;

[0082] Step 5: After training, select the model weight with the smallest loss function in the validation set and load it into the UAV aerial photography small target detection model, and evaluate it with the test set. The detection accuracy and visualization result graph are used to evaluate the performance of the model.

[0083] In order to comprehensively evaluate the performance of the improved algorithm, a series of comparative experiments were conducted to compare and analyze the optimized algorithm with the current mainstream object detection algorithms. In the experiment, the mAP50, mAP50-90, and Parmas indicators were mainly evaluated. The results are shown in Table 1.

[0084] Table 1 Experimental results of the improved model in the VisDrone2019 dataset

[0085]

[0086] from Figure 6 The PR curve of YOLOv11s and Figure 7 The comparison diagram of the PR curve diagram of the method of the present invention is shown and Figure 8 It can be seen that the method of the present invention has only 2.75M parameters, and the mAP50 reaches 46.97%, and the mAP50-95 reaches 29.17%. In particular, for small target categories such as pedestrians and people, the model performs well, with mAP50 as high as 53.97% and 44%, indicating that the model has strong small target detection capabilities.

[0087] Depend on Fig. 9 and Fig.10 It can be seen from the detection results that small targets are difficult to accurately identify and locate because they occupy fewer pixels in the image. Therefore, the original YOLOv11s algorithm has missed detection when facing small targets. However, through the method of the present invention, the features of small targets can be better extracted, and small targets can be detected and located more accurately. Therefore, the method of the present invention effectively improves the ability to detect and locate small targets.

[0088] In summary, the present invention can reduce the computational complexity of the model and improve the accuracy of small target detection in UAV aerial images.

[0089] The above description is only an embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of the claims of the present invention.

Claims

1. A method for detecting small targets in drone aerial photography based on an improved YOLOv11s algorithm, characterized in that: The method for detecting small targets in drone aerial photography specifically comprises the following steps: Step 1: Build a drone aerial image dataset and randomly divide the drone aerial image dataset into a training set, a validation set, and a test set according to a preset ratio; Step 2: construct a UAV aerial photography small target detection model based on the YOLOv11s algorithm, wherein the UAV aerial photography small target detection model includes a backbone feature extraction network, a neck fusion network, and a head detection network; Step 3, optimizing the drone aerial photography small target detection model in step 2, including optimizing the backbone feature extraction network, optimizing the neck fusion network, and optimizing the loss function of the head detection network; Step 4: Set the training parameters of the UAV aerial photography small target detection model, use the training set in step 1 to train the improved UAV aerial photography small target detection model, and gradually save the training weights during the training process; Step 5. After the training is completed, select the model weight with the smallest loss function of the validation set and load it into the UAV aerial photography small target detection model, and evaluate it with the test set.

2. The method for detecting small targets in drone aerial photography based on the improved YOLOv11s algorithm according to claim 1, characterized in that: In step 1, the drone aerial image dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:

1.

3. The method for detecting small targets in drone aerial photography based on the improved YOLOv11s algorithm according to claim 1, characterized in that: In step 3, the backbone feature extraction network in the UAV aerial photography small target detection model is optimized, specifically: the PPA_C3k2 module is introduced into the backbone feature extraction network of the YOLOv11s algorithm, and the PPA_C3k2 module includes a local branch, a global branch and a C3k2 branch. The local branch is responsible for capturing detail information, and the global branch is responsible for helping to obtain the context information of the global branch. The C3k2 branch further enhances the expression ability of the feature by introducing a convolutional jump connection; specifically, the PPA_C3k2 module adopts a parallel multi-branch method to realize the multi-branch extraction process, which specifically includes the following steps: Step 3.1.1, given the input feature tensor F′∈R H′×W′×C , adjust the input feature tensor through point-by-point convolution to obtain the point-by-point convolution output tensor F∈R H′×W′×C′ ; Step 3.1.2: Calculate the local branch output tensor F through the local branch, global branch and C3k2 branch respectively local ∈R H′×W′×C′ , global branch output tensor F global ∈R H′×W′×C′ and the C3k2 branch output tensor F C3K2 ∈R H′×W′×C′ ; Step 3.1.3: Output the local branch tensor F local , global branch output tensor F global and the C3k2 branch output tensor F C3K2 Add together to get the multi-branch fusion output tensor F ~ ∈R H′×W′×C′ ; Step 3.1.4: Multi-branch fusion output tensor F obtained by multi-branch feature extraction ~ ∈R H′×W′×C′ , use the attention module CBAM to perform adaptive feature enhancement, and get the PPA_C3k2 module output F PPA_C3K2 ∈R H′×W′×C′ .

4. The method for detecting small targets in drone aerial photography based on the improved YOLOv11s algorithm according to claim 3 is characterized in that: In step 3.1.2, the local branch outputs the tensor F local and the global branch output tensor F global The specific calculation method is: Step 3.1.2.

1. Divide the input feature tensor F′ into a set of spatially continuous blocks Step 3.1.2.2: Perform channel averaging to obtain the channel average output tensor Step 3.1.2.3, use the feedforward neural network FFN to perform linear calculations to obtain linear calculation features, then apply the activation function to obtain the probability distribution of the linear calculation features in the spatial dimension, multiply the obtained probability distribution by the linear calculation features, adjust the weights of the linear calculation features, obtain weighted features, and use the feature selection module (FeatureSelection) to select task-related features from tokens and channels for the weighted features. Specifically, set The weighted result is expressed as where t i ∈R d represents the i-th output token. The feature selection module selects each token to operate on, and the output is where s∈R c′ and P∈R c′×C′ is a task-specific parameter, sim(·,·) is the cosine similarity function in the range [0,1]; Step 3.1.2.4: Apply linear transformation of p to each token for channel selection, and then perform reshaping and interpolation operations to finally generate the local branch output tensor F. local ∈R H′×W′×C′ and the global branch output tensor F global ∈R H′×W′×C′ .

5. The method for detecting small targets in drone aerial photography based on the improved YOLOv11s algorithm according to claim 1, characterized in that: In step 1, the backbone feature extraction network uses four downsampling operations. The step size of the first downsampling layer is 1, and four feature maps of 320*320, 160*160, 80*80, and 40*40 are obtained to reduce the risk of information loss. The neck fusion network receives three feature maps of 160*160, 80*80, and 40*40. Through larger resolution, the network can better detect and locate small targets on larger feature maps.

6. The method for detecting small targets in drone aerial photography based on the improved YOLOv11s algorithm according to claim 1, characterized in that: The number of output channels of the first Conv module in the backbone feature extraction network is 16, the number of output channels of the second Conv module is 32, the number of output channels of the first PPA_C3k2 module is 64, the number of output channels of the third Conv module is 64, the number of output channels of the second PPA_C3k2 module is 128, the number of output channels of the fourth Conv module is 128, the number of output channels of the third PPA_C3k2 module is 128, the number of output channels of the fifth Conv module is 256, the number of output channels of the fourth PPA_C3k2 module is 256, the number of output channels of SPPF is 256, and the number of output channels of C2PSA is 256.

7. The method for detecting small targets in drone aerial photography based on the improved YOLOv11s algorithm according to claim 1, characterized in that: In step 3, the neck fusion network of the UAV aerial photography small target detection model is optimized as follows: the TriFPN structure is introduced into the YOLOv11s algorithm. The TriFPN structure includes three enhancement paths, namely, two top-down path fusions and one bottom-up path enhancement. The adaptive fusion module introduced into the TriFPN structure is AFConcat. The formula of the fusion mechanism of the adaptive fusion module is as follows: Among them, O is the output value, which represents the result after the adaptive weight fusion mechanism; w i represents the weight of the i-th input feature map, i represents the index number of the i-th input feature map; I i Is the input value, representing the i-th input feature map.

8. The method for detecting small targets in drone aerial photography based on the improved YOLOv11s algorithm according to claim 1, characterized in that: Optimizing the loss function of the head detection network in step 3 refers to using a combination of CIoU and point distance DotDistance (DotD) as a positioning loss function to enhance the small target positioning capability.

Citation Information

Patent Citations

  • Unmanned aerial vehicle aerial photography small target detection method based on improved YOLOv7 algorithm

    CN118212553A

  • Unmanned aerial vehicle aerial photography small target detection method based on improved YOLOv8

    CN119152390A

  • Green orange detection method based on self-supervised comparative learning and computer device

    CN119169472A

Cited By

  • Aerial photography small target detection method combining windmill convolution and weighted multi-branch fusion

    CN120510540A

  • Track small target detection method and device, electronic equipment and storage medium

    CN120807890A

  • Fan blade fault detection method and device, storage medium and computer equipment

    CN121010549A