An Unmanned Aerial Vehicle Aerial Photography Vehicle Detection Method Based on Improved YOLO11n
By improving the YOLO11n model, the SOEFPN feature pyramid and Dyhead-DCNv4 detection head were constructed, and the Inner-SIoU loss function was adopted, the problem of low detection accuracy in the detection of drone aerial photography vehicles was solved, achieving higher detection accuracy and real-time performance.
Patent Information
- Application Number
- CN202411779632.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-12-05
AI Technical Summary
In the detection of drone aerial vehicle, due to the large change in vehicle scale and large target vehicles, the detection accuracy is low, the real-time performance is poor, and the generalization is poor.
By improving the YOLO11n object detection model, a new feature pyramid SOEFPN and detection head Dyhead-DCNv4 are built, and a new loss function Inner-SIoU is adopted, the model is optimized to improve detection accuracy and real-time.
It improves the accuracy of detection of small target vehicles from the aerial viewing angle of the drone, enhances the detection accuracy of vehicles with larger scale changes, reduces the accuracy loss caused by small targets, and accelerates the convergence of the model.
Smart Images

Figure CN119810761B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to a method for detecting vehicles in UAV aerial photography based on improved YOLO11n. Background Art
[0002] With the continuous acceleration of the national urbanization process and the rapid growth of the number of motor vehicles, traditional traffic management methods are facing unprecedented challenges. In recent years, UAV technology has made great progress and has gradually become one of the key components of intelligent transportation systems. UAV aerial photography has the advantages of wide field of view, large shooting range, high flexibility, etc. As a supplementary perspective for vehicle detection, it has broad application prospects in the traffic field. However, there are problems such as large scale changes and many small target vehicles in UAV aerial photography of vehicles. Therefore, studying the method for detecting vehicles in UAV aerial photography is of great significance for the development of intelligent transportation systems.
[0003] Traditional vehicle detection methods mainly rely on manually designed features and classifiers, and detect vehicles through several steps such as feature extraction, sliding window, classification, and post-processing. Due to the limitations of manually designed features, the high computational complexity of the sliding window method, and the lack of end-to-end training in traditional vehicle detection algorithms, they have problems such as poor real-time performance, poor accuracy, and poor generalization in the vehicle detection process.
[0004] With the rapid development of deep learning, the target detection technology of improving convolutional neural networks has become a current research hotspot. Convolutional Neural Network (CNN), as a mainstream deep learning algorithm, includes various models such as R-CNN, Fast-R-CNN, Faster-R-CNN, YOLO, etc. These models have achieved good results in target detection in different application scenarios. Target detection algorithms can be divided into two categories: two-stage detection algorithms using region candidate boxes and one-stage detection algorithms not using region candidate boxes. Two-stage algorithms have high accuracy but slow detection speed. Although the improved algorithms have improved the speed to a certain extent, due to the setting of candidate regions, two-stage models cannot well meet the real-time detection requirements of UAV aerial photography of vehicles. Therefore, one-stage algorithms are usually used in UAV aerial photography vehicle detection. Compared with two-stage detection algorithms, one-stage target detection algorithms have improved real-time performance but lower accuracy. Especially in UAV aerial photography vehicle detection, there are problems such as large vehicle scale changes and many small target vehicles, which are prone to missed detection and false detection. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, the present invention provides a method for detecting vehicles in drone aerial photography based on improved YOLO11n. By improving the YOLO11n object detection model, the problem of low vehicle detection accuracy caused by large vehicle scale changes and a large number of small target vehicles is solved.
[0006] The technical solution provided by the present invention includes the following steps:
[0007] Step 1: Obtain drone aerial photography vehicle images to form a first data set; the images in the first data set can be taken by a drone camera or collected from the network.
[0008] Step 2: Add annotation information to the images in the first data set to form a second data set, and divide the second data set into a training set, a validation set, and a test set.
[0009] Preferably, the second data set is divided into a training set, a validation set, and a test set in a ratio of 8:1:1.
[0010] Step 3: Build a drone aerial photography vehicle detection model based on improved YOLO11n. The model includes a Backbone network, a Neck network, and a Head network. The construction of the model further includes steps 3.1 to 3.4:
[0011] Step 3.1: The Backbone network is composed of 5 Conv modules, 4 C3k2 modules, one SPPF module, and one C2PSA module; the training set and the validation set in the second data set are used as the input of the Backbone network.
[0012] Step 3.2: The Neck network is composed of 2 upsample modules, 4 Concat modules, 4 C3k2 modules, 2 Conv modules, 1 SPD-Conv module, and 1 CSP-OmniKernel module to form a new feature pyramid named SOEFPN.
[0013] The SPD-Conv module processes the features output by the P2 feature layer to obtain features containing small target information, and then transmits them to the P3 feature layer for feature fusion; the SPD-Conv module is composed of an SPD layer and a non-strided convolutional layer. First, the SPD layer is used to convert the spatial dimension of the feature map into the depth dimension, and then the non-strided convolutional layer is used to maintain the spatial dimension and reduce the number of channels.
[0014] In the CSP-OmniKernel module, OmniKernel is an image restoration network in a complete kernel form, consisting of three branches: a global branch, a large branch, and a local branch. It can reasonably expand the convolutional kernel to the feature size, effectively learn feature representations from global to local, thereby improving the detection performance of small targets; the introduction of the CSP idea divides the input features into two parts: one part passes through the OmniKernel module to generate adaptive features; the other part retains the input features as a lightweight path to reduce computational overhead;
[0015] Step 3.3: The Head network includes 3 Dyhead-DCNv4 detection heads, and the outputs of the 3 C3k2 modules in the neck network are respectively used as the inputs of the 3 Dyhead-DCNv4 detection heads;
[0016] The Dyhead-DCNv4 detection head replaces the deformable convolution DCNv2 with DCNv4 on the basis of the Dyhead detection head; compared with Dyhead, Dyhead-DCNv4 adds a dynamic weight generation module weight, and uses both offset and weight for convolution operations during convolution, improving the detection performance and reducing the computational overhead;
[0017] Step 3.4: Replace the loss function in the original model with the Inner-SIoU function, and the formula of the optimized loss function is:
[0018] ;
[0019] In formula (1), L SIoU represents the loss of the SIoU function; IoU inner represents the Inner-IoU loss function; IoU represents the intersection over union between the predicted box and the ground truth box; The calculation formula of L SIoU is:
[0020] ;
[0021] In formula (2), Δ represents the distance loss, and Ω represents the shape loss;
[0022] In formula (1), the calculation formula of IoU inner is:
[0023] ;
[0024] In formula (3), inter represents the area of the intersection region between the predicted box and the ground truth box; union represents the area of the union region between the predicted box and the ground truth box; the calculation formulas are respectively:
[0025] ;
[0026] ;
[0027] In formulas (4) and (5), and b r respectively represent the abscissa of the right boundary of the ground truth box and the predicted box; and b l respectively represent the abscissa of the left boundary of the ground truth box and the predicted box; and b b respectively represent the ordinate of the lower boundary of the ground truth box and the predicted box; and b t respectively represent the ordinate of the upper boundary of the ground truth box and the predicted box; and respectively represent the width and height of the ground truth box; w and h respectively represent the width and height of the predicted box; ratio corresponds to the scaling factor;
[0028] Step 4: Use the training set and the validation set to train the drone aerial vehicle detection model based on the improved YOLO11n, and save the trained model as the optimal model;
[0029] Further, the specific steps of Step 4 include Steps 4.1 to 4.4:
[0030] Step 4.1: Set the training parameters of the drone aerial vehicle detection model based on the improved YOLO11n;
[0031] Specifically, the training parameters include: the number of epochs Epoch is 200, the batch size batchsize is 16, the optimizer optimizer is SGD, the initial learning rate 1r0 is 0.01, the momentum momentum is 0.937, the weight decay weight_decay is 0.0005, and the number of threads workers is 4;
[0032] Step 4.2: Input the training set, the validation set and the corresponding labels into the drone aerial vehicle detection model based on the improved YOLO11n, use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters, and adjust the model parameters by minimizing the loss function to gradually approach the optimal solution;
[0033] Step 4.3: Use the optimizer SGD to update the model parameters, so that the model parameters are updated in the direction of gradient descent until the loss functions of the training set and the validation set no longer decrease, and at the same time the mean average precision mAP, recall rate R, and accuracy P of the evaluation metrics no longer increase;
[0034] Step 4.4: Save the trained model parameters as the optimal model;
[0035] Step 5: Use the test set to test the optimal model, evaluate the test results of the test set, and if the accuracy requirement is met, the final drone aerial vehicle detection model based on the improved YOLO11n is obtained;
[0036] Further, step 5 specifically includes steps 5.1 to 5.3:
[0037] Step 5.1: Input the test set into the optimal model;
[0038] Step 5.2: Calculate the model performance metrics: accuracy P, recall R, mean average precision mAP, number of parameters, and model size. The specific calculation formulas are as follows:
[0039] ;
[0040] ;
[0041] ;
[0042] ;
[0043] where P is the accuracy, R is the recall, mAP is the mean of the average precisions of all classes, AP is the average precision, m is the total number of vehicle classes, TP represents the number of positive samples correctly identified as positive samples, FP represents the number of negative samples misidentified as positive samples, and FN represents the number of positive samples misidentified as negative samples;
[0044] Step 5.3: When the performance metrics meet the accuracy requirements, the final drone aerial vehicle detection model based on the improved YOLO11n is obtained.
[0045] Compared with the prior art, the beneficial effects of the present invention are:
[0046] A drone aerial vehicle detection method based on the improved YOLO11n disclosed by the present invention constructs a new feature pyramid SOEFPN, improving the accuracy of small target vehicle detection under the drone aerial view; adopts a new detection head Dyhead-DCNv4, improving the detection accuracy for vehicles with large scale variations; and adopts a new loss function Inner-SIoU, avoiding large accuracy losses caused by small target vehicles and accelerating the convergence of the model. Description of the Drawings
[0047] Figure 1 is a flowchart of the drone aerial vehicle detection method based on the improved YOLO11n of the present invention;
[0048] Figure 2Schematic diagram of the UAV aerial vehicle detection model structure based on the improved YOLO11n of the present invention;
[0049] Figure 3 Schematic diagram of the SOEFPN feature pyramid structure;
[0050] Figure 4 Flow chart of the CSP-OmniKernel module;
[0051] Figure 5 Schematic diagram of the OmniKernel module structure;
[0052] Figure 6 Schematic diagram of the Dyhead-DCNv4 structure; Specific implementation manner
[0053] In order to make the technical solution, structural features, achieved purpose and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific implementation manners and with the accompanying drawings. It should be noted that the specific embodiments described herein are only used to explain the present invention more clearly and are not used to limit the present invention.
[0054] Figure 1 The following is a flowchart of a method for detecting UAV aerial vehicles based on the improved YOLO11n disclosed by the present invention, and its implementation process is as follows:
[0055] Step 1: Obtain UAV aerial vehicle images to form a first data set; the UAV aerial vehicle images in the first data set can be taken by a UAV camera or collected from the network;
[0056] In this embodiment, in order to better evaluate the detection effect of a method for detecting UAV aerial vehicles based on the improved YOLO11n disclosed by the present invention, the public data set VisDrone2019 was adopted for UAV aerial vehicle detection, and five common road vehicles, namely cars, trucks, lorries, buses and motorcycles, were selected for detection to form the first data set.
[0057] Step 2: Add annotation information to the images in the first data set. Since the public data set VisDrone2019 adopted in this embodiment already has annotation information, this step is skipped.
[0058] Convert the label files in the first dataset described in this embodiment into txt format label files required by YOLO11n to form a second dataset, and divide the second dataset into a training set, a validation set, and a test set according to a ratio of 8:1:1; the divided training set includes 6903 images, the test set includes 863 images, and the validation set includes 863 images. The image size in the second dataset in this embodiment is 640×640.
[0059] Step 3: Construct a drone aerial vehicle detection model based on the improved YOLO11n. The model includes a Backbone network, a Neck network, and a Head network. The structure is as Figure 2 shown. The specific process of model construction includes steps 3.1 to 3.4:
[0060] Step 3.1: The Backbone network is composed of 5 Conv modules, 4 C3k2 modules, one SPPF module, and one C2PSA module; the training set and the validation set in the second dataset are used as the input of the Backbone network;
[0061] Step 3.2: The Neck network is composed of 2 upsample modules, 4 Concat modules, 4 C3k2 modules, 2 Conv modules, 1 SPD-Conv module, and 1 CSP-OmniKernel module to form a new feature pyramid named SOEFPN. The structure is as Figure 3 shown;
[0062] The SPD-Conv module processes the features output by the P2 feature layer to obtain features containing small target information, and then transfers them to the P3 feature layer for feature fusion; the SPD-Conv module consists of an SPD layer and a non-strided convolutional layer. First, the SPD layer is used to convert the spatial dimension of the feature map into the depth dimension, and then the non-strided convolutional layer is used to maintain the spatial dimension and reduce the number of channels;
[0063] The structure of the CSP-OmniKernel module is as Figure 4 shown, where OmniKernel is an image restoration network in a fully kernel form, consisting of three branches: a global branch, a large branch, and a local branch. It can reasonably expand the convolutional kernel to the feature size and effectively learn the feature representations from global to local, thereby improving the detection performance of small targets. The structure is as Figure 5 shown; the introduction of the CSP idea divides the input features into two parts: one part passes through the OmniKernel module to generate adaptive features; the other part retains the input features as a lightweight path to reduce the computational overhead;
[0064] Step 3.3: The Head network includes 3 Dyhead-DCNv4 detection heads, and the outputs of the 3 C3k2 modules in the neck network are respectively used as the inputs of the 3 Dyhead-DCNv4 detection heads;
[0065] The Dyhead-DCNv4 detection head replaces the deformable convolution DCNv2 with DCNv4 on the basis of the Dyhead detection head; compared with the Dyhead, the Dyhead-DCNv4 adds a dynamic weight generation module weight, and uses both offset and weight for convolution operations during convolution, improving the detection performance and reducing the computational overhead. The structure is as Figure 6 shown;
[0066] Step 3.4: Replace the loss function in the original model with the Inner-SIoU function. The formula for the optimized loss function is:
[0067] ;
[0068] In formula (1), L SIoU represents the loss of the SIoU function; IoU inner represents the Inner-IoU loss function; IoU represents the intersection over union between the predicted box and the ground truth box; The calculation formula of L SIoU is:
[0069] ;
[0070] In formula (2), Δ represents the distance loss, and Ω represents the shape loss;
[0071] In formula (1), the calculation formula of IoU inner is:
[0072] ;
[0073] In formula (3), inter represents the area of the intersection region between the predicted box and the ground truth box; union represents the area of the union region between the predicted box and the ground truth box; The calculation formulas are respectively:
[0074] ;
[0075] ;
[0076] In formulas (4) and (5), and b r respectively represent the abscissas of the right boundaries of the ground truth box and the predicted box; and b l respectively represent the abscissas of the left boundaries of the ground truth box and the predicted box; and b brespectively represent the vertical coordinates of the lower boundaries of the ground truth box and the predicted box; and b t respectively represent the vertical coordinates of the upper boundaries of the ground truth box and the predicted box; and respectively represent the width and height of the ground truth box; w and h respectively represent the width and height of the predicted box; ratio corresponds to the scaling factor;
[0077] Step 4: Use the training set and the validation set to train the drone aerial vehicle detection model based on the improved YOLO11n, and save the trained model as the optimal model; the training process of the drone aerial vehicle detection model based on the improved YOLO11n further includes steps 4.1 to 4.4:
[0078] Step 4.1: Set the training parameters of the drone aerial vehicle detection model based on the improved YOLO11n;
[0079] In this embodiment, the training parameters include: the number of epochs Epoch is 200, the batch size batchsize is 16, the optimizer optimizer is SGD, the initial learning rate 1r0 is 0.01, the momentum momentum is 0.937, the weight decay weight_decay is 0.0005, and the number of threads workers is 4;
[0080] Step 4.2: Input the training set, the validation set, and the corresponding labels into the drone aerial vehicle detection model based on the improved YOLO11n, use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters, and adjust the model parameters by minimizing the loss function to gradually approach the optimal solution;
[0081] Step 4.3: Use the optimizer SGD to update the model parameters, so that the model parameters are updated in the direction of gradient descent until the loss functions of the training set and the validation set no longer decrease, and at the same time, the mean average precision mAP, recall rate R, and accuracy P of the evaluation metrics no longer increase;
[0082] Step 4.4: Save the trained model parameters as the optimal model;
[0083] Step 5: Use the test set to test the optimal model in step 4, evaluate the test results of the test set, and if the accuracy requirement is met, the final real-time drone vehicle detection model based on the improved YOLO11n is obtained. Specifically, step 5 further includes steps 5.1 to 5.3:
[0084] Step 5.1: Input the test set into the optimal model in step 4;
[0085] Step 5.2: Calculate the model performance metrics: The performance metrics specifically include accuracy P, recall R, mean average precision mAP, number of parameters, and model size. The specific calculation formulas are as follows:
[0086] ;
[0087] ;
[0088] ;
[0089] ;
[0090] Among them, P is the accuracy, R is the recall, mAP is the mean of the average precisions of all classes, AP is the average precision, m is the total number of vehicle classes, TP represents the number of positive samples correctly identified as positive samples, FP represents the number of negative samples misidentified as positive samples, and FN represents the number of positive samples misidentified as negative samples;
[0091] Step 5.3: When the performance metrics meet the accuracy requirements, obtain the final drone vehicle detection model based on the improved YOLO11n.
[0092] In this embodiment, in order to verify the effect of the improved model disclosed by the present invention, the present invention uses the YOLOv8n model, YOLOv10n model, YOLO11n model, YOLO11s model, and the detection model disclosed in this patent to conduct tests on the VisDrone2019 dataset. The evaluation index data is shown in Table 1: ;
[0093] As can be seen from Table 1, the drone aerial vehicle detection model disclosed by the present invention has improved in terms of accuracy P, recall R, mAP50, etc. compared with the original YOLO11n model, and is superior to other YOLO models of the same parameter order, and even superior to the YOLO11s model with a higher parameter order, meeting the real-time requirements and can be deployed to the drone platform.
[0094] The above is only one embodiment of the present invention, and it does not limit the patent scope of the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A drone aerial vehicle detection method based on improved YOLO11n, characterized in that: The specific steps include: Step 1: Obtain drone aerial images of vehicles to form a first data set; the images in the first data set can be taken by a drone camera or collected from the Internet; Step 2: adding annotation information to the images in the first data set to form a second data set, and dividing the second data set into a training set, a validation set, and a test set; Step 3: Construct a drone aerial vehicle detection model based on improved YOLO11n, the model includes a Backbone network, a Neck network and a Head network, and the construction of the model further includes steps 3.1 to 3.4: Step 3.1: The Backbone network consists of 5 Conv modules, 4 C3k2 modules, 1 SPPF module and 1 C2PSA module; the training set and the validation set in the second data set are used as the input of the Backbone network; Step 3.2: The Neck network consists of 2 upsample modules, 4 Concat modules, 4 C3k2 modules, 2 Conv modules, 1 SPD-Conv module and 1 CSP-OmniKernel module to form a new feature pyramid named SOEFPN; The SPD-Conv module processes the features output by the P2 feature layer to obtain features containing small target information, and then passes them to the P3 feature layer for feature fusion; the SPD-Conv module consists of an SPD layer and a non-stride convolution layer. The SPD layer is first used to convert the spatial dimension of the feature map into a depth dimension, and then the non-stride convolution layer is used to maintain the spatial dimension and reduce the number of channels; In the CSP-OmniKernel module, OmniKernel is a fully kernel-shaped image restoration network, which consists of three branches: global branch, large branch and local branch. It can reasonably expand the convolution kernel to the feature size and effectively learn feature representation from global to local. The introduction of CSP idea divides the input features into two parts: one part passes through the OmniKernel module to generate adaptive features; the other part retains the input features as a lightweight path to reduce computational overhead. Step 3.3: The Head network includes three Dyhead-DCNv4 detection heads, and the outputs of the three C3k2 modules in the neck network are respectively used as inputs of the three Dyhead-DCNv4 detection heads; The Dyhead-DCNv4 detection head replaces the deformable convolution DCNv2 with DCNv4 based on the Dyhead detection head; compared with Dyhead, Dyhead-DCNv4 adds a dynamic weight generation module weight, which uses offset and weight to perform convolution operations in the convolution, thereby improving detection performance and reducing computational overhead; Step 3.4: Replace the loss function in the original model with the Inner-SIoU function. The optimized loss function formula is: ; In formula (1), L SIoU Represents the loss of SIoU function; IoU inner represents the Inner-IoU loss function; IoU represents the intersection-over-union ratio between the predicted box and the true box; L SIoU The calculation formula is: ; In formula (2), Δ represents distance loss and Ω represents shape loss; In formula (1), IoU inner The calculation formula is: ; In formula (3), inter represents the area of the intersection of the predicted box and the true box; union represents the area of the union of the predicted box and the true box; The calculation formulas are: ; ; In formula (4) and formula (5), and b r Represents the right boundary horizontal coordinates of the real box and the predicted box respectively; and b l Represents the left boundary horizontal coordinates of the real box and the predicted box respectively; and b b Represents the lower boundary ordinates of the real box and the predicted box respectively; and b t Represent the upper boundary ordinates of the real box and the predicted box respectively; and Represent the width and height of the real box respectively; w and h represent the width and height of the predicted box respectively; ratio corresponds to the scaling factor; Step 4: Use the training set and the validation set to train the drone aerial vehicle detection model based on the improved YOLO11n, and save the trained model as the optimal model; the drone aerial vehicle detection model training process based on the improved YOLO11n further includes steps 4.1 to 4.4: Step 4.1: Setting the training parameters of the drone aerial vehicle detection model based on improved YOLO11n; Model training parameters include: number of iterations, batch size, optimizer, learning rate, momentum, weight decay, and number of threads; Step 4.2: Input the training set and validation set and the corresponding labels into the improved YOLO11n drone aerial vehicle detection model, use the back propagation algorithm to calculate the gradient of the loss function to the model parameters, and adjust the model parameters by minimizing the loss function to gradually approach the optimal solution; Step 4.3: Use the optimizer SGD to update the model parameters in the direction of gradient descent until the loss function of the training set and the validation set no longer decreases, and the evaluation indicators mean average precision mAP, recall rate R, and accuracy P no longer increase; Step 4.4: Save the trained model parameters as the optimal model; Step 5: Use the test set to test the optimal model, evaluate the test set test results, and meet the accuracy requirements, so as to obtain the final drone aerial vehicle detection model based on the improved YOLO11n.
2. The method for detecting vehicles in drone aerial photography based on improved YOLO11n according to claim 1, characterized in that: In step 3.2, a feature pyramid SOEFPN for small target vehicles is designed; first, the P2 feature layer is processed by SPD-Conv to obtain features containing small target information, and then the features are passed to the P3 feature layer for fusion, and then CSP-OmniKernel is used for feature integration to effectively learn feature representation from global to local, thereby improving the detection performance of small target vehicles.
3. The method for detecting vehicles in drone aerial photography based on improved YOLO11n according to claim 1, characterized in that: The step 5 further includes steps 5.1 to 5.3: Step 5.1: Input the test set into the optimal model; Step 5.2: Calculate the model performance indicators: accuracy P, recall R, mean average precision mAP, number of parameters, and model size. The specific calculation formula is as follows: ; ; ; ; Where P is the precision, R is the recall, mAP is the average precision of all categories, AP is the average precision, m is the total number of vehicle categories, TP is the number of positive samples correctly identified as positive samples, FP is the number of negative samples incorrectly identified as positive samples, and FN is the number of positive samples incorrectly identified as negative samples; Step 5.3: When the performance indicators meet the accuracy requirements, the final UAV aerial vehicle detection model based on improved YOLO11n is obtained.
Citation Information
Patent Citations
Face mask detection method based on improved Yolov5
CN117765595A
PCB defect detection method for improving YOLOv5s
CN118334398A