DA-YOLOv11-based unmanned aerial vehicle aerial image target detection method

By introducing an expanded residual module in the aerial image object detection of drone aerial image, improving the context anchor attention mechanism, replacing the upsampling operator and optimizing the loss function, the problem of low detection accuracy of small targets in complex environments is solved, and more efficient and accurate object detection is achieved.

CN120014487APending Publication Date: 2025-05-16DALIAN UNIV

Patent Information

Application Number
CN202510042069.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In complex environments, drone aerial image object detection has problems such as missed detection, false detection and low detection accuracy.

Method used

A drone aerial image object detection method based on DA-YOLOv11 is proposed. By introducing an expansion residual module (C3K2_DWR), the context anchor attention mechanism (ACSPPF) is improved, the traditional sampling operator is replaced with dynamic upsampling (DySample), and bounding box regression is optimized using the WIoU loss function.

Benefits of technology

It significantly improves the model's detection accuracy of small targets, reduces missed and missed detection, and improves the real-time and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014487A_ABST
    Figure CN120014487A_ABST
Patent Text Reader

Abstract

A DA-YOLOv11-based unmanned aerial vehicle aerial image target detection method belongs to the technical field of unmanned aerial vehicle aerial image target detection, and comprises the following steps: S1, obtaining and preprocessing a disclosed unmanned aerial vehicle aerial image data set; s2, carrying out improvement on the basis of a YOLOv11 network model, and constructing a DA-YOLOv11 network model; s3, training an unmanned aerial vehicle aerial image target detection network based on DA-YOLOv11; and S4, inputting the test set for testing and evaluation. An expansion residual module is introduced, and a C3K2DWR module is provided, so that the receptive field of the model for a small target is enhanced, and the features of the small target are better captured; a self-adaptive pyramid pooling module is provided in combination with an improved context anchor point attention mechanism to enhance the expression of small target features, and global information interaction is enhanced at the same time; traditional bilinear up-sampling is replaced by a dynamic up-sampling operator, so that the feature fusion capability is enhanced; a WIoU loss function is adopted to adjust the shape and position information of a target frame, and the sensitivity of the model to a small target is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection in unmanned aerial imagery, and in particular relates to a target detection method for unmanned aerial imagery based on DA-YOLOv11. Background Art

[0002] In recent years, with the continuous maturity of drone research and development technology, drones have been increasingly widely used in the civilian field due to their advantages such as simple operation, strong flexibility and low cost, such as forest fire prevention, power system inspection and intelligent transportation. Taking the construction industry as an example, drone aerial photography can be used to detect the safety of buildings, monitor the construction progress of construction sites, and detect whether workers are wearing safety helmets. In the field of traffic safety, drone aerial photography can be used to monitor vehicle and pedestrian flow, and provide data support for road planning and traffic management. In addition, drone aerial photography can also assist power inspections, improve inspection efficiency and reduce the risk of casualties.

[0003] With the breakthrough progress of deep learning technology, target detection algorithms represented by convolutional neural networks have comprehensively surpassed traditional algorithms, showing significant advantages in terms of robustness, accuracy, and running speed. Deep learning algorithms are based on a large number of samples and learn image features through parameter iteration. The features obtained have good generalization ability. Compared with traditional algorithms, they can perform detection tasks more effectively in various specific scenarios. The scenes of drone aerial images are complex, and the number and types of targets are large and densely distributed. Although the high-altitude perspective expands the field of view and obtains richer information, it also brings new problems such as small target size and target obstruction by complex terrain. Existing target detection algorithms can achieve high detection performance on medium and large targets, but the detection of small targets in complex environments still faces problems such as missed detection, false detection, and low detection accuracy.

[0004] The current mainstream deep learning target detection algorithms are mainly divided into two categories: two-stage algorithms represented by R-CNN and Faster R-CNN, and one-stage algorithms represented by the YOLO series. The two-stage algorithm has high computational complexity and is difficult to achieve real-time detection; while the one-stage algorithm directly predicts the target category and location, is faster, and is more suitable for the real-time requirements of drone aerial photography target detection.

[0005] However, the target detection task in UAV aerial images still faces many challenges due to its complex background environment, large number and types of targets, dense distribution of small targets and easy mutual occlusion. In order to improve the performance of target detection in UAV aerial images, many researchers have carried out active exploration. For example, Liao Guangping's research paper on small target detection algorithm in UAV aerial images based on improved SSD is based on the improvement of SSD algorithm, and proposed SSD-FC algorithm, which combines deep semantic information with shallow detail information through multi-scale feature fusion mechanism, improves the feature expression ability of small targets, and thus improves the detection accuracy, but this method increases the number of model parameters, which is not conducive to real-time detection. Wu Dong, Zhang Changliang, Pu Yuegang, et al.'s paper Improved Aerial Small Target Detection Algorithm Based on YOLOv7-tiny proposes an improved YOLOv7-tiny small target detection algorithm, which effectively improves the small target detection performance by introducing MobileViT block to enhance feature extraction, EVC block to optimize feature fusion, and MPDIoU loss function to improve the bounding box regression accuracy; Shen Xueli and Wang Lingchao's paper UAV Aerial Target Detection Based on YOLOv8n proposes the SFEYOLO algorithm to improve the detection accuracy of small targets in UAV aerial images. The algorithm improves the detection accuracy of the model for small targets through shallow feature enhancement, deformable convolution, multi-scale feature fusion (ASPPF) and medium-scale feature synthesis. The detection accuracy has reached 30.5% through verification on the VisDrone dataset, but the effect is still not ideal; Huang Haisheng and Rao Xuefeng's lightweight target detection for drone aerial photography scenes uses MobileNetV3 to replace the backbone network of YOLOv5 to reduce the complexity of the model, and combines the Convolutional Block Attention Module (CBAM) and SiLU activation function to build a lightweight and efficient YOLOv5.tiny network, which achieves a good balance between speed and accuracy in the task of aerial image target detection;

[0006] Ouyang Xi, Liu Qing, Wu Wei, et al. proposed a method for detecting aerial targets in drones based on the PCRS-YOLO network. The patent application number is CN118799766B. In order to solve the problems of false detection, missed detection, and low detection rate faced by small target detection in complex environments, a method for detecting aerial targets in drones based on the PCRS-YOLO network was proposed. This method is improved on the YOLOv8 basic model. The four-scale feature fusion structure is used in the neck network to overcome the problem of less RGB information of small targets. The Rep-CPCA attention mechanism is combined to help the network better capture important features, significantly improving the network's detection accuracy for small targets. However, the complex network structure and the fusion of multiple mechanisms of this invention make the floating-point computing amount reach 9.7G, which will inevitably increase the computing time of each detection and greatly increase the demand for computing resources.

[0007] Chen Jinguang, Zhao Sicheng, and Ma Lili applied for patent number CN202410900491.9 for a detection method for small and complex targets in drone aerial images. Based on the YOLOv8 model, they proposed a detection network for small and complex targets in drone aerial images. The invention introduces dynamic snake convolution to enhance the feature extraction capability of the backbone network, integrates an efficient multi-scale attention mechanism to optimize the feature transfer of the neck network, and uses the WIOU loss function and small target detection head to enhance the focus on small targets. Although the new loss function and small target detection head designed by them perform well in small target detection, in the complex and changeable actual aerial photography environment, when facing drastic changes in lighting conditions and severe target occlusion, the stability and accuracy of the model still need to be improved through verification and further optimization of a large amount of actual data.

[0008] Zhou Liming, Liu Zhehao, Zhao Hang, et al. proposed a multi-scale target detection method for drone aerial images based on coordinate and global information aggregation, patent application number: CN202310775421.0, which is used to detect drone images with complex and changeable shooting angles. This method combines coordinate information and global information to alleviate the interference of background factors in the feature extraction process, and improves the detection performance of multi-scale targets through the feature fusion network MFPPN. The effectiveness of the proposed model was verified by experiments on the VisDrone dataset, but the MFPPN structure designed by this method increases the complexity of the model, making the model parameters reach 35.8M, which cannot meet the real-time requirements.

[0009] He Ning, Wang Xin, Zhang Jingzun, et al. proposed a lightweight target detection method for drone aerial images, with patent application number: CN202310639688.7. This method uses a multi-scale feature extraction module (MSFEM) to extract features of different scales, fuses multi-scale information through a bidirectional dense feature pyramid network (BDFPN), and uses a tiny object detection head to improve the accuracy of small target detection. The detection accuracy reaches 33.4%. However, the performance of this method in dealing with extremely small targets still needs to be further optimized.

[0010] In summary, in the target detection task of drone aerial images, small targets have a limited pixel share, which leads to their weak features being easily lost in complex backgrounds after multiple convolution and pooling operations in the deep network, which has an adverse effect on the precise positioning and classification of small targets. Although current research is committed to introducing various attention mechanisms, such as spatial coordinate attention, channel attention, and self-attention mechanisms, to enhance the model's ability to capture key information, thereby improving the model's sensitivity to spatial position, channel, and semantic information, however, in complex scenes where small targets are densely distributed and vary in size, these methods still have difficulty in accurately locating the region of interest and extracting key features, which limits their detection performance in complex environments and leads to low detection accuracy. Summary of the invention

[0011] In order to solve the above problems, the present invention proposes: a method for detecting targets in drone aerial images based on DA-YOLOv11, comprising the following steps:

[0012] S1. Obtain a public drone aerial image dataset and perform preprocessing.

[0013] S2, based on the YOLOv11 network model, improve and build the DA-YOLOv11 network model;

[0014] S3, training the drone aerial image target detection network based on DA-YOLOv11;

[0015] S4. Input the test set for testing and evaluation.

[0016] Furthermore, in step S1, the data set is preprocessed:

[0017] Obtain the publicly available VisDrone2019 dataset, divide it into training set, validation set, and test set, and convert its format to YOLO format.

[0018] Furthermore, in step S2, YOLOv11n with a smaller number of parameters is used as the basic model for improvement to construct a DA-YOLOv11 network model. First, the C3K2 module in the backbone network and the neck network is improved, and a C3K2_DWR module is proposed by introducing a dilated residual module DWR;

[0019] Secondly, the contextual anchor attention mechanism CAA module is improved, and the ACSPPF adaptive pyramid pooling module is proposed in combination with adaptive pooling;

[0020] Again, the traditional bilinear upsampling is replaced by the dynamic upsampling DySample operator;

[0021] Finally, the WIoU loss function is used to optimize the original CIoU loss function.

[0022] Furthermore, the C3K2_DWR module is proposed:

[0023] By applying the DWR module to C3K2 to replace the standard convolution in Bottleneck, an improved C3K2_DWR module is proposed, which allows the model to freely sample the input feature map. It first uses 3x3 convolution to obtain feature information, and then adopts a three-branch structure to expand the receptive field; each branch performs 3x3 depth convolution at expansion rates of 1, 3, and 5 to extract semantic information; finally, the feature map obtains semantic residuals through a batch normalization layer, and uses 1x1 point-by-point convolution Conv to extract spatial features; the regional feature maps of different scales generated by regional residualization are combined with the application of multi-rate hole depth-separable convolution in semantic residualization to simultaneously obtain contextual information of different scales; its small-scale regional feature maps use convolution kernels with smaller receptive fields for fine feature extraction in semantic residualization; for large targets in aerial images, large-scale regional feature maps obtain overall features through convolution kernels with corresponding larger receptive fields.

[0024] Furthermore, the ACSPPF module is proposed by combining the adaptive pooling operation and the improved contextual anchor attention mechanism to improve CAA. A dilated convolution DilConv branch is added on the basis of CAA, so that the model can better recognize the contextual information of occluded targets and small objects.

[0025] The output eigenvalue F of the improved CAA module C Defined as:

[0026] f d =DW(DW(Conv(Avg(x))))+DilConv(x) (1)

[0027] F C =Sigmiod(Conv(fd )) (2)

[0028] Where DW is the depth-separable convolution; Avg is the average pooling;

[0029] The improved CAA module is divided into two branches. The main branch first uses average pooling and 1x1 convolution to extract local features. Secondly, it uses two depth-separable convolutions to capture long-distance contextual information in the horizontal and vertical directions respectively, and adds and fuses the convolution results in the two directions. The other branch is an expanded convolution with a dilation rate of 1, which is used to supplement the detail information lost by the main branch, thereby helping the model to better identify the contextual information of occluded targets and small objects. Next, a 1x1 convolution layer and a Sigmoid activation function are used to calculate the attention weights. The generated weight matrix is ​​used to perform element-by-element multiplication with the original feature map to highlight important features, thereby improving target detection accuracy. It is introduced into the SPPF module, and the ACSPPF module is proposed in combination with adaptive average pooling.

[0030] ACSPPF first passes through a 1x1 convolution layer, which halves the number of channels to obtain a new feature map y. Three maximum pooling operations are performed on y to obtain feature maps of different scales. At the same time, an adaptive average pooling operation is performed to compensate for the information loss caused by the maximum pooling operation and obtain the global feature. The global feature is expanded to the same size as other feature maps and concatenated with the feature maps obtained by three maximum pooling operations to obtain a new feature map x. A 1x1 convolution operation is performed on the concatenated feature map x to adjust the number of channels. x is input into the improved CAA attention mechanism for efficient multi-scale feature fusion and full mining of contextual information, thereby improving the model's detection accuracy for small targets.

[0031] Furthermore, the upsampling operator is optimized:

[0032] A dynamic upsampling operator DySample is used to perform sampling by finding the correct semantic clustering of each upsampling point;

[0033] DySample is an upsampling method based on point sampling. It learns the coordinates of the sampling points in the input feature map and then generates content-aware sampling points to resample the feature map. DySample adopts a differential sampling strategy to convert the input feature map into a continuous feature map in a specific way.

[0034] Furthermore, the loss function is optimized:

[0035] Use WIoUv3 as the bounding box regression loss function to replace the original CIoU. In WIoUv3, a small gradient gain is further added to reduce the harmful gradient effect of low-quality samples and weaken the competition between high-quality anchor boxes;

[0036] The formula of WIoUv3 is defined as follows:

[0037]

[0038] L WIoUv3 =r*L WIoUv1 (10)

[0039] Where r is the non-monotonic focusing factor; β is the discreteness, which indicates the quality of the anchor frame and is assigned a smaller gradient gain; δ and α are dynamically adjustable hyperparameters. In a dynamic update state, β also changes dynamically, that is, the anchor box quality division standard is dynamic. Therefore, WIoUv3 can make the most appropriate gradient gain allocation strategy at any time to improve the model detection accuracy.

[0040] Furthermore, in step S4, testing and evaluating

[0041] The performance of the proposed model is evaluated using three indicators: mean average precision (mAP), number of floating-point operations (GFLOPs), and the number of parameters. The model is then compared with the baseline model. The definitions of precision (P), recall (R), average precision (AP), and mAP are as follows:

[0042]

[0043]

[0044] In the formula, TP represents the number of correct detections; FP represents the number of false detections; FN represents the number of missed detections; and N represents the total number of categories.

[0045] The beneficial effects of the present invention are:

[0046] 1. The present invention introduces the dilated residual module and proposes the C3K2_DWR module, which enhances the receptive field of the model for small targets and better captures the characteristics of small targets;

[0047] 2. The present invention improves the contextual anchor attention mechanism and proposes an adaptive pyramid pooling module in combination with adaptive pooling to enhance the expression of small target features and strengthen global information interaction;

[0048] 3. The present invention replaces the traditional bilinear upsampling with a dynamic upsampling operator, thereby enhancing the feature fusion capability.

[0049] 4. The present invention adopts the WIoU loss function to adjust the shape and position information of the target box, thereby improving the sensitivity of the model to small targets.

[0050] The mAP50 of the proposed algorithm on the VisDrone2019 dataset is improved by 3.1%, 5.1%, 2.9% and 5.6% compared with YOLOv11n, YOLOv8, YOLOv9 and YOLOv10, respectively, indicating that the DA-YOLOv11 algorithm can effectively complete the task of detecting dense small targets in drone aerial images. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is the DA-YOLOv11 structure diagram;

[0052] Figure 2 This is the schematic diagram of the DWR module;

[0053] Figure 3 This is the schematic diagram of the improved CAA module;

[0054] Figure 4 This is the schematic diagram of the ACSPPF module;

[0055] Figure 5 This is the schematic diagram of the DySample module;

[0056] Figure 6 This is a visual comparison chart;

[0057] Figure 7 This is a heat map comparison chart. DETAILED DESCRIPTION

[0058] In order to make the technical means and objectives adopted by the present invention easy to understand, the present invention is further described below in combination with a specific implementation method. A method for detecting targets in drone aerial images based on DA-YOLOv11 includes the following steps:

[0059] S1. Obtain a public drone aerial image dataset and perform preprocessing.

[0060] S2, based on the YOLOv11 network model, improve and build the DA-YOLOv11 network model;

[0061] S3, training the drone aerial image target detection network based on DA-YOLOv11;

[0062] S4. Input the test set for testing and evaluation.

[0063] The present invention effectively improves the network's ability to extract features of small targets, significantly reduces the problems of missed detection and false detection, and thus improves the detection accuracy of the target.

[0064] In step S1, the data set is preprocessed

[0065] Obtain the publicly available VisDrone2019 dataset, which contains 10 categories, namely pedestrians (Pedestrian, Ped), people (People, Peo), bicycles (Bicycle, Bic), cars (Car), vans (Van), trucks (Truck, Tru), tricycles (Tri), awning tricycles (Awn), buses (Bus) and motorcycles (Mot). Divide it into training sets, validation sets and test sets, and convert its format to YOLO format.

[0066] In step S2, a DA-YOLOv11 network model is constructed.

[0067] YOLOv11 optimizes complex feature extraction with its C3K2 module's multi-scale convolution kernel and channel separation strategy; its C2PSA module combines the multi-head attention mechanism and shows significant advantages in processing complex backgrounds and diverse target feature extraction. Therefore, the present invention uses the YOLOv11n with a smaller number of parameters as the basic model to improve and construct the DA-YOLOv11 network model, whose structure is as follows: Figure 1 shown.

[0068] First, the C3K2 module in the backbone network and the neck network is improved. By introducing the deformable weighted residual module (DWR), the C3K2_DWR module is proposed. Secondly, the contextual anchor attention mechanism (CAA) module is improved, and the ACSPPF adaptive pyramid pooling module is proposed in combination with adaptive pooling. Thirdly, the traditional bilinear upsampling is replaced by the dynamic upsampling DySample operator. Finally, the WIoU loss function is used to optimize the original CIoU loss function.

[0069] S21, C3K2_DWR module proposed

[0070] Vehicles, pedestrians and other target objects in drone aerial images have the characteristics of large scale changes, diverse postures, and complex geometric deformations. Traditional convolution operations cannot effectively capture the fine features of small targets due to the fixed receptive field, and do not fully grasp the overall features of large targets, resulting in limited accuracy of target detection and difficulty in adapting to the feature extraction requirements of targets of different scales at the same time. To address the above problems, the DWR module is introduced, and its structure is as follows: Figure 2As shown in the figure. It first uses 3x3 convolution to obtain feature information, and then uses a three-branch structure to expand the receptive field. Each branch performs 3x3 depth convolution at expansion rates of 1, 3, and 5 to extract semantic information. Finally, the feature map is passed through a batch normalization layer to obtain semantic residuals, and a 1x1 point-by-point convolution Conv is used to extract spatial features. The regional feature maps of different scales generated by regional residualization, combined with the application of multi-rate hole depth separable convolution in semantic residualization, can simultaneously obtain contextual information of different scales. Its small-scale regional feature map can use the convolution kernel with a smaller receptive field to perform fine feature extraction in semantic residualization; for large targets in aerial images, the large-scale regional feature map can obtain overall features through the convolution kernel with a corresponding larger receptive field. This enables the model to detect targets of different scales more comprehensively and accurately, thereby significantly improving the detection accuracy.

[0071] By applying DWR to C3K2 to replace the standard convolution in Bottleneck, an improved C3K2_DWR module is proposed, which enables the model to freely sample the input feature map and thus better learn the feature information of the target in the aerial image. The structure of the improved C3K2_DWR module is as follows: Figure 1 shown.

[0072] S22, proposed ACSPPF module

[0073] The original model YOLOv11 uses the SPPF module to extract information from feature maps of different scales and fuse them together. However, given that the target sizes in aerial images vary greatly, there are many small targets and they are easy to occlude each other, the SPPF module uses continuous maximum pooling operations, which will cause some feature information to be lost and affect the detection accuracy of the model. In addition, the SPPF module mainly focuses on the extraction of local features and has limited ability to capture global information, which in turn affects the model's processing of complex scenes in aerial images. In order to solve the above problems, the ACSPPF module is proposed by combining the adaptive pooling operation and the improved contextual anchor attention mechanism. The improved CAA module is shown in Figure 2. Figure 3 As shown, the structure of ACSPPF is as follows Figure 4 shown.

[0074] CAA is a contextual anchor attention mechanism used to enhance feature expression capabilities in target detection. Although CAA enhances target features by learning contextual information around the target and suppresses interference from irrelevant information, thereby improving target recognition accuracy, its effect is still not ideal when processing severely occluded targets in aerial images, making it difficult to accurately identify targets. Therefore, a dilated convolution (DilConv) branch is added on the basis of CAA to help the model better identify occluded targets and contextual information of small objects.

[0075] The output eigenvalue F of the improved CAA module C Defined as:

[0076] f d =DW(DW(Conv(Avg(x))))+DilConv(x) (1)

[0077] F C =Sigmiod(Conv(f d )) (2)

[0078] Where DW is the depth-wise separable convolution and Avg is the average pooling.

[0079] The improved CAA module is divided into two branches. The main branch first uses average pooling and 1x1 convolution to extract local features. Secondly, it uses two depth-separable convolutions to capture long-distance contextual information in the horizontal and vertical directions respectively, and adds and fuses the convolution results in the two directions. The other branch is a dilated convolution with a dilation rate of 1, which is used to supplement the detail information lost by the main branch, thereby helping the model to better identify the contextual information of occluded targets and small objects. Next, a 1x1 convolution layer and a Sigmoid activation function are used to calculate the attention weights. The generated weight matrix is ​​used to perform element-by-element multiplication with the original feature map to highlight important features, thereby improving the accuracy of target detection. It is introduced into the SPPF module, and the ACSPPF module is proposed in combination with adaptive average pooling.

[0080] ACSPPF first passes through a 1x1 convolution layer, which halves the number of channels to obtain a new feature map y. Three maximum pooling operations are performed on y to obtain feature maps of different scales. At the same time, an adaptive average pooling operation is performed to compensate for the information loss caused by the maximum pooling operation to obtain global features. The global features are expanded to the same size as other feature maps and concatenated with the feature maps obtained by three maximum pooling operations to obtain a new feature map x. A 1x1 convolution operation is performed on the concatenated feature map x to adjust the number of channels. x is input into the improved CAA attention mechanism for efficient multi-scale feature fusion and full mining of contextual information, thereby improving the model's detection accuracy for small targets.

[0081] S23, optimize upsampling operator

[0082] Upsampling refers to the process of increasing the resolution or dimension of low-resolution images or data in some way, aiming to increase the details and information content of the image and improve the image quality. Background noise is common in aerial images, which makes it difficult to effectively distinguish target features from noise information during feature upsampling, thereby reducing target detection accuracy. Traditional interpolation upsampling methods only rely on spatial information and ignore semantic information, resulting in the target area and noise area being equally enlarged, which in turn causes the positioning accuracy of the target position to decrease.

[0083] In order to solve the above problems, this algorithm uses a dynamic upsampling operator DySample to replace traditional upsampling, and performs sampling by finding the correct semantic clustering of each upsampling point.

[0084] DySample is an upsampling method based on point sampling, such as Figure 5 As shown in the figure, it resamples the feature map by learning the coordinates of the sampling points in the input feature map and then generating content-aware sampling points. First, DySample adopts a differential sampling strategy to convert the input feature map into a continuous feature map through a specific method (such as bilinear interpolation). At each sampling stage, the difference between the current pixel and the adjacent pixels can be accurately determined; then, only the pixels with large differences are selected for sampling, which can effectively highlight the difference between the target and the background while reducing the amount of parameters, thereby reducing the interference of noise and irrelevant information.

[0085] S24. Optimize loss function

[0086] YOLOv11 uses CIoU as the bounding box regression loss function, and its calculation formula is as follows:

[0087]

[0088] Where IoU represents the ratio of the intersection area to the union area of ​​the predicted box and the true box; ρ 2 *(B,B gt ) represents the Euclidean distance between the center point of the predicted box and the real box, where B is the coordinate of the center point of the predicted bounding box, and B gt is the coordinate of the center point of the real bounding box; C is the diagonal length of the minimum enclosing box surrounding the predicted box and the real box; Used to measure the similarity of aspect ratio, where w and h are the width and height of the prediction box respectively, w gt and h gt are the width and height of the real box; is the weight coefficient.

[0089] Due to the dynamic acquisition process of aerial images, some low-quality samples are inevitably included. CIoU uses the bounding box aspect ratio as a penalty factor. For low-quality samples, CIoU penalizes its geometric factors too heavily, and the predicted aspect ratio is significantly different from the actual aspect ratio and cannot effectively reflect the actual aspect ratio.

[0090] Wise-IoUv3 uses a dynamic non-monotonic focusing mechanism to focus on low-quality anchor boxes. By evaluating the discreteness of the anchor boxes and a reasonable gradient gain strategy, the parameters of the low-quality anchor boxes are adjusted more accurately to make them closer to the true value, thereby improving the overall accuracy. Therefore, WIoUv3 is used as the bounding box regression loss function to replace the original CIoU. Currently, there are three versions of WIoU. In WIoUv1, distance attention is constructed based on distance measurement; in WIoUv2, a monotonic focusing mechanism is designed for cross entropy to better focus on low-quality samples; in WIoUv3, a small gradient gain is further added to reduce the harmful gradient effects of low-quality samples and weaken the competition between high-quality anchor boxes.

[0091] The formula of WIoUv1 is defined as follows:

[0092]

[0093] L IOU =1-IoU (5)

[0094] L WIoUv1 =R WIoU *L IoU (6)

[0095] Where IoU is the intersection over union ratio between the predicted box and the real box; R WIoU is a high-quality anchor box, W and H represent the width and height of the minimum bounding rectangle surrounding the predicted box and the real box respectively; (xx gt ) 2 +(yy gt ) 2 It is the square of the Euclidean distance between the center point of the predicted box and the real box, which is used to measure the deviation in position between the predicted box and the real box.

[0096] The formula of WIoUv2 is defined as follows:

[0097]

[0098] In the formula, is the monotonic focusing coefficient; is a sliding average.

[0099] The formula of WIoUv3 is defined as follows:

[0100]

[0101] L WIoUv3 =r*L WIoUv1 (10)

[0102] Where r is the non-monotonic focusing factor; β is the discreteness, which indicates the quality of the anchor box and assigns a smaller gradient gain; δ and α are dynamically adjustable hyperparameters. In a dynamic update state, β also changes dynamically, that is, the anchor box quality division standard is dynamic. Therefore, WIoUv3 can make the most appropriate gradient gain allocation strategy at any time to improve the model detection accuracy.

[0103] In step S3, the DA-YOLOv11 drone aerial image target detection network is trained

[0104] A training environment that meets the requirements is built on a device with the following hardware configuration: Intel@Xeon(R)Silver 42 10R CPU@2.40CHz×40 processor, 32GiB memory, and Quadro RTX 8000GPU.

[0105] The training related parameters are set as follows: the number of training epochs is set to 200, the batch size is set to 16, and the initial learning rate is set to 0.01.

[0106] In step S4, testing and evaluation

[0107] The performance of the proposed model is evaluated using three indicators: mean average precision (mAP), GigaFloating point operations (GFLOPs) and parameter count, and compared with the baseline model. The definitions of precision (P), recall (R), average precision (AP) and mAP are as follows:

[0108]

[0109] In the formula, TP represents the number of correct detections; FP represents the number of false detections; FN represents the number of missed detections; and N represents the total number of categories.

[0110] Experimental results of the method

[0111] (1) Ablation experiment

[0112] In order to verify the effectiveness of the adaptive pyramid pooling module (ACSPPF), dynamic upsampling operator (DySample), WIoU loss function and C3K2_DWR module for target detection in drone aerial images, the ablation experiment shown in Table 1 was designed. Among them, I is to use C3K2_DWR to replace the original C3K2 module in the backbone network and neck network, III is to use the ACPPF module to replace the SPPF module, II and IV are the introduced loss function WIoU and dynamic upsampling operator DySample respectively. It can be seen from the results that the accuracy of the improved parts has increased. VIII model is DA-YOLOv11. Although it is higher than the basic algorithm YOLOv11n in terms of parameters and calculation, its mAP50 is improved by 3.1% compared with it, which can better achieve the detection task of drone aerial images.

[0113] Table 1 Ablation experiment

[0114]

[0115] (2) Comparative experiment of detection methods

[0116] In order to better reflect the excellence and effectiveness of the proposed method, the DA-YOLOv11 method was compared with the typical method of small targets, and the experiment was verified on the VisDrone2019 dataset. Table 2 shows the experimental comparison results. As can be seen from Table 1, except for the four categories of pedestrians, people, bicycles and awning tricycles, the detection accuracy of the method of the present invention is better than that of other algorithms in each category. Compared with YOLOv11, the mAP, P and R of the present invention on VisDrone2019 are improved by 3.1%, 3.9% and 2.4% respectively, and the number of parameters is not significantly increased. Compared with YOLOv5, CenterNet, YOLOv8, YOLOv9 and YOLOv10, mAP50 is improved by 2.3%, 9.4%, 4.5%, 5.1%, 2.9% and 5.6% respectively. In summary, the proposed method has good performance in the detection accuracy of drone aerial images.

[0117] Table 2 Comparative experiments on the VisDrone2019 dataset

[0118]

[0119] (2) Visualization Analysis

[0120] In order to verify the detection effect of the DA-YOLOv11 network model, the present invention uses the YOLOv11 algorithm and the DA-YOLOv11 algorithm to perform inference on the VisDrone2019 dataset test set. The results are as follows: Figure 6As shown. Each row is a group of pictures. The four groups of pictures involved cover different typical shooting scenes, which are highly representative and diverse. The first group of pictures was taken at a road intersection during the day. In this scene, the traffic flow is large, the target types are numerous, and the background is complex and changeable, which poses a severe challenge to the algorithm's target recognition and positioning capabilities; the second group of pictures are taken from simple parking lots and roadsides. The parking lots have different parking postures and occlusions, which also test the algorithm's accurate detection capabilities; the third group shows the elevated bridge scene at night. The scene is affected by unfavorable factors such as dim light and low contrast, which greatly increases the difficulty and uncertainty of target detection; the fourth group of pictures comes from the shopping mall pedestrian street, where small targets are densely distributed, which also increases the difficulty of detection. From the detailed analysis of the marked areas in the model prediction graph, it can be seen that in the scenes of road intersections and shopping malls, the unimproved YOLOv11 algorithm has obvious false detection problems, which is specifically manifested in the failure to accurately identify the true category and location information of some target objects, and there are problems of missed detection in parking lots and viaduct scenes. When processing images of the same scene, the DA-YOLOv11 algorithm did not have similar false detection and missed detection problems, showing more reliable detection stability and accuracy. In addition, the DA-YOLOv11 algorithm has achieved significant improvements in the detection accuracy of multiple categories. For example, in the detection of car categories, its accuracy has increased by 1.1%, and in the detection of pedestrians, its accuracy has increased by 3.7%. Experimental results show that the DA-YOLOv11 algorithm has shown excellent performance advantages in the detection of drone aerial images, and can effectively cope with the challenges brought by complex practical application scenarios such as different lighting conditions, complex backgrounds, and target diversity.

[0121] In order to further verify the effectiveness of the proposed C3K2_DWR module and ACSPPF module, the baseline algorithm YOLOv11n and its improved algorithm with integrated C3K2_DWR and ACSPPF modules were used to perform reasoning for two scenarios: complex and changeable background and densely distributed small targets, and the following images were generated: Figure 7 The results show that the red area (high confidence area) in the heat map of the improved algorithm is significantly larger and more concentrated in both scenarios, indicating that the detection confidence is improved. Among them, in the scene where small targets are dense and prone to occlusion, the red area of ​​the heat map increases most significantly after the algorithm is improved using the ACSPPF module, indicating that the module effectively restores the feature information of partially occluded targets. In the scene with complex background and variable target scale, the red area of ​​the heat map increases most significantly after the algorithm is improved using the C3K2_DWR module, indicating that the introduction of the dilated residual module expands the receptive field and can capture more complete and clearer target feature information.

[0122] The present invention improves the accuracy of small target detection and reduces the phenomenon of missed detection and false detection by optimizing the network structure and loss function. In order to expand the receptive field of the network model for small targets, a C3K2_DWR module is proposed, so that the model can detect targets of different scales more comprehensively and accurately. In order to strengthen global information interaction, an adaptive pyramid pooling module ACSPPF is proposed. The module can strengthen context information interaction through the improved CAA attention mechanism, effectively handle occluded targets, further alleviate the problem of missed detection of occluded targets, and combine the adaptive pooling module to make up for the information loss problem caused by the maximum pooling operation; replace the traditional upsampling operator with a dynamic upsampling operator to reduce the impact of background noise and irrelevant information on the upsampling process; replace the CIOU loss function with the WIOU loss function, and improve the sensitivity of the model to small targets by adjusting the shape of the target box and the target position, thereby improving the detection accuracy.

[0123] The model was evaluated through a series of comparative experiments. The experimental results show that the DA-YOLOv11 algorithm has shown excellent performance advantages in target detection in drone aerial images, and can effectively cope with the challenges brought by complex practical application scenarios such as different lighting conditions, complex backgrounds, and target diversity.

[0124] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical solutions and concepts of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A method for detecting targets in drone aerial images based on DA-YOLOv11, characterized in that: The following steps are involved: S1. Obtain a public drone aerial image dataset and perform preprocessing. S2, based on the YOLOv11 network model, improve and build the DA-YOLOv11 network model; S3, training the drone aerial image target detection network based on DA-YOLOv11; S4. Input the test set for testing and evaluation.

2. The method for detecting targets in unmanned aerial images based on DA-YOLOv11 according to claim 1, characterized in that: In step S1, data set preprocessing: Obtain the publicly available VisDrone2019 dataset, divide it into training set, validation set, and test set, and convert its format to YOLO format.

3. The method for detecting targets in drone aerial images based on DA-YOLOv11 as claimed in claim 2, characterized in that: In the step S2, YOLOv11n with a small number of parameters is used as the basic model to improve and construct a DA-YOLOv11 network model. First, the C3K2 module in the backbone network and the neck network is improved, and the C3K2_DWR module is proposed by introducing the dilated residual module DWR; Secondly, the contextual anchor attention mechanism CAA module is improved, and the ACSPPF adaptive pyramid pooling module is proposed in combination with adaptive pooling; Again, the traditional bilinear upsampling is replaced by the dynamic upsampling DySample operator; Finally, the WIoU loss function is used to optimize the original CIoU loss function.

4. The method for detecting targets in drone aerial images based on DA-YOLOv11 as claimed in claim 3, characterized in that: Propose C3K2_DWR module: By applying the DWR module to C3K2 to replace the standard convolution in Bottleneck, an improved C3K2_DWR module is proposed, which allows the model to freely sample the input feature map. It first uses 3x3 convolution to obtain feature information, and then uses a three-branch structure to expand the receptive field; each branch performs 3x3 deep convolution at expansion rates of 1, 3, and 5 to extract semantic information; finally, the feature map passes through a batch normalization layer to obtain semantic residuals, and uses 1x1 point-by-point convolution Conv to extract spatial features; By generating regional feature maps of different scales through regional residualization and combining them with the application of multi-rate atrous depth-separable convolution in semantic residualization, contextual information of different scales can be simultaneously acquired. In semantic residualization, convolution kernels with smaller receptive fields are used to extract fine features in small-scale regional feature maps. For large targets in aerial images, large-scale regional feature maps obtain overall features through convolution kernels with correspondingly larger receptive fields.

5. The method for detecting targets in drone aerial images based on DA-YOLOv11 as claimed in claim 4, characterized in that: Combining the adaptive pooling operation with the improved contextual anchor attention mechanism, the ACSPPF module is proposed to improve CAA. On the basis of CAA, a dilated convolution DilConv branch is added to enable the model to better recognize the contextual information of occluded targets and small objects. The output eigenvalue F of the improved CAA module C Defined as: f d =DW(DW(Conv(Avg(x))))+DilConv(x) (1) F C =Sigmiod(Conv(f d )) (2) Where DW is the depth-separable convolution; Avg is the average pooling; The improved CAA module is divided into two branches. The main branch first uses average pooling and 1x1 convolution to extract local features, and then uses two depth-wise separable convolutions to capture long-distance context information in the horizontal and vertical directions respectively, and adds and fuses the convolution results in the two directions. The other branch is a dilated convolution with a dilation rate of 1, which is used to supplement the detail information lost by the main branch, thereby helping the model better identify the context information of occluded targets and small objects. Next, a 1x1 convolutional layer and a Sigmoid activation function are used to calculate the attention weights. The generated weight matrix is ​​used to perform element-by-element multiplication with the original feature map to highlight important features, thereby improving target detection accuracy. It is introduced into the SPPF module, and the ACSPPF module is proposed in combination with adaptive average pooling; ACSPPF first passes through a 1x1 convolution layer, which halves the number of channels to obtain a new feature map y. It then performs three maximum pooling operations on y to obtain feature maps of different scales. At the same time, it performs an adaptive average pooling operation to compensate for the information loss caused by the maximum pooling operation and obtain global features. The global feature is expanded to the same size as other feature maps, and is concatenated with the feature maps obtained by three maximum pooling operations to obtain a new feature map x. A 1x1 convolution operation is performed on the concatenated feature map x to adjust the number of channels, and x is input into the improved CAA attention mechanism for efficient multi-scale feature fusion and full mining of contextual information, thereby improving the model’s detection accuracy for small targets.

6. The method for detecting targets in drone aerial images based on DA-YOLOv11 as claimed in claim 5, characterized in that: Optimize the upsampling operator: A dynamic upsampling operator DySample is used to perform sampling by finding the correct semantic clustering of each upsampling point; DySample is an upsampling method based on point sampling. It learns the sampling point coordinates in the input feature map and then generates content-aware sampling points to resample the feature map. DySample adopts a differential sampling strategy to convert the input feature map into a continuous feature map in a specific way.

7. The method for detecting targets in drone aerial images based on DA-YOLOv11 as claimed in claim 6, characterized in that: Optimize the loss function: Use WIoUv3 as the bounding box regression loss function to replace the original CIoU. In WIoUv3, a small gradient gain is further added to reduce the harmful gradient effect of low-quality samples and weaken the competition between high-quality anchor boxes; The formula of WIoUv3 is defined as follows: L WIoUv3 =r*L WIoUv1 (10) Where r is the non-monotonic focusing factor; β is the discreteness, which indicates the quality of the anchor frame and is assigned a smaller gradient gain; δ and α are dynamically adjustable hyperparameters. In a dynamic update state, β also changes dynamically, that is, the anchor box quality division standard is dynamic. Therefore, WIoUv3 can make the most appropriate gradient gain allocation strategy at any time to improve the model detection accuracy.

8. The method for detecting targets in drone aerial images based on DA-YOLOv11 as claimed in claim 7, characterized in that: In step S4, testing and evaluation The performance of the proposed model is evaluated using three indicators: mean average precision (mAP), number of floating-point operations (GFLOPs), and the number of parameters. The model is then compared with the baseline model. The definitions of precision (P), recall (R), average precision (AP), and mAP are as follows: In the formula, TP represents the number of correct detections; FP represents the number of false detections; FN represents the number of missed detections; and N represents the total number of categories.

Citation Information

Patent Citations

  • Lightweight unmanned aerial vehicle aerial image target detection method

    CN116597331A

  • A UAV aerial target detection method based on PCRS-YOLO network

    CN118799766B

  • Detection method for tiny complex target in aerial image of unmanned aerial vehicle

    CN118865170A

Cited By

  • Chip surface defect target detection method based on improved YOLOv11

    CN120219395A

  • Terminal block drawing detection method based on two-stage optimization and multi-stage feature enhancement

    CN120260069A

  • Malaria pathogen detection system based on improved YOLO algorithm

    CN120376009A

  • Belt foreign matter visual identification system based on image processing

    CN120495993A

  • Lung CT image processing method and system based on multi-scale spatial semantic fusion

    CN120598908A