DST-YOLO-based unmanned aerial vehicle aerial photography small target detection method

By improving the network structure of YOLOv11, using the C3k2-DWRB module, Slim-Neck structure and dynamic task collaborative detection head, the missed detection and misdetection problem of small object detection from the perspective of the drone is solved, the detection accuracy and lightweight model are improved, and the performance and complexity are balanced.

CN119992049APending Publication Date: 2025-05-13NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510058430.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

There is a problem of missed detection and misdetection in small target detection from the perspective of drones, and mainstream target detection algorithms are insufficiently optimized in specific scenarios, which affects the accuracy of the model and real-time response capabilities.

Method used

The drone aerial small object detection method based on DST-YOLO, by improving the network structure of YOLOv11, replacing the C3k2 module with the C3k2-DWRB module, using the Slim-Neck structure for neck feature fusion, and using a dynamic task collaborative detection head to improve detection accuracy and model lightweighting.

Benefits of technology

The accuracy of small object detection from the perspective of the drone is improved, the feature extraction capability and computing efficiency of the model are enhanced, and the balance between model lightweighting and performance is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992049A_ABST
    Figure CN119992049A_ABST
Patent Text Reader

Abstract

The invention provides a DST-YOLO-based unmanned aerial vehicle aerial photography small target detection method, which comprises the following steps: step 1, constructing an improved network model by taking YOLOv11 as a basic network, step 2, obtaining a small target image data set under an open-source unmanned aerial vehicle visual angle, step 3, configuring a training environment required by the network model, 4, the network model obtained through training is utilized, and detection is carried out on a to-be-detected unmanned aerial vehicle aerial photography small target. YOLOv11 is improved, the detection accuracy is improved, meanwhile, model parameters are reduced, and the balance of model light weight and performance can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image target detection, and specifically is a method for detecting small targets in drone aerial photography based on DST-YOLO. Background Art

[0002] In recent years, drone technology has developed rapidly, making it more efficient in performing complex, dangerous, harsh or space-constrained tasks. Drones are not only widely used in the military field, but also show great potential and value in civil fields such as environmental monitoring, agricultural plant protection, and express delivery. Their popularity and application scope are growing rapidly. With high flexibility and low cost, they are now widely used in smart transportation, urban management, smart agriculture and forestry, and disaster inspection.

[0003] With the rapid development of deep learning technology, its application in the field of computer vision, especially in target detection and recognition, has become very common. There are two main types of target detection technology based on deep learning: one is a two-stage target detection algorithm, and the other is a single-stage target detection algorithm. The two-stage algorithm first generates candidate regions, and then performs classification and target recognition. Typical representatives include SPP-Net and algorithms based on convolutional neural networks (CNN). The single-stage algorithm completes candidate region generation and classification at the same time, and directly outputs target category and location information. It has the characteristics of fast speed and small amount of calculation, but due to uniform and dense sampling, it is easy to cause imbalance of positive and negative samples, affecting the detection accuracy. Representative algorithms include SSD, Retina-Net and YOLO series.

[0004] However, mainstream target detection algorithms focus more on generalization and are not optimized for specific scenarios. From the perspective of drones, the scale of objects varies greatly, small targets account for a high proportion, and there are occlusions. At the same time, image and video data are easily disturbed by factors such as light source, environment, weather, and equipment, which makes the accuracy and real-time response capabilities of the model face challenges. Therefore, developing an efficient and stable small target detection method and optimizing the balance between model complexity and detection accuracy is a research-worthy and challenging issue. Summary of the invention

[0005] In order to solve the above problems, the present invention provides a method for detecting small targets in drone aerial photography based on DST-YOLO. This method improves YOLOv11 to solve the problems of missed detection and false detection of small targets under the perspective of drones. While improving the detection accuracy, it reduces the model parameters and achieves a balance between model lightweight and performance.

[0006] To achieve the above objectives, the present invention adopts the following technical means:

[0007] The present invention is a method for detecting small targets in drone aerial photography based on DST-YOLO, and the method comprises the following steps:

[0008] Step 1: Based on YOLOv11 as the basic network, an improved network model is constructed. The C3k2 module in the YOLOv11 backbone network is replaced by the C3k2-DWRB module. The neck feature fusion network adopts the Slim-Neck structure, and the GSConv and VoV-GSCSPC modules are used to realize the neck feature fusion. The detection head uses the dynamic task collaborative detection head to achieve the interactive process between the classification task and the positioning task.

[0009] Step 2: Obtain an open source dataset of small target images from the perspective of a drone.

[0010] Step 3: Configure the training environment required by the network model and load the network model into the configured training environment.

[0011] Step 4: Use the trained network model to detect the small drone aerial targets to be detected.

[0012] The further improvement of the present invention is that in step 1, the C3k2-DWRB module refers to the design idea of ​​the Dilation-wise residual (DWR) module in the ultra-real-time semantic segmentation model DWR-Seg, and replaces the Bottleneck module in the C3k2 module to expand the receptive field, enhance the feature representation and the feature connection between network layers. At the same time, the Dilatedreparam block (DRB) module is introduced to replace the convolution module in DWR to improve the fusion capability of feature information and extract rich features more efficiently.

[0013] The further improvement of the present invention is that the neck feature fusion network Slim-neck connects two GSConv modules in sequence, and replaces the C3k2 structure in the neck network with the VoV-GSCSP module, and reduces the number of channels in the middle layer of the model to reduce the number of parameters, thereby reducing the model complexity and the demand for storage resources.

[0014] A further improvement of the present invention is that the GSConv module first converts the input feature map with C1 channels into a feature map with C2 / 2 channels by using standard convolution, and then generates another feature map with the same number of channels C2 / 2 by using depthwise separable convolution. The two feature maps are then merged by splicing, and then processed by channel shuffle operation to obtain an output feature map with C2 channels, thereby achieving complete fusion of standard convolution and depthwise separable convolution outputs, and achieving efficient integration of feature information extracted by different convolution forms.

[0015] The further improvement of the present invention is that the VoV-GCSP module first performs 1×1 ordinary convolution on the input feature map to achieve dimensionality reduction, and then divides the dimensionality reduction result into two branches. One branch contains two layers of GSConv operations, and the other branch performs Concat operation with the output of the two layers of GSConv. Finally, the spliced ​​result is processed by 1×1 convolution to generate an output feature map, thereby fully exploring the feature information of the shallow network and the deep network, achieving an efficient fusion effect, and effectively improving the feature utilization efficiency and information richness.

[0016] A further improvement of the present invention is that the dynamic task collaborative detection head reduces the number of parameters by sharing convolutions, and uses a feature extractor to learn interactive features between tasks from multiple convolutional layers to generate joint features. Among them, the localization branch combines deformable convolution DCNV2 and interactive features, and the classification branch implements dynamic feature selection based on interactive features.

[0017] Furthermore, in step 3, the training environment is as follows: the GPU is NVIDIA GeForce RTX 4070Laptop GPU, the CPU is AMD Ryzen 9 7940HX, the programming language is Python3.9, the deep learning framework is Pytorch2.0.0, and the GPU acceleration library is CUDA12.3. The model parameters during training are set as follows: the resolution is 640×640, the batch size is 16, and the number of iterations is 200.

[0018] The beneficial effects of the present invention are:

[0019] Enhanced receptive field and feature extraction capabilities: By integrating the DWR and DRB modules to improve the C3K2 structure of YOLOv11, the model's feature extraction and multi-scale perception capabilities can be improved. The DWR module expands the network's receptive field through dilated convolutions with different dilation rates, while the DRB module enhances the large-kernel convolution layer through structural reparameterization, allowing the model to more effectively capture the relationship between sparse features and distant pixels;

[0020] Optimized network structure and computational efficiency: The Slim-Neck module is used to improve the neck structure of YOLOv11, which can optimize the network structure and improve the computational efficiency. The reasonably designed convolutional structure reduces the computational complexity while maintaining the detection accuracy, making it more suitable for resource-constrained devices such as drones, and improving the feasibility and practicality of the model in practical applications.

[0021] Enhanced task interaction and target scale adaptability: The dynamic task collaborative detection head can significantly reduce the number of parameters by using shared convolutions. In addition, in order to deal with the problem of inconsistent target scales detected by each detection head, the features can be scaled. In addition, the detection head can learn task interaction features from multiple convolutional layers through feature extractors to obtain joint features, so that the detection head can better adapt to targets of different scales. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 The overall network structure diagram of the present invention.

[0023] Figure 2 This is the C3k2-DWRB module structure diagram.

[0024] Figure 3 This is the GSConv module structure diagram.

[0025] Figure 4 This is the VoV-GSCSP network structure diagram.

[0026] Figure 5 Structure diagram of collaborative detection head for dynamic tasks DETAILED DESCRIPTION

[0027] To more clearly describe the embodiments of the present invention, the following is an explanation in conjunction with the accompanying drawings. The present invention provides a method for detecting small targets in drone aerial photography based on DST-YOLO, and the method mainly comprises the following steps:

[0028] Step 1: Based on YOLOv11, build an improved network model and combine it with the attached Figure 1 Note that there are three sub-processes:

[0029] Step 1-1: In order to enhance the performance of feature representation and improve the model's ability to capture high-level detail features, the Dilation-wise Residual (DWR) module in the ultra-real-time semantic segmentation model DWR-Seg replaces the Bottleneck module in the C3k2 module. By introducing the dilated convolution technology, the receptive field is expanded and more contextual information is captured, thereby improving the expression of detail features. The DWR module adopts a residual design. First, a 3×3 convolution is used to generate regional residual features of different scales, and then the regional features are morphologically filtered using the dilated depth convolution branches with different dilation rates. Then, the features are fused using 1×1 point-by-point convolution. This design effectively integrates the feature information of different network layers, strengthens the connection between contextual features, and forms a richer feature representation. The formula is:

[0030] C1(x)=SiLU(BN(Conv(x)))

[0031] C2(x, d) = D d DConv(C1(x))

[0032]

[0033] Where x represents the input feature map, Conv() represents a 3×3 convolution, BN() represents a Batch Normalization layer, SiLU() represents an activation function, and D d DConv() represents a 3×3 convolution in the depth direction with a dilation rate d, and PConv() represents a point-by-point convolution. It represents the cascade operation of all d, and DWR() represents the DWR module.

[0034] Furthermore, in order to solve the problem of missing feature information, based on the above improvements, the DilatedReparam Block (DRB) module is introduced to improve the DWR module to generate the C3k2-DWRB module, whose structure is as follows: Figure 2 As shown in the figure, the DRB module first implements multi-level feature extraction and conversion operations with several convolutional layers, so that each convolutional layer can obtain feature expressions at different levels. Then, multiple connection layers are used to carry out multiple fusion processes on the features of different convolutional layers and connection layers, so as to enhance the diversity and richness of features, improve the model's ability to parse input data, and finally improve the model performance.

[0035] Step 1-2: The neck feature fusion network Slim-neck connects the two GSConv modules in sequence, and replaces the C3k2 structure in the neck network with the VoV-GSCSP module. By reducing the number of channels in the middle layer of the model, the number of parameters is reduced, thereby reducing the model complexity and the demand for storage resources.

[0036] Specifically, the GSConv module first converts the input feature map with C1 channels into a feature map with C2 / 2 channels by using standard convolution, and then generates another feature map with the same number of channels C2 / 2 by using depthwise separable convolution. The two feature maps are then merged by splicing, and then processed by channel shuffle operation to obtain an output feature map with C2 channels, thereby achieving a complete fusion of the outputs of standard convolution and depthwise separable convolution, and realizing efficient integration of feature information extracted by different convolution forms, such as Figure 3 As shown in the figure, this lightweight convolution structure effectively maintains the implicit connection between channels while having a lower time complexity. Its time complexity is:

[0037]

[0038] Where O represents the time complexity, U1·U2 is the size of the convolution kernel, W and H represent the width and height of the output feature map, respectively, C1 represents the number of channels of the input feature map and the convolution kernel, and C2 represents the number of channels of the output feature map.

[0039] The VoV-GSCSP module is a cross-stage network module constructed by a one-time aggregation method, which can efficiently fuse the feature graph information of different stages. Its structure is as follows: Figure 4 As shown in the figure. In the VoV-GSCSP network, GSConv is used to replace the traditional convolution operation, and the cascade enhancement function of two GSConv modules is used. After applying VoV-GSCSP to the neck network and replacing the original C3k2 structure, feature maps of different scales can be connected in sequence to generate extended feature vectors. This method improves the fusion effect of the backbone network in feature extraction, ensures that each layer of features can deeply capture deep semantic information and surface detail information, and effectively improves the network's feature extraction and fusion capabilities.

[0040] Step 1-3: The detection head uses the dynamic task to cooperate with the detection head to achieve the task interaction function between classification and positioning. The detection head is first processed by two shared convolution layers of group normalization (GN), and then the channel connection operation is performed to achieve dynamic interaction of image features. It is then divided into three branches: the upper branch is responsible for generating the mask and offset required for the deformable convolution (DCNv2) to enhance the network's adaptability to complex scenes and the accuracy of feature extraction; the lower branch sequentially uses a 1x1 Conv layer, a ReLU activation function, a 3×3 Conv layer, and a Sigmoid activation function to construct a dynamic feature selection module, and uses interactive features to generate dynamic feature selection weights; the middle branch is further subdivided into two parts: positioning and classification, both of which are successively decomposed by tasks (Task After that, the positioning branch uses DCNv2 and the mask and offset generated by the interactive features, combined with the convolutional bounding box to realize target positioning, and uses the scale layer (Scale) to perform scale scaling operations to deal with the situation where the scales of the targets detected by each detection head are different. The classification branch multiplies the dynamic feature selection weight with the features after task decomposition to improve the interaction ability of the two, and then completes the target classification task through the convolution classifier (Conv_Cls). With the help of the dynamic task collaborative detection head, the number of parameters can be significantly reduced, making the model more lightweight, while making full use of the interactive features of positioning and classification to effectively improve the accuracy of model detection. Figure 5 The network structure of the collaborative detection head for dynamic tasks.

[0041] Step 2: Obtain an open source dataset of small target images from the perspective of a drone.

[0042] This example uses the open source VisDrone dataset, which is a large-scale drone perspective dataset. The data is collected from 14 cities in China, covering diverse environments such as cities and rural areas, involving various types of targets such as pedestrians, vehicles, and bicycles, and includes different situations such as sparse scenes and crowded scenes. The data is collected by various drone cameras, covering a variety of environments, objects, and densities, and widely evaluates the performance of drone vision systems in different scenarios, providing important references and benchmarks for related research and applications.

[0043] Step 3: Configure the training environment required by the network model and load the network model into the configured training environment.

[0044] 6471 images were selected from the VisDrone dataset as training sample sets, 548 as validation sets, and 1610 as test sets. After completing the preprocessing of the dataset, the dataset path and the number of categories were written in the yaml file to build the configuration file corresponding to the dataset. The training environment is as follows: GPU is NVIDIA GeForce RTX 4070 Laptop GPU, CPU is AMD Ryzen 9 7940HX, programming language is Python3.9, deep learning framework is Pytorch2.0.0, GPU acceleration library is CUDA12.3. The model parameters are set as follows during training: resolution is 640×640, batch size is 16, and number of iterations is 200.

[0045] Step 4: Use the trained network model to detect the small drone aerial targets to be detected.

[0046] This embodiment illustrates the effect achieved by this invention in combination with data, as shown in Table 1. In order to further test the effect of the model, the YOLOv11n model officially provided by ultralytics was used as the benchmark model through ablation test. The experimental results are shown in Table 1. The model with the best detection effect proposed by the present invention is DST-YOLO. Compared with the original model, the value of mAP50 increased by 2.0 percentage points, the value of mAP50-90 increased by 1.3 percentage points, and the amount of calculation was reduced by 18%. It can be seen that this improvement strategy improves the accuracy of detection while reducing the model parameters to a certain extent, which can achieve a balance between model lightweight and performance.

[0047] Table 1 Experimental results of different improvement strategies

[0048]

[0049] The above is an implementation example of the invention. Modifications, equivalent replacement means or improvements based on the spirit and principles of the invention should all be included in the scope of the claims of the invention.

Claims

1. A method for detecting small targets in drone aerial photography based on DST-YOLO, characterized in that: The method comprises the following steps: Step 1: Based on YOLOv11 as the basic network, an improved network model is constructed. The C3k2 module in the YOLOv11 backbone network is replaced by the C3k2-DWRB module. The neck feature fusion network adopts the Slim-Neck structure, and the GSConv and VoV-GSCSPC modules are used to realize the neck feature fusion. The detection head uses the dynamic task collaborative detection head to achieve the interactive process between the classification task and the positioning task. Step 2: Obtain an open source dataset of small target images from the perspective of a drone. Step 3: Configure the training environment required by the network model and load the network model into the configured training environment. Step 4: Use the trained network model to detect small targets photographed by drones.

2. The method for detecting small targets in drone aerial photography based on DST-YOLO according to claim 1, characterized in that: In step 1, the C3k2-DWRB module refers to the design idea of ​​the Dilation-wiseresidual (DWR) module in the ultra-real-time semantic segmentation model DWR-Seg, and replaces the Bottleneck module in the C3k2 module to expand the receptive field, enhance the feature representation and the feature connection between network layers. At the same time, the Dilated reparam block (DRB) module is introduced to replace the convolution module in DWR to improve the fusion capability of feature information and extract rich features more efficiently.

3. The method for detecting small targets in drone aerial photography based on DST-YOLO according to claim 1, characterized in that: The neck feature fusion network Slim-neck connects two GSConv modules in sequence, and replaces the C3k2 structure in the neck network with a VoV-GSCSP module. It uses the method of reducing the number of channels in the middle layer of the model to reduce the number of parameters, thereby reducing the model complexity and the demand for storage resources.

4. The method for detecting small targets in drone aerial photography based on DST-YOLO according to claim 3 is characterized in that: The GSConv module first converts the input feature map with C1 channels into a feature map with C2 / 2 channels using standard convolution, and then generates another feature map with the same number of channels C2 / 2 using depthwise separable convolution. The two feature maps are then merged by splicing, and then processed by channel shuffle operation to obtain an output feature map with C2 channels, thereby achieving a complete fusion of the standard convolution and depthwise separable convolution outputs, and realizing efficient integration of feature information extracted by different convolution forms.

5. The method for detecting small targets in drone aerial photography based on DST-YOLO according to claim 3, characterized in that: The VoV-GCSP module first performs 1×1 ordinary convolution on the input feature map to achieve dimensionality reduction, and then divides the dimensionality reduction result into two branches. One branch contains two layers of GSConv operations, and the other branch performs a Concat operation with the output of the two layers of GSConv. Finally, the concatenated result is processed using 1×1 convolution to generate an output feature map, thereby fully exploring the feature information of the shallow network and the deep network, achieving an efficient fusion effect, and effectively improving feature utilization efficiency and information richness.

6. The method for detecting small targets in drone aerial photography based on DST-YOLO according to claim 1, characterized in that: The dynamic task collaborative detection head reduces the number of parameters by sharing convolutions, and uses a feature extractor to learn interactive features between tasks from multiple convolutional layers to generate joint features. Among them, the localization branch combines deformable convolution DCNV2 and interactive features, and the classification branch realizes dynamic feature selection based on interactive features.

7. The method for detecting small targets in drone aerial photography based on DST-YOLO according to claim 1, characterized in that: In step 3, the training environment is as follows: the GPU is NVIDIA GeForce RTX 4070 Laptop GPU, the CPU is AMD Ryzen 9 7940HX, the programming language is Python 3.9, the deep learning framework is Pytorch 2.0.0, and the GPU acceleration library is CUDA 12.

3. The model parameters during training are set as follows: the resolution is 640×640, the batch size is 16, and the number of iterations is 200.

Citation Information

Cited By

  • Graphite ore image segmentation method based on improved YOLO11-seg model

    CN120782783A

  • Road signboard identification method and system based on deep learning

    CN121033805A

  • Unmanned aerial vehicle target detection method and system based on YOLOv12

    CN121330561A

  • A method and system for detecting a target of a UAV based on YOLOv12

    CN121330561B

  • An aerial multi-view small target detection method and system based on improved YOLO

    CN122799083A