Unmanned aerial vehicle small target detection method

By improving the YOLOv5s model, deformable convolution module, C3-DWR module, weighted bidirectional feature pyramid module and P2 detection head were introduced, and the SIoU loss function was used to solve the missed detection and misdetection problems in the detection of small targets of drones, improving detection accuracy and reliability.

CN120182864APending Publication Date: 2025-06-20ZHEJIANG NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510246247.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

There are problems of missed detection and misdetection in small target detection, resulting in insufficient detection accuracy and reliability, affecting its efficient application.

Method used

By improving the YOLOv5s model, the BVD-YOLO model is obtained, and the deformable convolution module, C3-DWR module, weighted bidirectional feature pyramid module and P2 detection head are used, and the SIoU loss function based on the dual-frame vector angle is used to improve the model's detection performance for small targets.

Benefits of technology

It improves the accuracy and reliability of the drone's detection of small targets, achieves a good balance between computing resources and detection performance, and can efficiently complete the drone's small target detection task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182864A_ABST
    Figure CN120182864A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle small target detection method, and relates to the field of unmanned aerial vehicle small target detection, and the method comprises the steps: obtaining an image collected by an unmanned aerial vehicle, and building an unmanned aerial vehicle detection data set; preprocessing the images in the detection data set, wherein the preprocessing comprises screening available data and dividing the available data into a training set, a test set and a verification set; the YOLOv5s model is improved, and an improved BVD-YOLO model is obtained; carrying out iterative training on the BVD-YOLO model by utilizing the training set; testing the trained BVD-YOLO model by using the test set, and performing validity verification by using the verification set; and performing unmanned aerial vehicle small target detection by using the finally determined BVD-YOLO model. According to the invention, the accuracy and reliability of small target detection by the unmanned aerial vehicle are improved, and the small target detection task of the unmanned aerial vehicle can be completed more efficiently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of small target detection for unmanned aerial vehicles, and more specifically, to a method for small target detection of unmanned aerial vehicles. Background Art

[0002] With the rapid development of unmanned aerial vehicle detection technology, unmanned aerial vehicles have been widely used in many fields such as precision agriculture, military reconnaissance, and traffic monitoring. With the advantages of high efficiency, safety, and flexibility, unmanned aerial vehicles can achieve comprehensive and real-time image acquisition at high altitudes by carrying small high-definition cameras and various sensors. In the field of target detection, traditional machine learning algorithms rely on manually designed image features, which is time-consuming and laborious. Especially when facing targets with similar features, it is difficult to define an accurate standard to distinguish them. With the improvement of computer performance, deep learning algorithms that can autonomously learn image features have gradually become the mainstream. At present, deep learning algorithms are mainly divided into two-stage and one-stage learning algorithms. The two-stage algorithm first generates candidate boxes, and then uses non-maximum suppression technology to accurately screen the most accurate prediction boxes from the candidate boxes to achieve effective detection of real objects. The one-stage algorithm regards the processing process as a single regression problem, omitting the candidate box generation stage. Although the detection accuracy is slightly reduced, the detection efficiency has been significantly improved, which can meet the real-time detection requirements of target detection. Its main representatives include YOLO series algorithms and SSD algorithms, etc.

[0003] However, during the actual shooting process of unmanned aerial vehicles, they will be affected by factors such as light intensity, wind force, and height difference, resulting in a decrease in image resolution. In addition, vertical shooting at high altitudes makes the images mainly composed of small targets. The problems of missed detection and false detection of unmanned aerial vehicles in small target detection have become the key factors restricting their efficient application.

[0004] Therefore, how to improve the accuracy and reliability of unmanned aerial vehicle detection of small targets is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a method for small target detection of unmanned aerial vehicles, aiming to further improve the detection performance of unmanned aerial vehicle monitoring for small targets.

[0006] In order to achieve the above object, the present invention provides the following technical solutions:

[0007] The present invention discloses a method for small target detection of unmanned aerial vehicles, including:

[0008] Obtain the images collected by the unmanned aerial vehicle and establish an unmanned aerial vehicle detection data set;

[0009] Preprocess the images in the detection data set, and the preprocessing includes screening available data and dividing it into a training set, a test set, and a validation set;

[0010] Improve the YOLOv5s model to obtain the improved BVD - YOLO model;

[0011] Use the training set to iteratively train the BVD - YOLO model; for the trained BVD - YOLO model, use the test set for testing and the validation set for effectiveness verification;

[0012] Use the finally determined BVD - YOLO model for small target detection of drones.

[0013] Furthermore, the improvement of the YOLOv5s model includes:

[0014] Introduce a deformable convolution module into the backbone network of the YOLOv5s model;

[0015] Introduce a C3 - DWR module into the neck and backbone network of the YOLOv5s model;

[0016] Introduce a weighted bidirectional feature pyramid module into the neck of the YOLOv5s model;

[0017] Add a P2 detection head to the head of the YOLOv5s model;

[0018] Replace the loss function of the bounding box of the YOLOv5s model with the SIoU loss function based on the angle of the double - box vector.

[0019] Furthermore, the deformable convolution module includes a first two - dimensional convolutional layer, a resampling layer, a tensor transformation layer, a second two - dimensional convolutional layer, a normalization layer, and an activation function layer connected in sequence;

[0020] The first two - dimensional convolutional layer performs a two - dimensional convolution operation on the input feature map to generate an offset parameter tensor θ of size (B, 2×N, H, W);

[0021] The resampling layer performs bilinear interpolation and resampling operations on the offset parameter tensor θ to obtain a tensor X for final sampling θ , X θ with a size of (B, C, H, W, N);

[0022] The tensor X θ , respectively, undergoes convolution operations, normalization operations, and activation operations of the second two - dimensional convolutional layer, the normalization layer, and the activation function layer to obtain the output feature map of the deformable convolution module.

[0023] Furthermore, the coordinate transformation formula of the deformable convolution module is:

[0024]

[0025] The formula for bilinear interpolation is as follows:

[0026]

[0027] where P0 is the current position coordinate, and P n is the initial sampling coordinate of the deformable convolution module; is the set of original coordinates, is the set of modified sampling coordinates, p is any point in, is the set of coordinates of the four adjacent points adjacent to the p point, and q is any point in; max is the maximum value function; p x and p y are the x and y axis coordinates of the p point respectively, and q x and q y are the x and y axis coordinates of the q point respectively, and y(p) and y(q) are the input image pixel values of the p point and the q point respectively.

[0028] Furthermore, the C3-DWR module is obtained by integrating the DWR module into the C3 module. The DWR module includes a 3×3 convolutional layer, a batch normalization and activation function layer, a depthwise separable convolutional layer, a batch normalization layer, a 1×1 convolutional layer, and a residual structure connected in sequence;

[0029] The 3×3 convolutional layer convolves the input feature map to generate a simple feature map of three regions; and then performs a non-linear transformation through the batch normalization and activation function layer to complete the feature regionalization process;

[0030] The depthwise separable convolutional layer respectively uses depthwise separable convolutions with three different dilation rates to perform morphological filtering on the three-region features of different sizes to obtain multi-scale features;

[0031] The multi-scale features respectively undergo batch normalization, convolution, and splicing operations of the batch normalization layer, the 1×1 convolutional layer, and the residual structure to output a feature map.

[0032] Furthermore, the calculation formula of the weighted bidirectional feature pyramid module is as follows:

[0033]

[0034] where F i in and F i td and F i outrespectively represent the input feature map, intermediate feature map, and output feature map of the i-th layer; Conv represents the convolution operation, and Resize represents the upsampling or downsampling process of the feature map; ρ = 0.0001 is used to ensure numerical stability; w i represents the learning weight from the input feature map to the intermediate feature map of the i-th layer, and w i ′ represents the learning weight from the intermediate feature map to the output feature map of the i-th layer.

[0035] Further, the specific method of adding the P2 detection head is: based on the head structure of the YOLOv5s model, add a P2 detection head for the feature map with a size of 160×160.

[0036] Further, for the SIoU loss function based on the angle of the double-box vector, the calculation formula is:

[0037]

[0038] Among them, L SIoU represents the value of the loss function; Δ represents the distance loss; Ω represents the shape loss; IoU represents the IoU loss, and the calculation formula is: Among them, B and B GT respectively represent the predicted bounding box and the ground truth bounding box.

[0039] Further, the calculation formula of the shape loss is:

[0040]

[0041]

[0042] Among them, ω w and ω h are the determinants corresponding to the length and width of the predicted bounding box respectively, H GT and W GT respectively represent the length and width of the ground truth bounding box; H and W respectively represent the length and width of the predicted bounding box; max is the maximum value function; λ is the attention factor controlling the shape loss; e is the natural constant, and t is an intermediate parameter;

[0043] The calculation formula of the distance loss is:

[0044]

[0045] Among them, ρ x and ρ y are the determinants related to the horizontal distance and vertical distance of the centers of the ground truth bounding box and the predicted bounding box respectively, γ is a determinant, and γ = 2 - Λ; and are the x and y coordinate values of the center point of the ground truth bounding box respectively, B x and By They are the x and y coordinate values of the center point of the prediction box; C' h and C' w They are the lengths and widths of the minimum circumscribed rectangles of the ground truth box and the prediction box respectively; Λ is the angular loss, and the specific formula is:

[0046]

[0047] where C h is the distance between the center points of the ground truth box and the prediction box in the vertical direction, and σ is the distance between the center points of the ground truth box and the prediction box.

[0048] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method for detecting small targets of drones. By introducing deformable convolutions (AKConv) into the existing YOLOv5s model to replace the convolution modules of part of the backbone network, while enhancing the detection performance, the model complexity is reduced; the C3-DWR module is introduced into the neck and the backbone network to improve the model's recognition ability for overlapping targets at a small computational cost; the weighted bidirectional feature pyramid module (BiFPN) is introduced into the neck to achieve deeper feature fusion; an additional P2 detection head is added to enhance the model's detection ability for small targets; the SIoU loss function based on the angle of the double-box vector is used to make the regression processing of the bounding box more accurate. The present invention improves the accuracy and reliability of drone small target detection, can more efficiently complete the drone small target detection task, and achieves a better balance between computing resources and detection performance. Description of the Drawings

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0050] Figure 1 It is the overall flow schematic diagram of the embodiment of the present invention.

[0051] Figure 2 It is the schematic diagram of the BVD-YOLO model structure of the embodiment of the present invention.

[0052] Figure 3 It is the schematic diagram of the AKConv module process of the embodiment of the present invention.

[0053] Figure 4 It is the schematic diagram of the C3-DWR module structure of the embodiment of the present invention.

[0054] Figure 5Schematic diagrams of the FPN, PANet, and BiFPN structures according to embodiments of the present invention.

[0055] Figure 6 Schematic diagrams of the angular loss, distance loss, and shape loss of the SIoU loss function according to embodiments of the present invention.

[0056] Figure 7 Comparison diagram of detection results under drone detection according to embodiments of the present invention. Detailed implementation manners

[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0058] Embodiments of the present invention disclose a method for detecting small targets of drones, as Figure 1 shown, including:

[0059] Obtain images collected by the drone and establish a drone detection data set;

[0060] Preprocess the images in the detection data set. The preprocessing includes screening available data and dividing it into a training set, a test set, and a validation set;

[0061] Improve the YOLOv5s model to obtain an improved BVD-YOLO model;

[0062] Iteratively train the BVD-YOLO model using the training set; for the trained BVD-YOLO model, test it using the test set and verify its effectiveness using the validation set;

[0063] Use the finally determined BVD-YOLO model to detect small targets of drones.

[0064] In a specific embodiment, the YOLOv5s model is improved, and the improved network structure is as Figure 2 shown, including:

[0065] Introduce a deformable convolution module into the backbone network of the YOLOv5s model;

[0066] Introduce a C3-DWR module into the neck and backbone network of the YOLOv5s model;

[0067] Introduce a weighted bidirectional feature pyramid module into the neck of the YOLOv5s model;

[0068] Add a P2 detection head to the head of the YOLOv5s model;

[0069] Replace the loss function of the bounding box of the YOLOv5s model with the SIoU loss function based on the angle of the double-box vector.

[0070] In a specific embodiment, the deformable convolution module includes a first two-dimensional convolutional layer, a resampling layer, a tensor transformation layer, a second two-dimensional convolutional layer, a normalization layer, and an activation function layer connected in sequence;

[0071] The first two-dimensional convolutional layer performs a two-dimensional convolution operation on the input feature map to generate an offset parameter tensor θ of size (B, 2×N, H, W);

[0072] The resampling layer performs bilinear interpolation and resampling operations on the offset parameter tensor θ to obtain a tensor X for final sampling θ , X θ has a size of (B, C, H, W, N);

[0073] The tensor X θ , respectively undergoes convolution operations, normalization operations, and activation operations of the second two-dimensional convolutional layer, the normalization layer, and the activation function layer to obtain the output feature map of the deformable convolution module.

[0074] Specifically, standard convolution plays an important role in traditional image processing. It can effectively extract local information of images through a convolution kernel of a fixed size. However, in the face of the complex and changeable shooting environment of drones, the distribution of effective information of detection targets often varies significantly. Therefore, standard convolution is limited by its fixed sampling position and convolution kernel size. To overcome this defect, the deformable convolution module (AKConv) introduces the concept of deformable convolution and proposes flexible adjustment of sampling positions and convolution kernel sizes. The process of AKConv is as Figure 3 shown, where C is the number of channels; Conv2d is a two-dimensional convolution operation; θ is a learnable offset parameter tensor; B is the batch-size; N is the number of sampling points in the convolution kernel; H and W are the height and width of the input image respectively; Resample represents a resampling operation; X θ is a tensor for final adoption; Reshape represents a tensor transformation operation; Norm is a normalization operation; SiLU is an activation function.

[0075] Before performing the task, customize the initial convolution kernel size and sampling shape of AKConv according to the image features; then AKConv generates an offset parameter tensor θ of size (B, 2×N, H, W) through a two-dimensional convolution operation. Then, by adding θ to the original coordinate set to obtain the modified sampling coordinate set Due to the instability of the coordinates after offset, AKConv obtains the feature sampling tensor X at the corresponding position through bilinear interpolation and resampling operations. θ The size is (B, C, H, W, N). Finally, through a series of operations such as tensor transformation, the image features in X θ are obtained as the output.

[0076] In a specific embodiment, the coordinate transformation formula of the deformable convolution module is:

[0077]

[0078] The formula for bilinear interpolation is:

[0079]

[0080] Among them, P0 is the current position coordinate, and P n is the initial sampling coordinate of the deformable convolution module; is the set of original coordinates, is the set of modified sampling coordinates, p is any point in, is the set of coordinates of the four adjacent points adjacent to the p point, and q is any point in; max is the maximum value function; p x , p y are the x and y axis coordinates of the p point respectively, and q x , q y are the x and y axis coordinates of the q point respectively, and y(p), y(q) are the pixel values of the input images at the p point and q point respectively.

[0081] Specifically, in the standard 3x3Conv operation, P n is a 3x3 tensor. In the convolution operation, after adding it to the current position P0, the set of original sampling coordinates is obtained. However, AKConv is different from the standard convolution. Because of the uncertain number of initial sampling points (i.e., not necessarily 3x3, it may be 3 or 5), so P n uniformly determines the upper left corner coordinate as the center coordinate.

[0082] The main purpose of the bilinear interpolation formula is: Since the offset parameter tensor θ is to be added to the set of original coordinates , there will be a situation where decimal points appear in the coordinate set (i.e., instability). Therefore, the p point is not necessarily an integer coordinate, and there is no definite pixel value for non-integer coordinates. For this reason, the bilinear interpolation formula estimates the pixel value of the p point through the four integer coordinate points adjacent to the p point. At this time, the p point has not undergone convolution processing, so it is not the eigenvalue of the feature map, and its estimated value result will be stored in X θand perform convolution processing to obtain the eigenvalues of the feature map.

[0083] In a specific embodiment, the C3-DWR module is obtained by integrating the DWR module into the C3 module. The DWR module includes a 3×3 convolutional layer, a batch normalization and activation function layer, a depthwise separable convolutional layer, a batch normalization layer, a 1×1 convolutional layer, and a residual structure connected in sequence;

[0084] The 3×3 convolutional layer convolves the input feature map to generate a simple feature map of three regions; then, through the batch normalization and activation function layer, a non-linear transformation is performed to complete the feature regionalization process;

[0085] The depthwise separable convolutional layer uses depthwise separable convolutions with three different dilation rates to perform morphological filtering on the three-region features of different sizes to obtain multi-scale features;

[0086] The multi-scale features are respectively subjected to batch normalization, convolution, and splicing operations of the batch normalization layer, the 1×1 convolutional layer, and the residual structure, and the feature map is output.

[0087] Specifically, aiming at the problem of missed detection of dense crowds by drones, the present invention introduces a feature fusion module (DWR) based on multi-scale dilation rate convolution and integrates it into the C3 module. The structural diagram of the DWR module is as Figure 4 shown, where DConv represents depthwise separable convolution; D-n represents dilation by n times; the ‘+’ inside the circle represents the addition operation; c represents the base number of the feature map channels; RR and SR respectively represent the two major steps of regional residualization and semantic residualization. From Figure 4 it can be seen that DWR first performs regional residualization, and the input feature map is convolved through a 3×3 convolution to generate a simple feature map with three concise regional forms; then, combined with the non-linear transformation of batch normalization BN and the ReLU activation function, the feature regionalization process is completed. Subsequently, it enters the semantic residualization stage. In this stage, depthwise separable convolutions with three different dilation rates are respectively used to perform morphological filtering on the regional features of different sizes, so as to efficiently and accurately capture multi-scale features while expanding the receptive field. Finally, through the 1×1 convolution and the residual structure, an output feature with richer semantics is obtained. The formula of the DWR module is:

[0088]

[0089] where is the input feature map with 2c channels; Conv 3×3 represents a convolution operation of 3×3 size; BR represents first performing batch normalization operation and then activating through the ReLU activation function; R1, R2, R3 represent regional residualization operations, is the feature map corresponding to the regional residualization operation; S1, S2, and S3 represent the semantic residualization operations, is the feature map corresponding to the semantic residualization operation; respectively represent depthwise separable convolution operations with dilation factors of 1, 3, and 5 and a convolution kernel size of 3×3; Conv 1×1 represents a convolution operation with a size of 1×1, BN represents batch normalization operation, Concat represents channel concatenation operation, is the output feature map with 2c channels.

[0090] The DWR module cleverly decomposes the multi-scale feature extraction method into two stages: regional residualization and semantic residualization. In this design, the role of the depthwise separable convolution has changed. Instead of striving to obtain as complex semantic information as possible, it simply performs morphological filtering on each concisely expressed feature map using the desired receptive field. This change greatly simplifies the difficulty of extracting multi-scale context information, thereby improving the detection efficiency; at the same time, in the semantic residualization stage, the depthwise separable convolution enhances the recognition of overlapping objects by expanding the receptive field, ultimately effectively improving the missed detection problem in dense images.

[0091] In a specific embodiment, in order to improve the detection performance of the model for multi-scale objects, YOLOv5 introduces the traditional Feature Pyramid Network (FPN) and Path Aggregation Network (PANet). The former enhances the information transmission between feature maps through the top-down propagation path and lateral connections, enabling the network to effectively fuse the spatial detail information of the shallow layer and the semantic information of the deep layer, thereby greatly improving the detection accuracy; the latter adds an additional bottom-up path aggregation network on the basis of the former, breaking the limitation of the one-way propagation information flow of FPN. The Bidirectional Feature Pyramid Network (BiFPN) introduced in the present invention is precisely based on the essence of FPN and PANet, and realizes a more efficient integration of shallow detail information and deep semantic information by streamlining redundant components and adding repeated modules. The structural schematic diagrams of the above three feature fusion methods are as Figure 5 shown (a, b, and c in the figure respectively represent the three network structures of FPN, PANet, and BiFPN). In the figure, compared with PANet, BiFPN deletes nodes 1 and 2 with only one input edge, and additionally adds connection operations to the input and output nodes of the same layer, achieving a richer feature fusion effect at a lower cost. In addition, a repeatable module is proposed to achieve deeper feature fusion.

[0092] Traditional feature fusion often uses resampling techniques to achieve the superposition of feature maps of different scales. The formula is:

[0093]

[0094] Among them, Resize represents the use of upsampling or downsampling processing on the feature map; Conv represents the convolution operation, and F i in and F i out respectively represent the input and output feature maps of the i-th layer feature map.

[0095] In contrast, BiFPN adopts a fast normalization fusion algorithm to calculate the output feature map O, and the formula is:

[0096]

[0097] Among them, ρ = 0.0001 is used to ensure the numerical stability. Through the weighted method, BiFPN can automatically determine the weight distribution of feature maps at different scales during the fusion process, so that the model can flexibly optimize the feature fusion strategy according to different datasets and task requirements, avoiding the fixed weighted method from restricting the performance of the model. Combining Figure 5 with part c, we get the final calculation formula of the weighted bidirectional feature pyramid module as:

[0098]

[0099] Among them, F i in 、F i td and F i out respectively represent the input feature map, intermediate feature map, and output feature map of the i-th layer; Conv represents the convolution operation, and Resize represents the use of upsampling or downsampling processing on the feature map; ρ = 0.0001 is used to ensure the numerical stability; w i represents the learning weight from the input feature map of the i-th layer to the intermediate feature map, and w i ′ represents the learning weight from the intermediate feature map of the i-th layer to the output feature map.

[0100] In a specific embodiment, adding the P2 detection head specifically means: based on the head structure of the YOLOv5s model, adding a P2 detection head for the 160×160-sized feature map. To address the false detection problem existing in the original model, an additional P2 detection head for the 160×160-sized feature map is added to improve the detection accuracy of the model for small targets. Specifically, there is richer semantic information and spatial information on the 160×160 feature map for small targets, so it helps the model accurately identify small features and can effectively improve the false detection problem of small targets.

[0101] In a specific embodiment, in YOLOv5, in order to improve the detection accuracy of bounding boxes, the CIoU loss function is adopted. Although CIoU has successfully helped YOLOv5 achieve high detection performance, it still has two major defects. First, when the aspect ratios of the predicted box and the ground truth box are the same, the penalty term for the aspect ratio is always 0, which leads to the irrationality of the loss function setting; second, CIoU fails to consider the mismatch direction between the ground truth box and the predicted box, which will lead to a slowdown in the convergence speed and a decrease in the convergence effectiveness. To address the above problems, the present invention proposes to replace the CIoU loss function with the SIoU loss function, which converges faster and has higher accuracy, and redefines the bounding box loss function of the model as consisting of four parts: angular loss, distance loss, shape loss, and IoU loss.

[0102] The SIoU loss function based on the angle of the double-box vector has the following calculation formula:

[0103]

[0104] where, L SIoU represents the value of the loss function; Δ represents the distance loss; Ω represents the shape loss; IoU represents the IoU loss, and its calculation formula is: where, B and B GT represent the predicted box and the ground truth box respectively.

[0105] The calculation formula for the shape loss is:

[0106]

[0107] where, ω w , ω h are the determining factors corresponding to the length and width of the predicted box respectively; as shown in part b of Figure 6 , H GT and W GT represent the length and width of the ground truth box respectively, and H and W represent the length and width of the predicted box respectively; max is the maximum value function; λ is to control the attention of the shape loss, and its range is generally between 2 and 6; e is the natural constant, t is an intermediate parameter, and t is the representation form in the operation and does not participate in the calculation. As shown in , ω t respectively refer to ω w or ω h .

[0108] As shown in part a of Figure 6 , where C' h and C' w are the length and width of the minimum circumscribed rectangle of the two boxes respectively. The calculation formula for the distance loss is:

[0109]

[0110]

[0111] Among them, ρ x and ρ y are respectively the determinants related to the horizontal and vertical distances between the center points of the ground truth box and the predicted box, γ is the determinant, and γ = 2 - Λ; and are respectively the x and y coordinate values of the center point of the ground truth box, and B x and B y are respectively the x and y coordinate values of the center point of the predicted box; Λ is the angular loss, as shown in part a of Figure 6 , where BGT represents the ground truth box, B is the predicted box, and both α and β are optional calculation angles. The specific formula is:

[0112]

[0113] Among them, C h is the vertical distance between the center points of the ground truth box and the predicted box, and σ is the distance between the center points of the ground truth box and the predicted box.

[0114] IoU, as the ratio of the intersection area of two boxes to the common area, is the most classic model evaluation criterion; Δ and Ω are the distance loss and shape loss respectively. Introducing the distance loss can avoid the divergence problem of the bounding box during the training process, enabling the predicted box to be more accurately located at the target center. The shape loss provides more detailed control in optimizing the aspect ratio and shape convergence of the target box, and can further improve the accuracy and robustness of the model in the object detection task. The SIoU loss function combines the above four parts, bringing a faster convergence speed and higher regression accuracy to the bounding box, effectively enhancing the detection performance and real-time detection rate of the model.

[0115] In a specific embodiment, through specific experiments, target detection is performed on the images collected by the drone to verify the beneficial effects of the method of the present invention. Specifically:

[0116] Experimental environment: The experimental environment configuration is Windows 10 operating system, equipped with an Intel(R) Core(TM) i5-10400F CPU @ 2.90GHz processor, and the NVIDIA GeForce RTX 2080Ti GPU is used to accelerate the training and inference processes. The CUDA version is selected as 12.6, and the deep learning framework is Pytorch, with the version 2.4.0. The selected compiler and compilation language are Pycharm 2022 and Python 3.8.19 respectively.

[0117] Dataset: The present invention uses the VisDrone2019 dataset publicly available from the AI SKYEYE team of Tianjin University, which contains 8,575 static images collected by drones, covering various scenarios such as urban roads, highways, scenic area crowds, and county town markets. In this experiment, the dataset is divided into a training set of 6,417 images, a validation set of 548 images, and a test set of 1,610 images, involving a total of ten categories such as pedestrians, people, bicycles, and tricycles. The VisDrone2019 dataset performs excellently in terms of target diversity and the challenge of small target detection, providing rich real-world environment samples and a reliable evaluation basis for the algorithm model.

[0118] To demonstrate the excellent performance of BVD-YOLO, the present invention conducts experimental comparative analysis using classical target detection metrics and mainstream YOLO series models as shown in Table 1. The YOLOv5s model as the baseline achieves detection accuracies of 0.283 and 0.152 in mAP@0.5 and mAP@0.5-0.95 respectively. Compared with YOLOv5n, the improvements are 20.9% and 29.9% respectively. It can be seen that the increase in model size is beneficial to enhancing the detection performance of the model. Comparing the YOLOv6s model with the YOLOv5s model, the former improves by 7.4% and 15.8% respectively in mAP@0.5 and mAP@0.5-0.95, indicating that the improvement of YOLOv6 effectively enhances the detection performance of the model for small targets, but the substantial increase in the number of parameters is not conducive to the simple deployment of drones. The experimental results of YOLOv7 reach 0.404 and 0.22 respectively in mAP@0.5 and mAP@0.5-0.95, but its huge computational parameters limit its feasibility in real-time detection and are not suitable for the real-time detection task of drones. The YOLOv8 series of algorithms improves the C2f module and the decoupled head on the basis of the YOLOv5 model algorithm, effectively improving the accuracy while appropriately increasing the computational parameters, indicating its performance in such tasks.

[0119] Table 1 Comparative experimental results of BVD-YOLO and other YOLO models

[0120]

[0121]

[0122] To further prove the superiority of the BVD-YOLO model, BVD-YOLO was compared with the current mainstream object detection models, and the experimental results are shown in Table 2. Compared with BVD-YOLO, the typical two-stage algorithms, Faster-RCNN, Mask-RCNN, and Cascade-RCNN, lagged behind by 2.8%, 1.2%, and 10.7% respectively in mAP@0.5. The single-stage algorithm representative SSD512 has more excellent model size and detection speed compared with other mainstream algorithms, but its detection effect is poor, and it only achieved a detection accuracy of 0.280 in mAP@0.5, unable to achieve accurate identification of the target. Comparing with the anchor-free type detection algorithms Center-Net and FCOS, BVD-YOLO achieved significant improvements of 4.1% and 20% in mAP@0.5. Generally speaking, BVD-YOLO achieved a significant lead in detection accuracy with smaller parameter quantity and model complexity, confirming the superiority of the algorithm model of the present invention.

[0123] Table 2 Comparison between BVD-YOLO and current mainstream object detection models

[0124]

[0125] To verify the practical significance of the P2 detection head, this experiment used two additional fourth detection heads to conduct a comparative analysis with the original model, and the experimental results are shown in Table 3. Among them, YOLOv5s-P6 and YOLOv5s-P2 respectively represent adding a detection head on the P6 feature map (10×10) and the P2 feature map (160×160). Comparing YOLOv5s-P6 with the original model YOLOv5s, the addition of the P6 detection head led to a decrease in accuracy. The reason may be that the P6 feature Figure 1 is generally suitable for detecting features of larger-sized objects, causing the model to be more biased towards the extraction of large-scale features, thereby reducing the model's attention to the recognition of small targets, and thus greatly reducing the detection accuracy. While the P2 detection head added on the P2 feature map successfully helped the baseline model improve by 11.7% and 13.2% respectively in mAP@0.5 and mAP@0.5-0.95, significantly enhancing the detection performance and robustness of the model.

[0126] Table 3 Comparative experiments on adding different four-head detection heads

[0127]

[0128] In the BVD-YOLO model, the AKConv and C3-DWR modules are introduced in the Backbone and Head parts respectively. To verify their improvement effects on the model performance, in this invention, other modules are replaced at the Backbone and Head positions for comparison, and the relevant results are shown in Table 4. The experimental results show that AKConv not only improves the model performance, but also reduces the model complexity to a certain extent; while the C3-DWR module significantly improves the performance by introducing a small number of model parameters. Compared with the experimental effects of other advanced algorithms, the combination of AKConv and C3-DWR modules shows better advantages in performance.

[0129] Table 4 Comparative experiments on separately improving the Backbone and Head

[0130]

[0131] To confirm the effectiveness of the SIoU loss function for the model, this invention compares the impacts of different IoU loss functions on the overall model performance, as shown in Table 5. For the CIoU loss function used in the original model, the detection accuracies of mAP@0.5 and mAP@0.5-0.95 reach 0.327 and 0.178 respectively. Compared with different variants of the IoU (intersection over union) loss function, SIoU performs excellently in the overall performance, achieving detection accuracies of 0.330 and 0.180 on mAP@0.5 and mAP@0.5-0.95. The experiments show that the SIoU loss function is more adaptable to the complex environment of such tasks and helps to improve the detection accuracy of the model.

[0132] Table 5 Comparative experiments on IoU loss functions

[0133]

[0134] Through ablation experiments with different improved modules, to verify the outstanding contributions of each improved module to the BVD-YOLO model, the results are shown in Table 6. It can be seen from the table that YOLOv5s as the baseline model has detection accuracies of 0.283 and 0.152 on mAP@0.5 and mAP@0.5-0.95 respectively; first, the P2 detection head is introduced, and the improved model achieves significant improvements of 11.7% and 13.2% on mAP@0.5 and mAP@0.5-0.95; subsequently, the C3-DWR, AKConv, BiFPN, and SIoU modules are introduced respectively to obtain the final model BVD-YOLO, which reaches detection accuracies of 0.330 and 0.180 on mAP@0.5 and mAP@0.5-0.95. There is a significant improvement in performance compared with the baseline model.

[0135] Table 6 Ablation Experiment of BVD-YOLO

[0136]

[0137]

[0138] As Figure 7 shown, Figure 7 shows the visual comparison between YOLOv5s and BVD-YOLO algorithms in the detection task. From the effect diagrams of (a) and (b) in the comparison Figure 7 , it can be seen that there are still a large number of missed detections and false detections in image detection by YOLOv5s, and the detection accuracy is relatively low. In contrast, BVD-YOLO not only improves the confidence index of the detected objects in complex environments such as darkness, crowds, roads, and dense areas, but also improves the problems of missed detections and false detections of dense crowds and small targets, and effectively realizes the real-time detection of small UAV targets in various complex scenarios with more powerful detection performance.

[0139] Summary: Aiming at the problems of missed detections, false detections and low detection accuracy in small UAV target detection, the present invention proposes the BVD-YOLO algorithm improved based on YOLOv5. Specifically, deformable convolutions (AKConv) are introduced to replace some convolutional modules of the backbone network, which enhances the detection performance while reducing the model complexity; the C3-DWR module is introduced into the neck and the backbone network to improve the model's recognition ability of overlapping targets at a small computational cost; the weighted bidirectional feature pyramid module (BiFPN) is introduced into the neck to achieve deeper feature fusion; an additional P2 detection head is added to enhance the model's detection ability for small targets; the SIoU loss function based on the angle of the double-box vector is used to make the regression processing of the bounding box more accurate. The experimental results show that BVD-YOLO achieves a comprehensive detection accuracy of 0.330 mAP@0.5 and 0.180 mAP@0.5-0.95 on the Visdrone2019 dataset, which is 16.6% and 18.4% higher than that of the baseline model YOLOv5. Compared with mainstream object detection algorithms and YOLO series models, BVD-YOLO can complete the small UAV target detection task more efficiently, achieving a better balance between computational resources and detection performance. The BVD-YOLO algorithm can more efficiently realize the small UAV target detection task.

[0140] In this specification, each embodiment is described in a progressive manner. The key points of each embodiment are the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0141] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for detecting small targets of unmanned aerial vehicles, characterized in that: include: Obtain images collected by drones and build a drone detection dataset; Preprocessing the images in the detection data set, wherein the preprocessing includes screening available data and dividing the data into a training set, a test set, and a validation set; Improve the YOLOv5s model to obtain the improved BVD-YOLO model; Iteratively training the BVD-YOLO model using the training set; The trained BVD-YOLO model is tested using the test set, and the validity is verified using the validation set; The finalized BVD-YOLO model is used to detect small drone targets.

2. A method for detecting small targets of unmanned aerial vehicles according to claim 1, characterized in that: The improvements to the YOLOv5s model include: Introducing a deformable convolution module into the backbone network of the YOLOv5s model; Introduce C3-DWR module in the neck and backbone network of YOLOv5s model; Introduce a weighted bidirectional feature pyramid module at the neck of the YOLOv5s model; Add the P2 detection head to the head of the YOLOv5s model; The loss function of the bounding box of the YOLOv5s model is replaced by the SIoU loss function based on the double-box vector angle.

3. A method for detecting small targets of unmanned aerial vehicles according to claim 2, characterized in that: The deformable convolution module includes a first two-dimensional convolution layer, a resampling layer, a tensor transformation layer, a second two-dimensional convolution layer, a normalization layer, and an activation function layer connected in sequence; The first two-dimensional convolution layer performs a two-dimensional convolution operation on the input feature map to generate an offset parameter tensor θ of size (B, 2×N, H, W); The resampling layer performs bilinear interpolation and resampling operations on the offset parameter tensor θ to obtain a tensor X for final sampling θ , X θ The size is (B, C, H, W, N); The tensor X θ , through the convolution operation, normalization operation and activation operation of the second two-dimensional convolution layer, the normalization layer and the activation function layer respectively, an output feature map of the deformable convolution module is obtained.

4. A method for detecting small targets of unmanned aerial vehicles according to claim 3, characterized in that: The coordinate transformation formula of the deformable convolution module is: The formula for the bilinear interpolation is: Among them, P0 is the current position coordinate, P n is the initial sampling coordinate of the deformable convolution module; is the original coordinate set, is the modified sampling coordinate set, p is Any point in is the coordinate set of four adjacent points near point p, and q is Any point in; max is the maximum value function; p x 、p y are the x- and y-axis coordinates of point p, respectively, and q x ,q y are the x-axis and y-axis coordinates of point q, and y(p) and y(q) are the pixel values ​​of the input images at point p and point q respectively.

5. The method for detecting small targets of unmanned aerial vehicles according to claim 2, characterized in that: The C3-DWR module is obtained by integrating the DWR module into the C3 module, wherein the DWR module includes a 3×3 convolutional layer, a batch normalization and activation function layer, a depthwise separable convolutional layer, a batch normalization layer, a 1×1 convolutional layer and a residual structure connected in sequence; The 3×3 convolution layer convolves the input feature map to generate simple feature maps of three regions; and then performs nonlinear transformation through the batch normalization and activation function layer to complete feature regionalization processing; The depthwise separable convolution layer uses three depthwise separable convolutions with different dilation ratios to perform morphological filtering on three regional features of different sizes to obtain multi-scale features; The multi-scale features are respectively subjected to batch normalization, convolution and concatenation operations of the batch normalization layer, the 1×1 convolution layer and the residual structure to output a feature map.

6. A method for detecting small targets of unmanned aerial vehicles according to claim 2, characterized in that: The calculation formula of the weighted bidirectional feature pyramid module is: Among them, F i in 、F i td and F i out They represent the input feature map, intermediate feature map, and output feature map of layer i respectively; Conv represents the convolution operation, and Resize represents the upsampling or downsampling of the feature map; ρ = 0.0001 is used to ensure the stability of the value; w i represents the learning weight from the input feature map of the i-th layer to the intermediate feature map, w i ′ represents the learning weight from the intermediate feature map of the i-th layer to the output feature map.

7. The method for detecting small targets of unmanned aerial vehicles according to claim 2, characterized in that: The specific method of adding a P2 detection head is as follows: based on the head structure of the YOLOv5s model, a P2 detection head for a 160×160 feature map is added.

8. The method for detecting small targets of unmanned aerial vehicles according to claim 2, characterized in that: The SIoU loss function based on the double-frame vector angle is calculated as follows: Among them, L SIoU Represents the loss function value; Δ represents the distance loss; Ω represents the shape loss; IoU represents the IoU loss, and the calculation formula is: Among them, B, B GT Represent the predicted box and the true box respectively.

9. A method for detecting small targets of unmanned aerial vehicles according to claim 8, characterized in that: The calculation formula of the shape loss is: Among them, ω w ,ω h are the determining factors corresponding to the length and width of the prediction box, respectively, GT With W GT Represent the length and width of the real box respectively; H and W represent the length and width of the predicted box respectively; max is the maximum value function; λ is the attention degree for controlling shape loss; e is a natural constant, and t is an intermediate parameter; The calculation formula of the distance loss is: Among them, ρ x , y are the determining factors related to the horizontal distance and vertical distance of the center point of the real box and the predicted box, respectively. γ is the determining factor, γ = 2-Λ; and are the x and y coordinates of the center point of the real frame, B x With B y are the x and y coordinates of the center point of the prediction box respectively; C' h With C' w are the length and width of the minimum bounding rectangle of the real box and the predicted box respectively; Λ is the angle loss, and the specific formula is: Among them, C h is the vertical distance between the center of the real box and the predicted box, and σ is the distance between the center of the real box and the predicted box.

Citation Information

Cited By

  • Small target identification method based on improved YOLOv5

    CN115661607A

  • A small target recognition method based on improved YOLOv5

    CN115661607B