An object detection method for drone-captured images based on an improved YOLOv5x model
By constructing the RSTS-YOLOv5 model, the image features extracted from the YOLOv5 structure are combined with the implicit context information extracted by the RSTS module, and multi-scale data enhancement is carried out, which solves the detection problem caused by too small objects and lack of obvious features in the images captured by the drone, and achieves the high-precision and low leakage detection rate object detection effect.
Patent Information
- Application Number
- CN202211042295.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-29
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-08-29
AI Technical Summary
The objects in the images captured by existing drones are too small and the characteristics are extremely insignificant, making it difficult to detect, the detection accuracy is relatively low and the missed detection rate is high.
Based on the YOLOv5x improved model, the RSTS-YOLOv5 model is constructed, and the image features in the partially extracted single-frame image retained in the YOLOv5 structure are combined with the implicit context information extracted by the RSTS module. Through multi-scale data enhancement, the clustering characteristics of images are used to improve detection accuracy.
The drone has improved the detection accuracy of objects in small target images, reduced the missed detection rate and error detection rate, and achieved simple operation, high detection efficiency and high accuracy.
Smart Images

Figure CN115482475B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target recognition. More specifically, it is a method for object detection in drone-captured images based on an improved YOLOv5x model. Background Art
[0002] In recent years, with the development of the drone industry and deep learning technology, object detection in drone-captured images has been widely applied in daily life and industrial production. Specifically, this technology has been widely used in scenarios such as smart cities, plant protection, and wildlife protection. However, since most objects in drone-captured images are small objects with low resolution and almost no visual information, it is difficult to extract discriminative features and they are easily affected by environmental factors. Moreover, due to reasons such as a higher clustering probability of objects in drone-captured images and multiple objects on the feature map aggregating into one point, the current detection schemes have relatively low detection accuracy and a high miss detection rate.
[0003] To solve the above problems, the present invention proposes a method for object detection in drone-captured images based on an improved YOLOv5x model, and uses the open-source dataset VisDrone dataset for training and testing. By constructing RSTS-YOLOv5, the image features extracted from a single frame of image in the part retained in the YOLOv5 structure are combined with the implicit context information extracted by the RSTS module, making up for the problem of difficult detection caused by the overly small size and extremely unclear features of objects in drone-captured images. Utilizing the clustering characteristics of drone-captured images, multi-scale data augmentation is performed, with the effects of simple operation, high detection efficiency, high accuracy, low miss detection rate, and low false detection rate. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for object detection in drone-captured images based on an improved YOLOv5x model, which combines the image features and color information in a single frame of image with the implicit context information between features, making up for the problem of difficult detection caused by the overly small size and extremely unclear features of objects in drone-captured images, and improving the detection accuracy of the deep learning model for object detection in small-target images captured by drones.
[0005] To achieve the above-mentioned invention purpose, the present invention provides a method for object detection in drone-captured images based on an improved YOLOv5x model, including the following steps:
[0006] Step 1: Obtain a target detection dataset for drone-captured images;
[0007] Step 2: Convert the format of the dataset in Step 1. Use a Python script to convert the detection labels into YOLO-format txt label files. Each line in the file represents a bounding box, and each line contains the following parameters: object class, the abscissa x of the center of the bounding box, the ordinate y of the center of the bounding box, the width width of the bounding box, and the height height of the bounding box;
[0008] Step 3: Perform multi-scale data augmentation on the dataset with the converted format obtained in Step 2. First, uniformly set the length of the dataset images to 1024, and then perform multi-scale data augmentation with corresponding scales of 1120, 1248, 1344, and 1472 respectively;
[0009] Step 4: Build the RSTS module (Res-Swin Transformer Stage), and the specific steps are as follows;
[0010] Step 4-1: Through Patch Merging, select 1 element every 2 pixels in both the row and column directions of the input feature map, then splice them together as 1 tensor, and finally unfold it; at this time, the channel dimension will become 4 times the original, because both H and W are reduced by 2 times. At this time, adjust the channel dimension to twice the original through a fully connected layer again;
[0011] Step 4-2: Downsample the original tensor to a size of 1 / 2 of the original in terms of length and width through convolution operations, and keep the number of channels unchanged;
[0012] Step 4-3: Normalize the tensor obtained in Step 4-1 along the channel dimension through LayerNorm;
[0013] Step 4-4: Divide the tensor obtained in Step 4-3 into multiple windows, that is, change the dimensions of the input tensor so that it changes from x∈R H×W×C spatial to x∈R N×Ws×Ws×C , where N = h×w / (Ws×Ws), H and W represent the height and width of the vector space, Ws represents the number of windows, and h, w, and C represent the original height, width, and number of channels respectively;
[0014] Step 4-5: Perform multi-head self-attention mechanism calculations on each window to obtain the feature map encoded with relative position information;
[0015] Step 4-6: Normalize the output of Step 4-4 through LayerNorm, then change the dimensions to the original input dimensions through a fully connected layer, and fuse the output tensor with the output of Step 4-4 as a residual to obtain a new feature map;
[0016] Steps 4-7: Window migration is achieved by shifting the feature maps output in Steps 4-6 and setting a mask for the self-attention mechanism calculation process. Then, Steps 4-3 to 4-7 are repeated, and finally, the original positions of the feature maps are restored;
[0017] Step 4-8: Fuse the residual obtained in Step 4-2 with the output obtained in Step 4-7;
[0018] Step 5: Build the RSTS-YOLOv5 model based on the RSTS module constructed in Step 4. The specific steps are as follows:
[0019] Step 5-1: Replace the C3 layer in the 10th layer of the backbone in the original YOLOv5 with the RSTS module. The specific parameters of this module are: 3 pairs of Swin Transformer blocks, window size of 4, and the number of heads of the multi-head attention mechanism is 12;
[0020] Step 5-2: After the 17th layer, continue to upsample the feature maps to make the feature maps continue to expand. At the same time, at the 20th layer, fuse the obtained feature maps with the feature maps of the 3rd layer of the YOLOv5 backbone network to obtain larger feature maps for small target detection;
[0021] Step 5-3: Add RSTS modules with 1, 3, 2, and 1 pairs of Swin Transformer blocks respectively at the 23rd, 27th, 31st, and 35th layers of the YOLOv5 backbone network. The window size is 4, and the number of heads of the multi-head attention mechanism is 12;
[0022] Step 6: For all the training sample sets in the dataset, uniformly perform mosaic data augmentation, and then train the model obtained in Step 5 until the training is completed;
[0023] Step 7: For all the test sample sets in the dataset, conduct batch testing, and use the model that passes the test for real-time object detection.
[0024] Furthermore, the specific steps of Step 3 are as follows:
[0025] Step 3-1: Calculate the average coordinate position (x mean , y mean ) of all the target boxes included in an image:
[0026]
[0027]
[0028] Among them, x i , y iThey respectively represent the coordinate positions of the center points of each target box, and n represents the total number of target boxes;
[0029] Step 3-2: Calculate the cropping scale corresponding to each scale size
[0030]
[0031]
[0032] Among them, l xo , l yo They respectively represent the length and width dimensions of the original image, and its l xo takes values of 1120, 1248, 1344, 1472, and k represents the magnification factor, that is, the magnified scale / the original scale;
[0033] Step 3-3: Taking the average coordinate position of all target boxes as the center point, with the cropping scale being the length of the extended pixel points from the center point, crop out the image area, obtain a new image, and magnify it to a unified value of 1024 for the length to get a new static image;
[0034] Step 3-4: Calculate the pixel coordinate positions and their lengths and widths of the target boxes after the image is cropped and magnified. For the targets with more than 50% of their area still located in the new image, retain their target boxes and update their label values according to the image content to generate new label data.
[0035] The present invention discloses a method for object detection in drone-captured images based on an improved YOLOv5x model. By constructing RSTS-YOLOv5, the image features extracted from a single frame of image in the part retained in the YOLOv5 structure are combined with the implicit context information extracted by the RSTS module, making up for the problem of difficult detection caused by too small objects and extremely unclear features in drone-captured images. Utilizing the clustering characteristics of drone-captured images, multi-scale data augmentation is carried out, having the effects of simple operation, high detection efficiency, high accuracy, low missed detection rate and low false detection rate. Brief Description of the Drawings
[0036] Figure 1 is the flow chart of the method for object detection in drone-captured images based on the improved YOLOv5x model of the present invention.
[0037] Figure 2 is the schematic diagram of multi-scale data augmentation designed by the present invention.
[0038] Figure 3 is the structural diagram of the RSTS module designed by the present invention.
[0039] Figure 4 is the network structure diagram of the RSTS-YOLOv5 constructed by the present invention.
[0040] Figure 5 It is an example of the recognition result diagram. Specific implementation manners
[0041] The following describes the specific implementation manners of the present invention with reference to the accompanying drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.
[0042] Embodiment 1:
[0043] Step 1: Obtain the drone-captured image target detection dataset VisDrone.
[0044] Step 2: Convert the format of the dataset in Step 1. Use a Python script to convert the detection labels into YOLO-format txt label files. Each line in the file represents a target box, and the parameters in one line are: object-class, x, y, width, height.
[0045] Step 3: Perform multi-scale data augmentation on the dataset with the converted format obtained in Step 2. Before the training process, first uniformly set the length of the dataset images to 1024, and perform multi-scale data augmentation with corresponding scale sizes of 1120, 1248, 1344, and 1472 respectively;
[0046] Step 3-1: Calculate the average coordinate position (x mean , y mean ) of all target boxes in an image:
[0047]
[0048]
[0049] Among them, x i , y i respectively represent the coordinate positions of the center points of each target box;
[0050] Step 3-2: Calculate the cropping scale corresponding to each scale size
[0051]
[0052]
[0053] Among them, l xo , l yorepresent the length and width dimensions of the original image respectively, and k represents the magnification factor, that is, the magnified scale / the original scale;
[0054] Step 3-3: Using the average coordinate position of all target boxes as the center point, intercept an image area with a scale that extends pixel points from the center point, obtain a new image, and magnify it to a unified length value of 1024 to get a new static image;
[0055] Step 3-4: Calculate the pixel coordinate positions and their lengths and widths of the target boxes after the image is cropped and magnified. For targets with more than 50% of their area still located in the new image, retain their target boxes and update their label values to generate new label data.
[0056] Step 4: Construct a Res-Swin Transformer Stage (RSTS) module. One RSTS contains a Patch Merging layer, multiple pairs of Swin Transformer blocks that appear in pairs, and a convolutional layer for dimension matching;
[0057] Step 4-1: Through Patch Merging, select 1 element every 2 pixels in both the row and column directions of the input feature map, then splice them together as 1 tensor, and finally expand. At this time, the channel dimension will become 4 times the original (because both H and W are reduced by 2 times). At this time, adjust the channel dimension to twice the original through a fully connected layer;
[0058] Step 4-2: Downsample the original tensor to a length and width that is 1 / 2 of the original tensor through convolution operations, and keep the number of channels unchanged;
[0059] Step 4-3: Normalize the tensor obtained in Step 4-1 along the channel dimension through LayerNorm;
[0060] Step 4-4: Divide the tensor obtained in Step 4-3 into multiple windows, that is, change the dimensions of the input tensor so that it changes from x ∈ R H×W×C spatial to x ∈ R N×Ws×Ws×C , where N = h × w / (Ws × Ws); Ws represents the number of windows, and h, w, C represent the original height, width, and number of channels respectively;
[0061] Step 4-5: Perform multi-head self-attention mechanism calculations on each window to obtain a feature map encoded with relative position information;
[0062] Step 4-6: Normalize the output of Step 4-4 through LayerNorm, then change the dimension to the original dimension of the input through a fully connected layer, and fuse the output tensor with the output of Step 4-4 as the residual to obtain a new feature map;
[0063] Step 4-7: Indirectly implement window migration by shifting the feature map output in Step 4-6 and setting a mask for the attention module, then repeat Steps 4-3 to 4-7, and finally restore the original position of the feature map;
[0064] Step 4-8: Fuse the residual obtained in Step 4-2 with the output obtained in Step 4-7.
[0065] Step 5: Build the RSTS-YOLOv5 model based on the RSTS module constructed in Step 4. The specific steps are as follows:
[0066] Step 5-1: Replace the C3 layer in the 10th layer of the backbone in the original YOLOv5 with the RSTS module. The specific parameters of this module are: 3 pairs of Swin Transformer blocks, the window size is 4, and the number of heads of the multi-head attention mechanism is 12.
[0067] Step 5-2: After the 17th layer, continue to perform upsampling and other processing on the feature map to make the feature map continue to expand. At the same time, at the 20th layer, fuse the obtained feature map with the feature map of the 3rd layer in the backbone network to obtain a larger feature map for small target detection.
[0068] Step 5-3: Add RSTS modules with 1, 3, 2, and 1 pairs of Swin Transformer blocks respectively at the 23rd, 27th, 31st, and 35th layers. The window size is 4, and the number of heads of the multi-head attention mechanism is 12.
[0069] Step 6: For all training sample sets in the dataset, uniformly perform mosaic data augmentation, set the number of iterations to 150, batch-size to 1, warm-up learning rate to 0.1, initial learning rate to 0.01. After 150 iterations of training, the loss value and accuracy tend to be stable, and save the best parameter model at this time.
[0070] Step 7: For all test sample sets in the dataset, perform batch testing. The input image resolution size is 1536, the confidence threshold of the screening box is 0.001, the IOU threshold of NMS is 0.65, and calculate its mean average precision mAP for evaluation.
[0071] Embodiment 2:
[0072] Different from Embodiment 1, for the object detection scenario of UAV captured images with clustering characteristics, the multi-scale data augmentation proposed in this patent has good effects; however, for the object detection scenario where objects are randomly distributed more evenly, this step should be skipped. The specific steps are as follows:
[0073] Step 1: Obtain the UAV captured image target detection dataset VisDrone.
[0074] Step 2: Convert the format of the dataset in Step 1. Use a Python script to convert the detection labels into YOLO format txt label files. Each line in the file represents a target box, and the parameters included in one line are: object-class, x, y, width, height.
[0075] Step 3: Construct a Res-Swin Transformer Stage (RSTS) module. One RSTS contains a Patch Merging layer, multiple pairs of Swin Transformer blocks that appear in pairs, and a convolutional layer for dimension matching;
[0076] Step 3-1: Through Patch Merging, select 1 element every 2 pixels in both the row and column directions of the input feature map, then splice them together as 1 tensor, and finally expand. At this time, the channel dimension will become 4 times the original (because both H and W are reduced by 2 times). At this time, adjust the channel dimension to twice the original through a fully connected layer;
[0077] Step 3-2: Downsample the original tensor to a length and width that are 1 / 2 of the original tensor through convolution operations, and keep the number of channels unchanged;
[0078] Step 3-3: Normalize the tensor obtained in Step 3-1 along the channel dimension through LayerNorm;
[0079] Step 3-4: Divide the tensor obtained in Step 3-3 into multiple windows, that is, change the dimensions of the input tensor so that it changes from x ∈ R H×W×C spatial to x ∈ R N×Ws×Ws×C , where N = h × w / (Ws × Ws); Ws represents the number of windows, and h, w, C represent the original height, width, and number of channels respectively;
[0080] Step 3-5: Perform multi-head self-attention mechanism calculations on each window to obtain the feature map encoded with relative position information;
[0081] Step 3-6: Normalize the output of Step 3-4 through LayerNorm, then change the dimension to the original dimension of the input through a fully connected layer, and fuse the output tensor with the output of Step 3-4 as the residual to obtain a new feature map;
[0082] Step 3-7: Indirectly implement window migration by shifting the feature map and setting a mask for the attention module, then repeat Steps 3-3 to 3-7, and finally restore the original position of the feature map;
[0083] Step 3-8: Fuse the residual obtained in Step 3-2 with the output obtained in Step 3-7.
[0084] Step 4: Build the RSTS-YOLOv5 model based on the RSTS module constructed in Step 4. The specific steps are as follows:
[0085] Step 4-1: Replace the C3 layer in the 10th layer of the backbone in the original YOLOv5 with the RSTS module. The specific parameters of this module are: 3 pairs of Swin Transformer blocks, the window size is 4, and the number of heads of the multi-head attention mechanism is 12.
[0086] Step 4-2: After the 17th layer, continue to perform upsampling and other processing on the feature map to make the feature map continue to expand. At the same time, at the 20th layer, fuse the obtained feature map with the feature map in the 3rd layer of the backbone network to obtain a larger feature map for small target detection.
[0087] Step 4-3: Add RSTS modules with 1, 3, 2, and 1 pairs of Swin Transformer blocks at the 23rd, 27th, 31st, and 35th layers respectively. The window size is 4, and the number of heads of the multi-head attention mechanism is 12.
[0088] Step 5: For all training sample sets in the dataset, uniformly perform mosaic data augmentation. Set the number of iterations to 150, batch-size to 1, warm-up learning rate to 0.1, and initial learning rate to 0.01. After 150 iterations of training, the loss value and accuracy tend to be stable, and save the best parameter model at this time.
[0089] Step 6: For all test sample sets in the dataset, perform batch testing. The input image resolution size is 1536, the confidence threshold of the screening box is 0.001, and the IOU threshold of NMS is 0.65. Calculate its mean average precision mAP for evaluation.
Claims
1. A method for object detection in drone - captured images based on an improved YOLOv5x model, comprising the following steps: Step 1: Obtain the object detection dataset of the images captured by the drone; Step 2: Convert the format of the dataset in Step 1. Use a Python script to convert the detection labels into YOLO-format txt label files. Each line in the file represents a bounding box, and the parameters in one line include: object class, the abscissa x of the center of the bounding box, the ordinate y of the center of the bounding box, the width width of the bounding box, and the height height of the bounding box; Step 3: Perform multi-scale data augmentation on the dataset with the converted format obtained in Step 2. First, uniformly set the length of the dataset images to 1024, and then perform multi-scale data augmentation with corresponding scales of 1120, 1248, 1344, and 1472 respectively; Step 4: Construct the RSTS module, and the specific steps are as follows; Step 4-1: Through Patch Merging, select 1 element every 2 pixels in both the row and column directions of the input feature map, then splice them together as 1 tensor, and finally unfold it; at this time, the channel dimension will become 4 times the original, because both H and W are reduced by 2 times. At this time, adjust the channel dimension to twice the original through a fully connected layer; Step 4-2: Downsample the original tensor to a size of 1 / 2 of the original in terms of length and width through convolution operations, while keeping the number of channels unchanged; Step 4-3: Normalize the tensor obtained in Step 4-1 along the channel dimension through LayerNorm; Step 4-4: Divide the tensor obtained in Step 4-3 into multiple windows, that is, change the dimensions of the input tensor so that it changes from x ∈ R H×W×C space to x ∈ R N×Ws×Ws×C , where N = h × w / (Ws × Ws), H and W represent the height and width of the vector space, Ws represents the number of windows, and h, w, and C represent the original height, width, and number of channels respectively; Step 4-5: Perform multi-head self-attention mechanism calculation on each window to obtain the feature map encoded with relative position information; Step 4-6: Normalize the output of Step 4-4 through LayerNorm, then change the dimension to the original input dimension through a fully connected layer, and fuse the output tensor with the output of Step 4-4 as the residual to obtain a new feature map; Step 4-7: Realize window migration by shifting the feature map output in Step 4-6 and setting a mask for the self-attention mechanism calculation process, then repeat Steps 4-3 to 4-7, and finally restore the original position of the feature map; Step 4-8: Fuse the residual obtained in Step 4-2 with the output obtained in Step 4-7; Step 5: Build the RSTS-YOLOv5 model based on the RSTS module constructed in Step 4. The specific steps are as follows: Step 5-1: Replace the C3 layer of the 10th layer in the backbone of the original YOLOv5 with the RSTS module. The specific parameters of this module are: 3 pairs of Swin Transformer blocks, the window size is 4, and the number of heads of the multi-head attention mechanism is 12; Step 5-2: After the 17th layer, continue to upsample the feature map to make the feature map continue to expand. At the same time, at the 20th layer, fuse the obtained feature map with the feature map of the 3rd layer of the YOLOv5 backbone network to obtain a larger feature map for small object detection; Step 5-3: Add the RSTS module of the SwinTransformer block composed of pairs of 1, 3, 2, and 1 respectively to the 23rd, 27th, 31st, and 35th layers of the YOLOv5 backbone network. Its window size is 4, and the number of heads of the multi-head attention mechanism is 12; Step 6: For all training sample sets in the dataset, perform mosaic data augmentation uniformly, and then train the model obtained in Step 5 until the training is completed; Step 7: For all test sample sets in the dataset, conduct batch testing, and use the model that passes the test for real-time object detection.
2. The method for object detection in drone - captured images based on an improved YOLOv5x model according to claim 1, wherein, The specific steps of Step 3 are as follows: Step 3-1: Calculate the average coordinate position (x mean , y mean ) of all target bounding boxes in an image: where x i , y i represent the coordinate positions of the center points of each target box respectively, and n represents the total number of target boxes; Step 3-2: Calculate the cropping scale corresponding to each scale size Among them, l xo , l yo respectively represent the length and width dimensions of the original image, and its l xo takes values of 1120, 1248, 1344, 1472, and k represents the magnification factor, that is, the magnified scale / the original scale; Step 3-3: Take the average coordinate position of all target boxes as the center point, intercept the scale as the length of the extended pixel points from the center point, crop out the image area, obtain a new image, and enlarge it to a unified value of 1024 for the length to get a new static image; Step 3-4: Calculate the pixel coordinate positions and their lengths and widths of the target boxes corresponding to the image after cropping and enlargement. For targets with more than 50% of their area still located in the new image, retain their target boxes and update their label values according to the image content to generate new label data.