An Unmanned Aerial Vehicle Aerial Photography Small Target Detection Method Based on Improved YOLOv4
By improving the backbone feature extraction network of the YOLOv4 model and strengthening the feature extraction network, the problems of small target detection of drone aerial images are solved, and the detection accuracy is improved and the number of model parameters is reduced, which is suitable for object detection tasks of drone images.
Patent Information
- Application Number
- CN202210701583.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-06-21
AI Technical Summary
The existing YOLOv4-based drone aerial image object detection model has problems such as high detection difficulty, large model parameters and high complexity when processing small targets. It is difficult to meet hardware needs and performs poorly on the VisDrone dataset.
By improving the backbone feature extraction network MobileCSPDarknet-tiny, the fifth Resblock_body is removed and the Resblock part is replaced with a deep separable module, and the Mob_Resblock_body is constructed; in the enhanced feature extraction network ASPP+Bi-PANet, ASPP is added and BiFPN ideas are integrated to enhance feature extraction capabilities; a modified YOLOv4 model is built and trained on the VisDrone dataset.
The network parameters are reduced and detection accuracy is improved. The mAP is increased by 7.36%, and the parameter volume is only 24.49% of the original model. It has the characteristics of low false alarm rate and high small target recognition rate, and is suitable for object detection tasks of drone images.
Smart Images

Figure CN115063701B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning and target detection, and in particular to a method for detecting small targets in unmanned aerial vehicle aerial photography based on improved YOLOv4. Background Art
[0002] In recent years, drone technology has been developing continuously. Small-sized and flexible drones have been widely used in life and production. As the tasks undertaken by drones become increasingly diverse, target detection in drone aerial images is not only conducive to ensuring the flight safety of drones and the smooth execution of tasks, but also expands the understanding of application scenarios and enhances the autonomous flight capabilities of intelligent drones. With the rapid development of artificial intelligence, the use of deep learning methods to process images and detect targets has gradually become mainstream. Target detection based on deep learning methods is mainly divided into a two-stage method based on candidate regions and a one-stage method based on regression. The YOLO series of models, as representatives of the one-stage method, has developed rapidly. Among them, YOLOv4 has become one of the best detection algorithms due to its high accuracy and fast speed.
[0003] VisDrone was collected by the AISKYEYE team of the Machine Learning and Data Mining Laboratory of Tianjin University based on different drone platforms at different heights and locations. The images are taken from 14 different regions in China. The images include various scenes, various weather and lighting conditions. It is a very challenging dataset for algorithm design. In view of the special environment of drone aerial photography, target detection based on YOLOv4 has the following difficulties: 1) The target size is small, the density is high, the background is complex, the contradiction between target information and deep network feature extraction is prominent, and detection is difficult; 2) The model has a large number of parameters and high complexity, which makes it difficult to meet hardware requirements. The target detection model for conventional scenes does not perform well on drone aerial images. Therefore, how to further compress the model while improving detection accuracy is the key difficulty of this work. Summary of the invention
[0004] The purpose of the present invention is to provide a small target detection method for UAV aerial photography based on improved YOLOv4, aiming to improve the target detection accuracy of UAV aerial images and provide a guarantee for the UAV to successfully complete the task.
[0005] The technical solution to achieve the purpose of the present invention is:
[0006] A method for detecting small targets in drone aerial photography based on improved YOLOv4 includes the following steps:
[0007] Establish the backbone feature extraction network MobileCSPDarknet-tiny:
[0008] The backbone feature extraction network MobileCSPDarknet-tiny is based on CSPDarknet53. The 5th Resblock_body is removed, and the Resblock part in the remaining 4 Resblock_bodies is replaced with a depthwise separable module to construct Mob_Resblock_body. The outputs of the last layers of the 2nd, 3rd, and 4th Mob_Resblock_bodies are used as the outputs of the backbone feature extraction network, named out1, out2, and out3 in sequence.
[0009] Build the enhanced feature extraction network ASPP+Bi-PANet:
[0010] The enhanced feature extraction network ASPP+Bi-PANet first adds ASPP to the three outputs out1, out2, and out3 of the backbone feature extraction network MobileCSPDarknet-tiny respectively. Then, based on PANet, the idea of BiFPN is integrated, and the channel stacks after upsampling and downsampling are respectively replaced with weighted summation.
[0011] Build the improved YOLOv4 model:
[0012] The improved YOLOv4 model mainly includes an input end, the backbone feature extraction network MobileCSPDarknet-tiny, the enhanced feature extraction network ASPP+Bi-PANet, and the classification and regression layer YOLO Head.
[0013] Use the divided VisDrone dataset to train this model.
[0014] Compared with the prior art, the significant advantages of the present invention are:
[0015] (1) In the improvement of the backbone feature extraction network, by replacing the Resblock part in Resblock_body with a depthwise separable module to construct Mob_Resblock_body, the network parameters can be effectively reduced, realizing the lightweight of the network. By removing the 5th Resblock_body in CSPDarknet53 and using the outputs of the last layers of the 2nd, 3rd, and 4th Mob_Resblock_bodies as the outputs of the backbone feature extraction network, not only can the network parameters be further effectively reduced, but also it is beneficial to solve the problem of serious information loss of small targets in aerial images in the deep layer of the network, effectively improving the detection accuracy.
[0016] (2) In the improvement of the enhanced feature extraction network, by replacing APP with ASPP and adding ASPP to each output of the backbone feature extraction network, the information loss caused by pooling operations can be effectively reduced, the receptive field can be increased, and the detection accuracy can be improved. By integrating the BiFPN idea into PANet, while retaining the advantages of bidirectional fusion, more features can be fused through skip connections, especially the original outputs of the backbone feature extraction network (i.e., out1 / out2 / out3) are included in each part of the fusion, effectively reducing the information loss of aerial small targets during the repeated feature extraction operations in the enhanced feature extraction network, which is beneficial to the improvement of detection accuracy.
[0017] (3) By building an improved YOLOv4 model and training and testing it on the VisDrone dataset, compared with the original YOLOv4, the mAP of the present invention is increased by 7.36%, and the number of parameters is only 24.49% of the original model. It has the characteristics of low false alarm rate and high small target recognition rate, and is suitable for the target detection task of drone images. Brief Description of the Drawings
[0018] Figure 1 Schematic diagram of the principle of the present invention;
[0019] Figure 2 Comparison diagram of the VisDrone dataset with the PASCAL VOC dataset and the MS COCO dataset;
[0020] Figure 3 Statistical chart of the number of each target in the VisDrone dataset;
[0021] Figure 4 Statistical chart of the number of targets with absolute sizes of each category in the VisDrone dataset;
[0022] Figure 5 Statistical chart of the number of targets with relative sizes of each category in the VisDrone dataset;
[0023] Figure 6 Overall structure diagram of the improved YOLOv4;
[0024] Figure 7 Structure diagram of the improved Mob_Resblock_body module;
[0025] Figure 8 Structure diagram of the depthwise separable module
[0026] Figure 9 Structure diagram of the SELayer;
[0027] Figure 10 Structure diagram of the improved ASPP module;
[0028] Figure 11 It is the structural diagram of the three - convolution block;
[0029] Figure 12 It is the structural diagram of the five - convolution block;
[0030] Figure 13 It is the graph of the improvement of AP for each category of the improved YOLOv4 (%). Specific implementation manners
[0031] The present invention will be further introduced below in conjunction with the accompanying drawings and specific embodiments.
[0032] Combined with Figure 1 , a method for small - target detection of UAV aerial - photography images based on improved YOLOv4 of the present invention includes the following steps:
[0033] Step 1: Divide the data set, and count and analyze the VisDrone data set;
[0034] 1.1 Select the VisDrone data set for model training. The data set includes a total of 10,209 images, among which there are 6,471 training images, 548 validation images, and 3,190 test images. This method only uses 6,471 training images and 548 validation images with publicly disclosed labels, a total of 7,019 images, for model training and prediction, and re - divides the training set, validation set, and test set according to the ratio of 7:2:1;
[0035] 1.2 Count the average number of labeled boxes per image in the data set and compare it with the PASCAL VOC data set and the MS COCO data set, as Figure 2 shown;
[0036] 1.3 Count the number of each target and its proportion in the data set, as Figure 3 shown;
[0037] 1.4 Define the target with a target - box pixel area less than 32 * 32 as an absolute - size small target, the target with a target - box pixel area greater than 96 * 96 as an absolute - size large target, and the target between the two as an absolute - size medium target. Respectively count the proportion of the absolute size of each category in the data set, as Figure 4 shown;
[0038] 1.5 Define the target with a target - box area accounting for less than 0.01 (1%) of the whole image as a relative - size small target, the target with a target - box area accounting for more than 0.1 (10%) of the whole image as a relative - size large target, and the target between the two as a relative - size medium target. Respectively count the proportion of the relative size of each category in the data set, as Figure 5 shown;
[0039] 1.6 Analyze and summarize the statistical data: The average number of annotations per image in the VisDrone dataset is as high as 52.9, far higher than the commonly used VOC or COCO datasets, with a large target density; there is a serious imbalance in the number of targets in each category in the dataset; small targets with an absolute size account for 61% in the dataset, and small targets with a relative size account for as high as 97%. The large number of small targets makes the contradiction between the target size and the features extracted by the deep network prominent, and models designed based on normal sizes generally perform poorly on the VisDrone dataset.
[0040] Step 2: Establish the backbone feature extraction network MobileCSPDarknet-tiny;
[0041] Combined with Figure 6 , the backbone feature extraction network MobileCSPDarknet-tiny is based on CSPDarknet53, removes the 5th Resblock_body, and replaces the Resblock part in the remaining 4 Resblock_bodies with depthwise separable modules to construct Mob_Resblock_body. The outputs of the last layers of the 2nd, 3rd, and 4th Mob_Resblock_bodies are used as the outputs of the backbone feature extraction network, named out1, out2, and out3 in sequence;
[0042] 2.1 Combined with Figure 8 , the depthwise separable module includes two cases:
[0043] (1) When the number of input channels is equal to the number of output channels, the depthwise separable module sequentially includes a depthwise convolution with a kernel size of 3*3, a BN layer, a first activation function, an SELayer, a regular 2D convolution with a kernel size of 1*1, a BN layer, and a second activation function;
[0044] (1.1) The first activation function is h-swish = x * ReLU6(x + 3) / 6, where x is the input of the activation function, and ReLU6 = min(max(x, 0), 6);
[0045] (1.2) Combined with Figure 9 , the SELayer sequentially includes average pooling, a fully connected layer, a ReLU activation function, a fully connected layer, and a third activation function;
[0046] (1.2.1) The third activation function is h-sigmoid = ReLU6(x + 3) / 6, where x is the input of the activation function, and ReLU6 = min(max(x, 0), 6);
[0047] (1.3) The second activation function is Mish = x * tanh(ln(1 + e x ))), where x is the input of the activation function;
[0048] (2) When the number of input channels is not equal to the number of output channels, the depthwise separable module sequentially performs a common two-dimensional convolution with a kernel size of 1*1, a BN layer, a first activation function, a depthwise convolution with a kernel size of 3*3, a BN layer, a SELayer, a first activation function, a common two-dimensional convolution with a kernel size of 1*1, a BN layer, and a second activation function;
[0049] (2.1) The first activation function is h-swish = x * ReLU6(x + 3) / 6, where x is the input of the activation function and ReLU6 = min(max(x, 0), 6);
[0050] (2.2) Combined with Figure 9 , the SELayer sequentially performs average pooling, a fully connected layer, a ReLU activation function, a fully connected layer, and a third activation function;
[0051] (2.2.1) The third activation function is h-sigmoid = ReLU6(x + 3) / 6, where x is the input of the activation function and ReLU6 = min(max(x, 0), 6);
[0052] (2.3) The second activation function is Mish = x * tanh(ln(1 + e x ))), where x is the input of the activation function;
[0053] 2.2 Combined with Figure 7 , the number of Mob_Resblock parts in the 4 Mob_Resblock_body are 2, 8, 8, and 4 respectively;
[0054] Step 3: Establish a strengthened feature extraction network ASPP + Bi-PANet;
[0055] Combined with Figure 6 , the strengthened feature extraction network ASPP + Bi-PANet first adds ASPP to the three outputs out1, out2, and out3 of the backbone feature extraction network MobileCSPDarknet-tiny respectively. Then, on the basis of PANet, integrating the BiFPN idea, it replaces the channel stacking after upsampling and the channel stacking after downsampling with weighted addition respectively. The specific steps are as follows:
[0056] 3.1 Add ASPP to each of the three outputs out1, out2, and out3 of the backbone feature extraction network MobileCSPDarknet-tiny respectively;
[0057] (1) Combine Figure 10 , and the ASPP replaces the max pooling with pooling kernel sizes of 5, 9, and 13 in SPP with dilated convolutions with dilation rates of 2, 4, and 6 respectively;
[0058] (2) Combine Figure 6 , out1 passes through a normal 2D convolution with a kernel size of 1*1, ASPP, and a normal 2D convolution with a kernel size of 1*1 in sequence, and the output is named out1_1; out2 passes through a normal 2D convolution with a kernel size of 1*1, ASPP, and a normal 2D convolution with a kernel size of 1*1 in sequence, and the output is named out2_1; out3 passes through three convolutions, ASPP, three convolutions, and a normal 2D convolution with a kernel size of 1*1 in sequence, and the output is named out3_1;
[0059] Combine Figure 11 , and the three convolutions are sequential normal 2D convolutions with a kernel size of 1*1, a normal 2D convolution with a kernel size of 3*3, and a normal 2D convolution with a kernel size of 1*1 in sequence;
[0060] 3.2 Combine Figure 6 , based on PANet, fuse the BiFPN idea, and replace the channel stacking after upsampling and the channel stacking after downsampling with weighted summation respectively;
[0061] (1) out3_1 passes through a normal convolution with a kernel size of 1*1 and upsampling with the algorithm nearest, and the output is ou3_2. This process is defined as the first upsampling. The weights after the first upsampling include three parts, namely w 0 *out2_1, w 1 *out3_2, w 2 *out2. After the three parts are added together, they pass through five convolutions, and the output is out2_3;
[0062] (1.1) The w 0 , w 1 , w 2 are weight matrices, and their values are obtained through training;
[0063] (1.2) Combine Figure 12 , and the five convolutions are sequential normal 2D convolutions with a kernel size of 1*1, a normal 2D convolution with a kernel size of 3*3, a normal 2D convolution with a kernel size of 1*1, a normal 2D convolution with a kernel size of 3*3, and a normal 2D convolution with a kernel size of 1*1 in sequence;
[0064] (2) out2_3 undergoes a regular convolution with a kernel size of 1*1 and an upsampling with the algorithm of nearest, and outputs ou2_2. This process is defined as the second upsampling. The weights after the second upsampling include three parts, namely w 0 *out1_1, w 1 *out2_2, w 2 *out1. After adding these three parts, it undergoes five convolutions and outputs out1_2;
[0065] (2.1) The w 0 , w 1 , w 2 are weight matrices, and their values are obtained through training;
[0066] (2.2) Combining Figure 12 , the five convolutions are sequentially a regular two-dimensional convolution with a kernel size of 1*1, a regular two-dimensional convolution with a kernel size of 3*3, a regular two-dimensional convolution with a kernel size of 1*1, a regular two-dimensional convolution with a kernel size of 3*3, and a regular two-dimensional convolution with a kernel size of 1*1;
[0067] (3) out1_2 undergoes a downsampling with a kernel size of 3*3 and a stride of 2, and outputs out1_3. This process is defined as the first downsampling. The weights after the first downsampling include three parts, namely w 0 *out1_3, w 1 *out2_3, w 2 *out2. After adding these three parts, it undergoes five convolutions and outputs out2_5;
[0068] (3.1) The w 0 , w 1 , w 2 are weight matrices, and their values are obtained through training;
[0069] (3.2) Combining Figure 12 , the five convolutions are sequentially a regular two-dimensional convolution with a kernel size of 1*1, a regular two-dimensional convolution with a kernel size of 3*3, a regular two-dimensional convolution with a kernel size of 1*1, a regular two-dimensional convolution with a kernel size of 3*3, and a regular two-dimensional convolution with a kernel size of 1*1;
[0070] (4) out2_5 undergoes a downsampling with a kernel size of 3*3 and a stride of 2, and outputs out2_4. This process is defined as the second downsampling. The weights after the second downsampling include three parts, namely w 0 *out2_4, w 1 *out3_1, w 2*out3. After the three parts are added together, it goes through five convolutions, and the output is out3_3;
[0071] (4.1) The w 0 、w 1 、w 2 are weight matrices, and their values are obtained through training;
[0072] (4.2) Combining Figure 12 , the five convolutions are in sequence: a normal two-dimensional convolution with a kernel size of 1*1, a normal two-dimensional convolution with a kernel size of 3*3, a normal two-dimensional convolution with a kernel size of 1*1, a normal two-dimensional convolution with a kernel size of 3*3, and a normal two-dimensional convolution with a kernel size of 1*1;
[0073] Step 4: Build an improved YOLOv4 model and train the model using the VisDrone dataset;
[0074] Combining Figure 6 , the YOLOv4 model mainly includes an input end, a backbone feature extraction network CSPDarknet53, an enhanced feature extraction network SPP+PANet, and a classification and regression layer YOLO Head; the improved YOLOv4 model mainly includes an input end, a backbone feature extraction network MobileCSPDarknet-tiny, an enhanced feature extraction network ASPP+Bi-PANet, and a classification and regression layer YOLO Head;
[0075] 4.1 Build an improved YOLOv4 model: Pass the out1_2 obtained in Step 3 through YOLO HEAD1; pass out2_5 through YOLO HEAD2; pass out3_3 through YOLO HEAD3;
[0076] (1) The YOLO HEAD1 sequentially goes through a normal two-dimensional convolution with a kernel size of 3*3, a BN layer, a LeakyReLU activation function, and a normal two-dimensional convolution with a kernel size of 1*1;
[0077] (2) The YOLO HEAD2 sequentially goes through a normal two-dimensional convolution with a kernel size of 3*3, a BN layer, a LeakyReLU activation function, and a normal two-dimensional convolution with a kernel size of 1*1;
[0078] (3) The YOLO HEAD3 sequentially goes through a normal two-dimensional convolution with a kernel size of 3*3, a BN layer, a LeakyReLU activation function, and a normal two-dimensional convolution with a kernel size of 1*1;
[0079] 4.2 Use the dataset in Step 1 to train the improved YOLOv4 model, and evaluate it using mAP and Total params as model evaluation metrics;
[0080] Recall = TP / (TP + FN)
[0081] Precision = TP / (TP + FP)
[0082] Recall (recall rate) and Precision (accuracy rate) are recognized evaluation metrics in the field of object detection; where TP represents the number of correctly recognized positive samples, FN represents the number of misclassified or unrecognized positive samples, and FP represents the number of negative samples misrecognized as objects. For each category, a curve can be drawn based on Recall and Precision, and the area of this curve with the coordinate axes is the AP value. mAP is the average of the APs of all categories; Total params is the total number of model parameters.
[0083] The platform hardware conditions for this embodiment are: the processor is a 12-core Intel i9-9900K CPU @ 3.60GHz, the graphics processor is a GeForce RTX 2080Ti, and the host is a Windows system with 11G of RAM. The compilation environment is Python 3.7, and the PyTorch version is 1.5.1. The specific experimental steps are as follows:
[0084] (1) Dataset statistics and analysis: The VisDrone dataset was selected for this training, which contains a total of 7,019 aerial drone images. The object categories include: bus, van, truck, bicycle, tricycle, awning-tricycle, pedestrian, people, motor, car. There is a serious imbalance in the number of targets in each category of the dataset. The total number of the two categories of car and pedestrian, which account for a relatively large proportion in the dataset, occupies more than half of the total number of labeled instances. While the total number of the three categories of bus, tricycle, and awning-tricycle, which account for a relatively small proportion in the dataset, accounts for less than 7% in the dataset. It can be seen from the statistical results that the proportion of small targets with an absolute size less than 32*32 pixels in the VisDrone dataset is 61%, and the proportion of small targets with a relative size less than 0.01 times the image size is as high as 97%. The large number of small targets makes the contradiction between the target size and the features extracted by the deep network prominent. Due to the limitation of hardware devices, the image is generally resized to 416*416 or 608*608 when fed into the model, which further increases the difficulty of small target detection. And due to the existence of model downsampling, the object detection algorithm based on CNN will further reduce the object information volume, resulting in weak expression ability of deep features for small targets. The models designed based on normal sizes generally perform poorly on the VisDrone dataset. The average number of annotations per image in the dataset is as high as 54.4, which is much higher than the commonly used VOC or COCO datasets. The target density is large, and the occurrence of occlusion makes the features of the target seriously lacking and it is very difficult to recover through reasoning. In addition, the two categories of pedestrain (a person who maintains an upright posture or is walking) and people (people in other situations) are very easy to be confused, which brings difficulties to the correct detection of the target.
[0085] (2) Establish an improved YOLOv4 model, and set the training parameters of the improved YOLOv4: The freezing training step is 20, and the initial learning rate is 1e-3; the unfreezing training step is 4, and the initial learning rate is 1e-4. The model is trained for 350 rounds in total;
[0086] (3) Analysis of training results: The small target detection results of aerial drone images based on the improved YOLOv4 are evaluated by mAP and Total params. Figure 13Regarding the improvement of the AP of each category of the improved YOLOv4 compared to the unimproved YOLOv4, the Total params of the original model is 64363101, and the Total params of the improved YOLOv4 model is 15761675. Compared with the original YOLOv4, the performance of this method on the VisDrone dataset has an mAP increase of 7.36%, and the number of parameters is only 24.49% of the original model. It has the characteristics of low false alarm rate and high recognition rate of small targets, and is suitable for the target detection task of drone images.
[0087] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as a limitation of the present invention itself. Various changes in form and detail may be made without departing from the spirit and scope of the present invention defined by the appended claims.
Claims
1. An improved YOLOv4-based small target detection method for UAV aerial photography, characterized in that, it includes the following steps: Establish the backbone feature extraction network MobileCSPDarknet-tiny: The backbone feature extraction network MobileCSPDarknet-tiny is based on CSPDarknet53, removes the 5th Resblock_body, and replaces the Resblock part in the remaining 4 Resblock_bodies with depthwise separable modules to construct Mob_Resblock_body. The outputs of the last layers of the 2nd, 3rd, and 4th Mob_Resblock_bodies are used as the outputs of the backbone feature extraction network, named out1, out2, and out3 in sequence; Establish the enhanced feature extraction network ASPP+Bi-PANet: The enhanced feature extraction network ASPP+Bi-PANet first adds ASPP to the three outputs out1, out2, and out3 of the backbone feature extraction network MobileCSPDarknet-tiny respectively. Then, on the basis of PANet, it integrates the BiFPN idea and replaces the channel stacking after upsampling and the channel stacking after downsampling with weighted addition respectively; Establish an improved YOLOv4 model: The improved YOLOv4 model includes an input end, the backbone feature extraction network MobileCSPDarknet-tiny, the enhanced feature extraction network ASPP+Bi-PANet, and the classification and regression layer YOLO Head; Use the divided VisDrone dataset to train this model, and the dataset includes multiple images; When the number of input channels is equal to the number of output channels, the depthwise separable module sequentially performs depthwise convolution with a kernel size of 3*3, a BN layer, a first activation function, an SELayer, ordinary two-dimensional convolution with a kernel size of 1*1, a BN layer, and a second activation function; The first activation function is h-swish = x * ReLU6(x + 3) / 6, where ReLU6 = min(max(x, 0), 6); the SELayer sequentially includes average pooling, a fully connected layer, a ReLU activation function, a fully connected layer, and a third activation function; the third activation function is h-sigmoid = ReLU6(x + 3) / 6, where ReLU6 = min(max(x, 0), 6); the second activation function is Mish = x * tanh(ln(1 + e x )), where x is the input of the activation function; When the number of input channels is not equal to the number of output channels, the depthwise separable module sequentially performs ordinary two-dimensional convolution with a kernel size of 1*1, a BN layer, a first activation function, depthwise convolution with a kernel size of 3*3, a BN layer, an SELayer, a first activation function, ordinary two-dimensional convolution with a kernel size of 1*1, a BN layer, and a second activation function; The first activation function is h-swish = x * ReLU6(x + 3) / 6, where ReLU6 = min(max(x, 0), 6); the SELayer sequentially includes average pooling, a fully connected layer, a ReLU activation function, a fully connected layer, and a third activation function; the third activation function is h-sigmoid = ReLU6(x + 3) / 6, where ReLU6 = min(max(x, 0), 6); the second activation function is Mish = x * tanh(ln(1 + e x )), where x is the input of the activation function; Add ASPP to the three outputs out1, out2, and out3 of the backbone feature extraction network MobileCSPDarknet-tiny respectively. The specific process is as follows: The ASPP replaces the max pooling with pooling kernel sizes of 5, 9, and 13 in SPP with dilated convolutions with dilation rates of 2, 4, and 6 respectively; out1 sequentially passes through a common two-dimensional convolution with a kernel size of 1*1, ASPP, and a common two-dimensional convolution with a kernel size of 1*1, and the output is named out1_1; out2 sequentially passes through a common two-dimensional convolution with a kernel size of 1*1, ASPP, and a common two-dimensional convolution with a kernel size of 1*1, and the output is named out2_1; out3 sequentially passes through three convolutions, ASPP, three convolutions, and a common two-dimensional convolution with a kernel size of 1*1, and the output is named out3_1; The three convolutions sequentially include a common two-dimensional convolution with a kernel size of 1*1, a common two-dimensional convolution with a kernel size of 3*3, and a common two-dimensional convolution with a kernel size of 1*1; Replace the channel stack after upsampling and the channel stack after downsampling with weighted addition respectively. The specific process is as follows: out3_1 undergoes ordinary convolution with a convolution kernel size of 1*1 and upsampling with the algorithm nearest, and outputs ou3_2. This process is defined as the first upsampling. The weights after the first upsampling include three parts, namely w 0 *out2_1, w 1 *out3_2, w 2 *out2. After the three parts are added together, they undergo five convolutions and output out2_3; out2_3 undergoes ordinary convolution with a convolution kernel size of 1*1 and upsampling with the algorithm of nearest, and outputs ou2_2. This process is defined as the second upsampling. The weights after the second upsampling include three parts, namely w 0 *out1_1, w 1 *out2_2, w 2 *out1. After the three parts are added together and go through five convolutions, the output is out1_2; out1_2 is downsampled with a convolutional kernel of size 3*3 and a stride of 2, and out1_3 is output. This process is defined as the first downsampling. The weights after the first downsampling include three parts, namely w 0 *out1_3, w 1 *out2_3, w 2 *out2. After the three parts are added together, they go through five convolutions and out2_5 is output; out2_5 is downsampled with a convolution kernel of size 3*3 and a stride of 2, and the output is out2_4. This process is defined as the second downsampling. The weights after the second downsampling include three parts, namely w 0 *out2_4, w 1 *out3_1, w 2 *out3. After adding these three parts, it goes through five convolutions and outputs out3_3; where w 0 , w 1 , w 2 are weight matrices; the five - time convolution is sequentially a normal two - dimensional convolution with a convolution kernel size of 1*1, a normal two - dimensional convolution with a convolution kernel size of 3*3, a normal two - dimensional convolution with a convolution kernel size of 1*1, a normal two - dimensional convolution with a convolution kernel size of 3*3, and a normal two - dimensional convolution with a convolution kernel size of 1*1.
2. The method for detecting small targets in UAV aerial photography based on improved YOLOv4 according to claim 1, characterized in that, The number of Mob_Resblock parts in the 4 Mob_Resblock_body are 2, 8, 8, and 4 respectively.
3. The method for detecting small targets in UAV aerial photography based on improved YOLOv4 according to claim 1, characterized in that, Build an improved YOLOv4 model. The specific process is as follows: Build an improved YOLOv4 model: pass the obtained out1_2 through YOLO HEAD1; out2_5 through YOLO HEAD2; out3_3 through YOLO HEAD3; Use the dataset to train the improved YOLOv4 model, and evaluate it with mAP and Total params as the model evaluation indicators; Recall = TP / (TP + FN) Precision = TP / (TP + FP) Where TP represents the number of correctly identified positive samples, FN represents the number of positive samples that are misclassified or not identified, FP represents the number of negative samples misidentified as targets. For each category, a curve can be drawn according to Recall and Precision, and the area of the curve and the coordinate axes is the AP value, and mAP is the average value of the APs of all categories; Total params is the total number of model parameters.
4. The method for detecting small targets in UAV aerial photography based on improved YOLOv4 according to claim 3, characterized in that, The YOLO HEAD1 sequentially includes a common two-dimensional convolution with a kernel size of 3*3, a BN layer, a LeakyReLU activation function, and a common two-dimensional convolution with a kernel size of 1*1; the YOLO HEAD2 sequentially includes a common two-dimensional convolution with a kernel size of 3*3, a BN layer, a LeakyReLU activation function, and a common two-dimensional convolution with a kernel size of 1*1; the YOLO HEAD3 sequentially includes a common two-dimensional convolution with a kernel size of 3*3, a BN layer, a LeakyReLU activation function, and a common two-dimensional convolution with a kernel size of 1*1.
Citation Information
Patent Citations
Infrared image weak and small target detection method based on improved YOLO v3
CN112101434A
Railway intrusion foreign matter detection method, device and terminal
CN113205510A