A Drone Aerial Photography Small Target Detection Method Based on Improved YOLOv7 Algorithm

By improving the neck structure and feature fusion method of the YOLOv7 network, the accuracy of the detection of small targets by aerial photography by drone is enhanced, and the problem of low detection accuracy of the YOLOv7 algorithm is solved, and the detection accuracy is achieved.

CN116597326BActive Publication Date: 2025-07-22XIDIAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310525931.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-11
Publication Date
2025-07-22
Estimated Expiration
2043-05-11

AI Technical Summary

Technical Problem

The existing YOLOv7 algorithm has the problem of low detection accuracy in drone aerial photography small target detection, especially in target detection of specific sizes.

Method used

By improving the neck network structure of the YOLOv7 network, introducing spatial attention SGE and bidirectional cascade structure, adding a detection head network of 160×160 feature maps, performing feature fusion, and improving the spatial information and positioning capabilities of small targets.

Benefits of technology

The accuracy of detection of small targets by drone aerial photography was significantly improved, especially the detection capability of smaller targets, and mAP increased from 41.9% to 44.2%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116597326B_ABST
    Figure CN116597326B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting small targets in UAV aerial photography based on an improved YOLOv7 algorithm, which includes the following steps; Step 1, obtain the UAV aerial photography dataset and convert it into the YOLO format, and divide the training set, validation set, and test set; Step 2, build an improved YOLOv7 network; the improved YOLOv7 network is an improvement that modifies the neck network structure of the YOLOv7 network and introduces spatial attention SGE therein; Step 3, use the improved YOLOv7 network as the detection model, and use the training set and validation set to train and validate the monitoring model to obtain the final detection model; Step 4, use the final detection model, take the UAV aerial photography image as the input, and perform small target detection. The present invention is used to improve the accuracy of small target detection in UAV aerial photography.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision target detection, and particularly relates to a method for detecting small targets in UAV aerial photography based on an improved YOLOv7 algorithm. Background Art

[0002] The detection of small targets in UAV aerial photography is to use target detection technology to assist in the identification and positioning of small targets in UAV aerial photography videos. This technology can be applied to the automated inspection of UAVs, such as forest inspection, border inspection, etc. At present, the target detection algorithms are mainly based on deep learning target detection algorithms. Such algorithms use convolutional neural networks to extract deep features of images and have strong generalization ability. Therefore, they are widely used in various target detection tasks.

[0003] The deep learning-based algorithms for detecting small targets in UAV aerial photography are divided into two-stage and one-stage algorithms. The two-stage algorithms are mainly the R-CNN algorithm proposed by R. Girshick et al. and its improved algorithms. Such algorithms need to first extract region candidate boxes in the training stage and then use CNN to extract features. In the inference stage, they also need to first extract region candidate boxes and then use the features generated in the training stage to judge whether the candidate boxes belong to the target or the background through a classifier. When the candidate box belongs to the target, the candidate box is fine-tuned to complete the detection of the target. Its advantage is high detection accuracy, and the disadvantage is that it cannot achieve real-time effects.

[0004] The one-stage algorithms are mainly the YOLO algorithm proposed by Joseph Redmon et al. and its improved algorithms. Such algorithms directly identify and locate targets through the CNN network. The main idea is to divide the image into an S×S grid. If the center point of an object is located in a certain grid, the grid is responsible for predicting the category and position of the object. Its advantage is that it can detect in real time, but the detection accuracy is slightly inferior to that of the R-CNN series. However, with the continuous improvement of the YOLO series algorithms, their detection accuracy has reached or even exceeded that of the two-stage algorithms.

[0005] The YOLO series of cutting-edge algorithms have good real-time performance and are suitable for the real-time detection of small targets in UAV aerial photography. However, applying the cutting-edge version YOLOv7 of the YOLO series to the detection of small targets in UAV aerial photography has the problem of low detection accuracy.

[0006] The existing CN202210701583.5, a method for detecting small targets in UAV aerial photography based on an improved YOLOv4, lightweighted the YOLOv4 backbone network and improved the neck network feature fusion structure, improving the detection accuracy. However, the sizes of the three scale feature maps of 13×13, 52×52, and 104×104 output are very different, and it is easy to miss detecting targets of a specific size. Summary of the Invention

[0007] To overcome the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a method for detecting small targets in UAV aerial photography based on an improved YOLOv7 algorithm, which is used to improve the detection accuracy of small targets in UAV aerial photography.

[0008] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0009] A method for detecting small targets in UAV aerial photography based on an improved YOLOv7 algorithm, comprising the following steps;

[0010] Step 1, obtain the UAV aerial photography dataset and convert it into the YOLO format, and divide the training set, validation set, and test set;

[0011] Step 2, build an improved YOLOv7 network; the improved YOLOv7 network is an improvement that modifies the neck network structure of the YOLOv7 network and introduces spatial attention SGE therein;

[0012] Step 3, use the improved YOLOv7 network as the detection model, and use the training set and validation set to train and validate the monitoring model to obtain the final detection model;

[0013] Step 4, use the final detection model to input the UAV aerial photography image for small target detection.

[0014] The UAV aerial photography dataset in Step 1 is the VisDrone dataset. Download the VisDrone2022 UAV dataset on the official website, and use the training set images and their labels to train the network.

[0015] In Step 2, build an improved YOLOv7 network, and build a backbone network, a neck network, and a four-detection-head network for 160×160 feature maps respectively;

[0016] The backbone network is used to extract features from the dataset images;

[0017] The neck network is used to fuse and enhance the image features extracted by the backbone network;

[0018] The four-detection-head network for 160×160 feature maps is used to predict the targets in the image using the features enhanced by the neck network. The backbone network is the same as the YOLOv7 backbone network, and its structure from left to right is: input layer -> convolutional layer 1 -> convolutional layer 2 -> convolutional layer 3 -> convolutional layer 4 -> layer aggregation network structure 1 -> downsampling layer 1 -> layer aggregation network structure 2 -> downsampling layer 2 -> layer aggregation network structure 3 -> downsampling layer 3 -> layer aggregation network structure 4; where:

[0019] The image size of the input layer is 640×640. All convolutional layers are of the CBS convolutional structure, which is the basic building block of the entire network and is composed of Conv convolution, batch normalization BN, and SiLU activation function in series. Without special instructions, all convolutional operations are convolutional operations with padding, and the size of the feature map before and after convolution will not change;

[0020] The convolutional kernel size of convolutional layer 1 is 3×3, the sliding step is 1, and the number of channels is 32;

[0021] The convolutional kernel size of convolutional layer 2 is 3×3, the sliding step is 2, and the number of channels is 64. This convolutional layer is a convolutional operation without padding and serves as a downsampling layer, making the feature map become 320×320;

[0022] The convolutional kernel size of convolutional layer 3 is 3×3, the sliding step is 1, and the number of channels is 64;

[0023] The convolutional kernel size of convolutional layer 4 is 3×3, the sliding step is 2, and the number of channels is 128. This convolutional layer is a convolutional operation without padding and serves as a downsampling layer, making the feature map become 160×160;

[0024] The layer aggregation network structure 1 is composed of 7 CBS convolutional structures. The specific connection method is as follows: First, 5 CBS convolutional structures are connected in series. From left to right, there is 1 CBS convolutional structure with a convolutional kernel size of 1×1, a sliding step of 1, and 64 channels, and 4 CBS convolutional structures with a convolutional kernel size of 3×3, a sliding step of 1, and 64 channels. Then, the 1st, 3rd, and 5th CBS convolutional structures are concatenated with another CBS convolutional structure with a convolutional kernel size of 1×1, a sliding step of 1, and 64 channels. Finally, a CBS convolutional structure with a convolutional kernel size of 1×1, a sliding step of 1, and 256 channels is connected in series after the concatenation result;

[0025] The downsampling layer 1 is composed of 3 CBS convolutional structures and 1 max-pooling structure. The specific connection method is as follows: It is divided into upper and lower parts. The upper part is a max-pooling structure connected in series with 1 CBS convolutional structure with a convolutional kernel size of 1×1, a sliding step of 1, and 128 channels. The lower part is 1 CBS convolutional structure with a convolutional kernel size of 1×1, a sliding step of 1, and 128 channels connected in series with 1 non-padded CBS convolutional structure with a convolutional kernel size of 3×3, a sliding step of 2, and 128 channels. Finally, the upper and lower parts are concatenated, making the feature map become 80×80;

[0026] The layer aggregation network structure 2 has the same structure as the layer aggregation network structure 1. The difference is that the number of channels of the first 6 CBS convolutional structures is 128, and the number of channels of the last 1 CBS convolutional structure is 512;

[0027] The structure of the downsampling layer 2 is the same as that of the downsampling layer 1, except that the number of channels of the 3 CBS convolutional structures is 256, making the feature map become 40×40;

[0028] The structure of the layer aggregation network structure 3 is the same as that of the layer aggregation network structure 1, except that the number of channels of the first 6 CBS convolutional structures is 256, and the number of channels of the last 1 CBS convolutional structure is 1024;

[0029] The structure of the downsampling layer 3 is the same as that of the downsampling layer 1, except that the number of channels of the 3 CBS convolutional structures is 512, making the feature map become 20×20;

[0030] The layer aggregation network structure 4 is exactly the same as the layer aggregation network structure 3.

[0031] The neck network is an E-PAN structure improved from the YOLOv7 neck network PANet. The structure of this neck network is: Spatial Pyramid Pooling structure -> Feature Pyramid 1 -> Feature Pyramid 2; where:

[0032] Feature Pyramid 1 is mainly composed of three upsampling layers and three layer aggregation network structures. The three upsamplings make the size of the feature map change from the 20×20 feature map after passing through the Spatial Pyramid Pooling structure to 40×40, 80×80, and 160×160, which are used to perform feature fusion with the 40×40, 80×80, and 160×160 feature maps in the backbone network respectively. 160×160 is the shallow feature map of the input in the backbone network after two downsamplings, and 40×40 is the deep feature map of the input in the backbone network after four downsamplings. Here, the upsampling layer uses transposed convolution. The difference between the layer aggregation network structure here and the layer aggregation network structure in the backbone network is that the 5 concatenated CBS convolutional structures are all concatenated with another 1 CBS convolutional structure. Feature Pyramid 1 undergoes three feature fusions, and each feature fusion structure is a BiC bidirectional cascaded structure. Taking the first feature fusion as an example, this BiC structure is composed of three parts concatenated together. The first part is the output of the layer aggregation network structure 1 concatenated with a CBS convolutional structure with a kernel size of 1×1 and a stride of 1. The second part is the output of convolutional layer 3 concatenated with a CBS convolutional structure with a kernel size of 1×1 and a stride of 1 and a non-padding CBS convolutional structure with a kernel size of 3×3 and a stride of 2. The third part is the output of the Spatial Pyramid Pooling structure passing through a CBS convolutional structure with a kernel size of 1×1 and a stride of 1 concatenated with a transposed convolution upsampling;

[0033] Feature Pyramid 2 is mainly composed of three downsampling layers and three layer aggregation network structures, and also undergoes three feature fusions. Each feature fusion structure is the same as the feature fusion structure in the YOLOv7 network.

[0034] For the E-PAN structure, first, input the original 20×20 feature map, 40×40 feature map, and 80×80 feature map of YOLOv7 into the neck network, and also input the 160×160 feature map of the backbone network into the neck network to construct a two-way fusion feature pyramid with three upsamplings and three downsamplings.

[0035] Secondly, replace the fusion structure of the shallow feature map and the deep feature map of the top-down feature pyramid FPN of the neck network with a two-way cascaded structure. This two-way cascaded structure replaces the upsampling method with a transposed convolution and adds another shallow feature map on the basis of the original fusion structure of the shallow feature map and the deep feature map for feature fusion.

[0036] The four-detection-head network of the 60×160 feature map mainly includes four reparameterized convolutions. The convolution kernels of the four convolutions are all 3×3, the sliding step is 1, and the number of channels is 128, 256, 512, and 1024 respectively.

[0037] In step 3 of training the network model, input the configured network environment into the VisDrone training set in step 1, train until convergence on the training set of the VisDrone dataset, and obtain the mAP of the YOLOv7 and the improved YOLOv7 models on the test set of the VisDrone dataset. Among them, the full name of mAP is Mean Average Precision, which is the result of summing the average precisions AP of all target classes and dividing by the number of classes. mAP uses the mean of the average detection precisions of all target classes to measure the algorithm performance.

[0038] The configured network environment is used to create a software environment for running the algorithm in the server.

[0039] The specific environment configuration is to install cuda10.1 and cudnn7603 on the server side for GPU-accelerated training, the artificial intelligence framework Pytorch1.7.1 for code support, and the libraries required for the operation of YOLOv7. Matplotlib is used for data visualization drawing, numpy is used for array and matrix operations, and opencv is used for image processing.

[0040] In step 2, prepare the UAV aerial photography dataset. The quality of the dataset directly affects the generalization performance of the algorithm. Therefore, it is necessary to prepare a relatively representative UAV aerial photography dataset for model training.

[0041] The UAV aerial photography dataset uses the VisDrone2022 dataset, which is divided into three parts: training set, validation set, and test set, with 6,471, 548, and 1,610 images respectively. To use this dataset for network training, it is necessary to convert the dataset format to the YOLO format required by the YOLOv7 network.

[0042] Train the original YOLOv7 algorithm and the improved YOLOv7 algorithm on the prepared dataset, and record the mean average precision (mAP) of the two models on the dataset. This performance metric is used to measure the performance of the model, and the performance improvement of the improved YOLOv7 algorithm is reflected by comparing the mAP of the two.

[0043] Advantages of the present invention:

[0044] First, since a 160×160 feature map is added to the input of the neck network, the shallow feature map with more spatial information is added to the bidirectional feature fusion of the neck network, and the spatial information of small targets is further enhanced.

[0045] Second, since the neck network uses a bidirectional cascade structure, two shallow feature maps are used for feature fusion with a deep feature map, and the localization ability of small targets is further enhanced.

[0046] Third, since the neck network uses four feature maps for feature fusion, the output feature map of the detection head network also changes from three to four. The newly added 160×160 output feature map can detect smaller targets, thereby improving the detection accuracy of small targets. Brief Description of the Drawings

[0047] Figure 1 It is the overall flowchart of the present invention.

[0048] Figure 2 It is the network structure diagram of YOLOv7.

[0049] Figure 3 It is the network structure diagram of the improved YOLOv7 proposed by the present invention.

[0050] Figure 4 It is the structure diagram of the original neck PANet and the new neck E-PAN.

[0051] Figure 5 It is the mAP of YOLOv7 and the improved YOLOv7 on the test set after training. Detailed Implementation Manner

[0052] The present invention will be further described in detail below with reference to the drawings and embodiments.

[0053] As Figures 1 - 5 shown, the implementation steps of this embodiment are as follows.

[0054] Step 1, Environment configuration.

[0055] Install cuda10.1 and cudnn7603 on the server side for GPU-accelerated training, the artificial intelligence framework Pytorch1.7.1 for code support, and some other libraries required for the operation of YOLOv7, such as matplotlib for data visualization drawing, numpy for array and matrix operations, opencv for image processing, etc.

[0056] Step 2, Prepare the VisDrone dataset.

[0057] Download the VisDrone2022 drone dataset from the official website, convert the dataset format to the YOLO format, and use 6471 images and their labels in the training set to train the network.

[0058] Step 3, Build an improved YOLOv7 network.

[0059] As Figure 3 shown, the improved YOLOv7 network consists of a backbone network, a neck network, and a detection head network, and its implementation is as follows:

[0060] 3.1) Build the backbone network of the improved YOLOv7 network:

[0061] This backbone network is the same as the YOLOv7 backbone network, and its structure from left to right is: input layer -> convolutional layer 1 -> convolutional layer 2 -> convolutional layer 3 -> convolutional layer 4 -> layer aggregation network structure 1 -> downsampling layer 1 -> layer aggregation network structure 2 -> downsampling layer 2 -> layer aggregation network structure 3 -> downsampling layer 3 -> layer aggregation network structure 4.

[0062] Among them:

[0063] All convolutional layers are CBS convolutional structures, which are the basic building blocks of the entire network, consisting of Conv convolution, batch normalization BN, and SiLU activation function in series. Without special explanation, all convolutional operations are convolutional operations with padding, and the size of the feature map before and after convolution will not change;

[0064] The convolutional kernel size of convolutional layer 1 is 3×3, the sliding stride is 1, and the number of channels is 32;

[0065] The convolutional kernel size of convolutional layer 2 is 3×3, the sliding stride is 2, and the number of channels is 64. This convolutional layer is a convolutional operation without padding and serves as a downsampling function;

[0066] The convolutional kernel size of convolutional layer 3 is 3×3, the sliding stride is 1, and the number of channels is 64;

[0067] The convolutional kernel size of convolutional layer 4 is 3×3, the sliding stride is 2, the number of channels is 128, and this convolutional layer is a non-padded convolution, which plays a role in downsampling;

[0068] The layer aggregation network structure 1 consists of 7 CBS convolutional structures. The specific connection method is as follows: First, 5 CBS convolutional structures are connected in series. From left to right, there is 1 CBS convolutional structure with a convolutional kernel size of 1×1, a sliding stride of 1, and a channel number of 64, and 4 CBS convolutional structures with a convolutional kernel size of 3×3, a sliding stride of 1, and a channel number of 64. Then, the 1st, 3rd, and 5th CBS convolutional structures are concatenated with another CBS convolutional structure with a convolutional kernel size of 1×1, a sliding stride of 1, and a channel number of 64. Finally, a CBS convolutional structure with a convolutional kernel size of 1×1, a sliding stride of 1, and a channel number of 256 is connected in series after the concatenation result;

[0069] The downsampling layer 1 consists of 3 CBS convolutional structures and 1 max-pooling structure. The specific connection method is as follows: It is divided into upper and lower parts. The upper part is a max-pooling structure connected in series with 1 CBS convolutional structure with a convolutional kernel size of 1×1, a sliding stride of 1, and a channel number of 128. The lower part is 1 CBS convolutional structure with a convolutional kernel size of 1×1, a sliding stride of 1, and a channel number of 128 connected in series with 1 non-padded CBS convolutional structure with a convolutional kernel size of 3×3, a sliding stride of 2, and a channel number of 128. Finally, the upper and lower parts are concatenated;

[0070] The layer aggregation network structure 2 has the same structure as the layer aggregation network structure 1. The difference is that the number of channels of the first 6 CBS convolutional structures is 128, and the number of channels of the last 1 CBS convolutional structure is 512;

[0071] The downsampling layer 2 has the same structure as the downsampling layer 1. The difference is that the number of channels of the 3 CBS convolutional structures is 256;

[0072] The layer aggregation network structure 3 has the same structure as the layer aggregation network structure 1. The difference is that the number of channels of the first 6 CBS convolutional structures is 256, and the number of channels of the last 1 CBS convolutional structure is 1024;

[0073] The downsampling layer 3 has the same structure as the downsampling layer 1. The difference is that the number of channels of the 3 CBS convolutional structures is 512;

[0074] The layer aggregation network structure 4 is exactly the same as the layer aggregation network structure 3.

[0075] 3.2) Build the neck network of the improved YOLOv7 network:

[0076] This neck network is an E-PAN structure improved from the PANet of the YOLOv7 neck network,Figure 4 These are two structural schematic diagrams. The neck network structure is: Spatial Pyramid Pooling structure -> Feature Pyramid 1 -> Feature Pyramid 2. Among them:

[0077] Feature Pyramid 1 is mainly composed of three upsampling layers and three layer aggregation network structures. Here, the upsampling layer uses transposed convolution. The difference between the layer aggregation network structure here and the backbone network layer aggregation network structure is that the 5 concatenated CBS convolution structures are all concatenated with another 1 CBS convolution structure. Feature Pyramid 1 undergoes three feature fusions, and each feature fusion structure is like Figure 4 (b) The BiC bidirectional cascaded structure shown on the right. Taking the first feature fusion as an example, this BiC structure is composed of three parts concatenated. The first part is the output of layer aggregation network structure 1 concatenated with a CBS convolution structure with a convolution kernel size of 1×1 and a stride of 1; the second part is the output of convolution layer 3 concatenated with a CBS convolution structure with a convolution kernel size of 1×1 and a stride of 1 and a non-padding CBS convolution structure with a convolution kernel size of 3×3 and a stride of 2; the third part is the output of the Spatial Pyramid Pooling structure passing through a CBS convolution structure with a convolution kernel size of 1×1 and a stride of 1 concatenated with a transposed convolution upsampling.

[0078] Feature Pyramid 2 is mainly composed of three downsampling layers and three layer aggregation network structures. Similarly, it undergoes three feature fusions, and each feature fusion structure is the same as the feature fusion structure of the YOLOv7 network.

[0079] 3.3) Build the detection head network of the improved YOLOv7 network:

[0080] This detection head network mainly includes four reparameterized convolutions. The convolution kernels of the four convolutions are all 3×3 and the stride is 1, and the number of channels is 128, 256, 512, and 1024 respectively.

[0081] Step 4, train the network model.

[0082] Train on the training set of the VisDrone dataset until convergence, and obtain the mAP of the YOLOv7 and the improved YOLOv7 models on the test set of the VisDrone dataset. As Figure 5 shown, by comparison, the mAP increases from 41.9% to 44.2%, indicating that the improved YOLOv7 greatly improves the accuracy of small target detection in UAV aerial photography.

[0083] In summary, a small target detection algorithm for UAV aerial photography based on improved YOLOv7 proposed by the present invention can better suit the small target detection task of UAV aerial photography and effectively improve the small target detection accuracy. The above description is only a specific example of the present invention, and does not limit the scope of the present invention. It is only for the convenience of those skilled in the art to understand that various deformations and improvements made to the technical solution of the present invention should fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for detecting small targets in UAV aerial photography based on an improved YOLOv7 algorithm, characterized in that, It includes the following steps; Step 1: Obtain the UAV aerial photography dataset and convert it into the YOLO format, and divide it into a training set, a validation set, and a test set; Step 2: Build an improved YOLOv7 network; the improved YOLOv7 network is an improvement that modifies the neck network structure of the YOLOv7 network and introduces the spatial attention SGE therein; Step 3: Use the improved YOLOv7 network as the detection model, and use the training set and the validation set to train and validate the detection model to obtain the final detection model; Step 4: Use the final detection model to perform small target detection with the UAV aerial photography image as the input; The neck network is the E-PAN structure improved from the YOLOv7 neck network PANet, and the neck network structure is: spatial pyramid pooling structure -> feature pyramid 1 -> feature pyramid 2; where: Feature pyramid 1 is mainly composed of three upsampling layers and three layer aggregation network structures. The three upsamplings make the feature map size change from the 20×20 feature map after passing through the spatial pyramid pooling structure to 40×40, 80×80, and 160×160, which are used to perform feature fusion with the 40×40, 80×80, and 160×160 feature maps in the backbone network respectively. 160×160 is the shallow feature map of the input in the backbone network after two downsamplings, and 40×40 is the deep feature map of the input in the backbone network after four downsamplings. Here, the upsampling layer uses transposed convolution. The difference between the layer aggregation network structure here and the layer aggregation network structure in the backbone network is that the 5 concatenated CBS convolution structures are all concatenated with another 1 CBS convolution structure. Feature pyramid 1 undergoes three feature fusions, and each feature fusion structure is a BiC bidirectional cascade structure. Taking the first feature fusion as an example, this BiC structure is composed of three parts spliced together. The first part is the output of the layer aggregation network structure 1 concatenated with a CBS convolution structure with a convolution kernel size of 1×1 and a stride of 1; the second part is the output of convolution layer 3 concatenated with a CBS convolution structure with a convolution kernel size of 1×1 and a stride of 1 and a CBS convolution structure without padding with a convolution kernel size of 3×3 and a stride of 2; the third part is the output of the spatial pyramid pooling structure passing through a CBS convolution structure with a convolution kernel size of 1×1 and a stride of 1 concatenated with a transposed convolution upsampling; Feature pyramid 2 is mainly composed of three downsampling layers and three layer aggregation network structures, and also undergoes three feature fusions. Each feature fusion structure is the same as the feature fusion structure in the YOLOv7 network.

2. The method for detecting small targets in UAV aerial photography based on the improved YOLOv7 algorithm according to claim 1, characterized in that, In step 2, build an improved YOLOv7 network, and build the backbone network, the neck network, and the four detection head networks for the 160×160 feature map respectively; The backbone network is used to extract features from the dataset images; The neck network is used to fuse and enhance the image features extracted by the backbone network; The four detection head networks for the 160×160 feature map are used to predict the targets in the image using the features enhanced by the neck network.

3. The method for detecting small targets in UAV aerial photography based on the improved YOLOv7 algorithm according to claim 2, characterized in that, The backbone network is the same as the YOLOv7 backbone network, and its structure from left to right is: input layer -> convolutional layer 1 -> convolutional layer 2 -> convolutional layer 3 -> convolutional layer 4 -> layer aggregation network structure 1 -> downsampling layer 1 -> layer aggregation network structure 2 -> downsampling layer 2 -> layer aggregation network structure 3 -> downsampling layer 3 -> layer aggregation network structure 4; where: The image size of the input layer is 640×640. All convolutional layers are CBS convolutional structures, which are the basic building blocks of the entire network and are composed of Conv convolution, batch normalization BN, and SiLU activation function in series. Without special instructions, all convolutional operations are convolutional operations with padding, and the size of the feature map before and after convolution will not change; The convolutional kernel size of convolutional layer 1 is 3×3, the sliding stride is 1, and the number of channels is 32; The convolutional kernel size of convolutional layer 2 is 3×3, the sliding stride is 2, and the number of channels is 64. This convolutional layer is a non-padded convolution that serves as a downsampling function, making the feature map become 320×320; The convolutional kernel size of convolutional layer 3 is 3×3, the sliding stride is 1, and the number of channels is 64; The convolutional kernel size of convolutional layer 4 is 3×3, the sliding stride is 2, and the number of channels is 128. This convolutional layer is a non-padded convolution that serves as a downsampling function, making the feature map become 160×160; Layer aggregation network structure 1 is composed of 7 CBS convolutional structures. The specific connection method is: First, 5 CBS convolutional structures are connected in series. From left to right, there is 1 CBS convolutional structure with a convolutional kernel size of 1×1, a sliding stride of 1, and 64 channels, and 4 CBS convolutional structures with a convolutional kernel size of 3×3, a sliding stride of 1, and 64 channels. Then, the 1st, 3rd, and 5th CBS convolutional structures are concatenated with another CBS convolutional structure with a convolutional kernel size of 1×1, a sliding stride of 1, and 64 channels. Finally, a CBS convolutional structure with a convolutional kernel size of 1×1, a sliding stride of 1, and 256 channels is connected in series after the concatenation result; Downsampling layer 1 is composed of 3 CBS convolutional structures and 1 max-pooling structure. The specific connection method is: It is divided into upper and lower parts. The upper part is a max-pooling structure connected in series with 1 CBS convolutional structure with a convolutional kernel size of 1×1, a sliding stride of 1, and 128 channels. The lower part is 1 CBS convolutional structure with a convolutional kernel size of 1×1, a sliding stride of 1, and 128 channels connected in series with 1 non-padded CBS convolutional structure with a convolutional kernel size of 3×3, a sliding stride of 2, and 128 channels. Finally, the upper and lower parts are concatenated, making the feature map become 80×80; Layer aggregation network structure 2 has the same structure as layer aggregation network structure 1, except that the number of channels of the first 6 CBS convolutional structures is 128, and the number of channels of the last 1 CBS convolutional structure is 512; Downsampling layer 2 has the same structure as downsampling layer 1, except that the number of channels of the 3 CBS convolutional structures is 256, making the feature map become 40×40; The structure of layer aggregation network structure 3 is the same as that of layer aggregation network structure 1, except that the number of channels of the first 6 CBS convolution structures is 256, and the number of channels of the last 1 CBS convolution structure is 1024; The structure of downsampling layer 3 is the same as that of downsampling layer 1, except that the number of channels of 3 CBS convolution structures is 512, making the feature map become 20×20; Layer aggregation network structure 4 is exactly the same as layer aggregation network structure 3.

4. A method for detecting small targets in UAV aerial photography based on an improved YOLOv7 algorithm according to claim 1, characterized in that, For the E-PAN structure, first, the original 20×20 feature map, 40×40 feature map, and 80×80 feature map of YOLOv7 are input into the neck network, and the 160×160 feature map of the backbone network is also input into the neck network to construct a two-way fusion feature pyramid with three upsamplings and three downsamplings; Secondly, the fusion structure of the shallow feature map and the deep feature map of the top-down feature pyramid FPN of the neck network is replaced with a two-way cascaded structure. This two-way cascaded structure replaces the upsampling method with a transposed convolution and adds another shallow feature map for feature fusion on the basis of the original fusion structure of the shallow feature map and the deep feature map.

5. The method for detecting small targets in UAV aerial photography based on the improved YOLOv7 algorithm according to claim 4, characterized in that, The four-detection-head network of the 60×160 feature map mainly includes four reparameterized convolutions. The convolution kernels of the four convolutions are all 3×3, the sliding step is 1, and the number of channels is 128, 256, 512, and 1024 respectively.

6. A method for detecting small targets in UAV aerial photography based on an improved YOLOv7 algorithm according to claim 1, characterized in that, The UAV aerial photography dataset in step 1 is the VisDrone dataset, and the training set images and their labels are used to train the network.

7. A method for detecting small targets in UAV aerial photography based on the improved YOLOv7 algorithm according to claim 6, characterized in that, In step 3, the network model is trained. The network environment configured in step 1 is input into the VisDrone training set in step 1, and training is carried out on the training set of the VisDrone dataset until convergence to obtain the mAP of YOLOv7 and the improved YOLOv7 model on the test set of the VisDrone dataset; among them, the full name of mAP is the mean average precision, which is the result of summing the average precision AP of all target classes and dividing by the number of classes. mAP uses the mean of the average detection precision of all target classes to measure the performance of the algorithm; The original YOLOv7 algorithm and the improved YOLOv7 algorithm are trained on the prepared dataset, and the mean average precision mAP of the two models on the dataset is recorded. This performance metric is used to measure the performance of the models, and the performance improvement of the improved YOLOv7 algorithm is reflected by comparing the mAP of the two.

8. A method for detecting small targets in UAV aerial photography based on an improved YOLOv7 algorithm according to claim 7, characterized in that, The configured network environment is used to create a software environment for running the algorithm in the server; The specific environment configuration is to install cuda10.1 and cudnn7603 on the server side for GPU-accelerated training, the artificial intelligence framework Pytorch1.7.1 for code support, and the libraries required for the operation of YOLOv7, matplotlib for data graphical drawing, numpy for array and matrix operations, and opencv for image processing.

Citation Information

Patent Citations

  • Unmanned aerial vehicle aerial photography small target detection method based on improved YOLOv4

    CN115063701A