Unmanned aerial vehicle detection method based on YOLOv5 improvement

Through the improved drone detection method based on YOLOv5, the model is pre-trained and fine-tuned by using data augmentation and full convolution mask autoencoder, the problems of slow infrared image detection speed and poor accuracy of drone under complex backgrounds are solved, and more efficient and accurate detection effects are achieved.

CN120147907APending Publication Date: 2025-06-13JIANGSU NORTH LAKE OPTOELECTRONICS CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510240764.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In complex backgrounds, the infrared image detection speed of drones is slow, the detection accuracy is poor, and is greatly affected by building facilities and background noise.

Method used

Based on YOLOv5's improved drone detection method, by acquiring drone infrared image data sets in complex backgrounds, data augmentation and annotation are carried out, a drone detection model based on YOLOv5 is constructed, and a full convolution mask autoencoder is used for pre-training and fine-tuning, optimizing the Backbone, Neck and Head parts of the model.

Benefits of technology

The speed and accuracy of infrared image detection of drones are improved, and the accuracy, detection speed and calculation amount of the model are optimized, which is suitable for drone detection tasks in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147907A_ABST
    Figure CN120147907A_ABST
Patent Text Reader

Abstract

The invention relates to an unmanned aerial vehicle detection method, system and device based on YOLOv5 improvement and a storage medium, and relates to the field of unmanned aerial vehicle detection. The method is based on an ImageNet data set, and comprises the following steps: obtaining an unmanned aerial vehicle infrared image data set shot under a background containing a shielding object; preprocessing and marking the unmanned aerial vehicle infrared image data set to obtain an unmanned aerial vehicle infrared image enhancement data set, wherein the unmanned aerial vehicle infrared image enhancement data set comprises a training set; an unmanned aerial vehicle detection model based on YOLOv5 is constructed; training the unmanned aerial vehicle detection model by using a full convolution mask auto-encoder to obtain a pre-training model; according to the pre-training model and the training set, performing layer-by-layer greedy pre-training and fine tuning on the unmanned aerial vehicle detection model to obtain an unmanned aerial vehicle detection improved model; and detecting the unmanned aerial vehicle infrared image data set by using the unmanned aerial vehicle detection improved model to obtain a detection result. The technical effect of the invention is to improve the speed and accuracy of infrared image detection of the unmanned aerial vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of drone detection, and in particular to a drone detection method, system, device, and storage medium improved based on YOLOv5. Background Art

[0002] YOLOv5 is the most stable version of the YOLO (You Only Look Once) series in engineering applications.

[0003] In recent years, with the development of computer image processing and drone technology, various object detections have been widely studied and successfully applied in fields such as video surveillance, agricultural information, power line detection, archaeological research, and reconnaissance. The continuous development of drone detection technology is playing an increasingly important role in airspace security, public safety, and the protection of important facilities.

[0004] However, when drones perform infrared image detection in complex backgrounds such as forests, houses, utility poles, and thick cloud cover, affected by obstacles such as building facilities and background noise, the detection speed of drone target infrared images is slow and the detection accuracy is poor. Summary of the Invention

[0005] In order to improve the speed and accuracy of drone infrared image detection, this application provides a drone detection method, system, device, and storage medium improved based on YOLOv5.

[0006] In a first aspect, this application provides a drone detection method improved based on YOLOv5, adopting the following technical solutions: Obtain a drone infrared image dataset taken under complex backgrounds and occlusions; Preprocess and label the drone infrared image dataset to obtain a drone infrared image enhanced dataset, and the drone infrared image enhanced dataset includes a training set; Construct a drone detection model based on YOLOv5; Use a fully convolutional mask autoencoder to train the ImageNet dataset to obtain a pre-trained model; According to the pre-trained model and the training set, layer-by-layer greedy pre-train and fine-tune the drone detection model to obtain a drone detection improved model; Use the drone detection improved model to detect the drone infrared image dataset to obtain a detection result.

[0007] Through the above technical solutions, for the detection of drones in complex backgrounds, a deep learning model based on YOLOv5 was trained. In the training stage, a fully convolutional masked autoencoder was used for pre-training, and the model was gradually fine-tuned through transfer learning. Finally, a model with the best accuracy, detection speed, and computational complexity was obtained.

[0008] In a specific feasible implementation, the preprocessing and annotation of the drone infrared image dataset to obtain the enhanced drone infrared image dataset include: Performing data augmentation processing on the drone infrared image dataset and annotating it to obtain the enhanced drone infrared image dataset; Randomly dividing the enhanced drone infrared image dataset into a training set according to a certain proportion.

[0009] Through the above technical solutions, data augmentation processing such as rotation, cropping, and adding noise is performed on the obtained drone dataset to enhance the filled data, making the model trained according to the image dataset more accurate.

[0010] In a specific feasible implementation, the construction of the drone detection model based on YOLOv5 includes: Constructing an initial drone detection model based on the improved YOLOv5, where the initial drone detection model includes a Backbone part, a Neck part, and a Head part; Replacing the Backbone part with ConvNeXtV2 with efficient channel attention and dynamic convolution; Replacing the feature extraction structure of the Neck part with the C2f-faster structure; Using a lightweight decoupled detection head in the Head part to obtain the drone detection model.

[0011] Through the above technical solutions, by adding the ECA attention mechanism to ConvNextV2, the computational complexity is reduced while the accuracy is improved for complex backgrounds. At the same time, the C2f-faster is used to optimize feature extraction, and the decoupled detection head is used to further optimize the speed and accuracy of the model.

[0012] In a specific feasible implementation, the ConvNeXtV2 includes ConvNeXt Block, and the ConvNeXtV2 with efficient channel attention and dynamic convolution includes layer normalization, multi-stage ConvNeXt Block sampling, feature weight adjustment under the ECA attention mechanism, and regularization.

[0013] Through the above technical solution, the Backbone part is replaced with ConvNeXtV2 with the ECA attention mechanism added. By introducing the channel attention mechanism, the model can better focus on important features and suppress unimportant features, thereby improving the model detection accuracy.

[0014] In a specific feasible implementation, the UAV infrared image enhancement dataset is also randomly divided into a validation set according to a ratio. The pre-trained model includes pre-trained weights, and the UAV detection model includes YOLOv5 pre-trained weights. The method of obtaining the improved UAV detection model by greedily pre-training and fine-tuning the UAV detection model layer by layer according to the pre-trained model and the training set includes: Adding the pre-trained weights and the YOLOv5 pre-trained weights to the UAV detection model to obtain the first UAV detection model. Using the training set and the validation set to unfreeze and train the Backbone part of the first UAV detection model layer by layer to obtain the second UAV detection model. Adjusting the hyperparameters of the second UAV detection model using a Bayesian optimizer to obtain the improved UAV detection model.

[0015] Through the above technical solution, layer-by-layer greedy pre-training is an effective deep neural network training strategy. By training layer by layer, it reduces the training complexity, alleviates the gradient problem, and learns the hierarchical features of the data, improving the training efficiency of the UAV detection model while ensuring the training quality.

[0016] In a specific feasible implementation, the UAV infrared image enhancement dataset is also randomly divided into a test set according to a ratio. After obtaining the improved UAV detection model by adjusting the UAV detection model according to the pre-trained model and the training set, it further includes: Inputting the test set into the improved UAV detection model to obtain test results. Evaluating the improved UAV detection model according to the test results.

[0017] In a specific feasible implementation, the method of evaluating the improved UAV detection model according to the test results includes: Obtaining the true results of the test set. Calculating evaluation metrics according to the test results and the true results. Judging whether the evaluation metrics meet the requirements. If the evaluation metrics meet the requirements, it is determined that the improved UAV detection model can be put into use. Otherwise, it is determined that the improved UAV detection model cannot be put into use.

[0018] Through the above technical solution, the obtained improved UAV detection model is verified using the test set to ensure the accuracy of the improved UAV detection model.

[0019] In a second aspect, the present application provides a UAV detection system improved based on YOLOv5, adopting the following technical solution: The system includes: A dataset acquisition module, configured to acquire a UAV infrared image dataset captured under complex backgrounds and occlusions; A dataset preprocessing module, configured to preprocess and annotate the UAV infrared image dataset to obtain an enhanced UAV infrared image dataset, where the enhanced UAV infrared image dataset includes a training set; A detection model construction module, configured to construct a UAV detection model based on YOLOv5; A pre-trained model construction module, configured to train the ImageNet dataset using a fully convolutional mask autoencoder to obtain a pre-trained model; A detection model training module, configured to greedily pre-train and fine-tune the UAV detection model layer by layer according to the pre-trained model and the training set to obtain an improved UAV detection model; An actual detection module, configured to use the improved UAV detection model to detect the UAV infrared image dataset to obtain a detection result.

[0020] In a third aspect, the present application provides a computer device, adopting the following technical solution: It includes a memory and a processor, and a computer program capable of being loaded and executed by the processor, such as the above-mentioned UAV detection method improved based on YOLOv5, is stored on the memory.

[0021] In a fourth aspect, the present application provides a computer-readable storage medium, adopting the following technical solution: It stores a computer program capable of being loaded and executed by the processor, such as the above-mentioned UAV detection method improved based on YOLOv5.

[0022] In summary, the present application has the following beneficial technical effects: The present invention trains a deep learning model based on YOLOv5 for UAV detection in complex backgrounds. This model uses the ConvNeXtV2 structure with an ECA attention mechanism added in the Block in the main part, uses the C2f-faster structure to replace the C3 structure in the feature fusion part, and uses a lightweight decoupled detection head. In the training stage, a fully convolutional mask autoencoder is used for pre-training, and the model is gradually fine-tuned through transfer learning. Finally, a model with the best effects in terms of accuracy, detection speed, and computational complexity is obtained. Description of the Drawings

[0023] Figure 1 It is the flowchart of the improved UAV detection method based on YOLOv5 in the embodiments of this application. Figure 2 It is the structure diagram of C2f-faster in the embodiments of this application.

[0024] Figure 3 It is the structure diagram of the decoupled head in the embodiments of this application.

[0025] Figure 4 It is the structure diagram of ConvNeXtv2 in the embodiments of this application.

[0026] Figure 5 It is the structure diagram of ConvNeXt Block in the embodiments of this application.

[0027] Figure 6 It is the comparison chart of evaluation indexes of different model structures in the embodiments of this application.

[0028] Figure 7 It is the structure block diagram of the improved UAV detection method based on YOLOv5 in the embodiments of this application.

[0029] Reference numerals: 701, data set acquisition module; 702, data set preprocessing module; 703, detection model construction module; 704, pre-trained model construction module; 705, detection model training module; 706, actual detection module. Detailed implementation manners

[0030] The following further describes this application in detail Figures 1 - 7 with reference to the accompanying drawings.

[0031] The embodiments of this application disclose an improved UAV detection method based on YOLOv5, which is used to improve the speed and accuracy of UAV infrared image detection.

[0032] YOLOv5 is the most stable version of the YOLO (You Only Look Once) series in engineering applications.

[0033] In recent years, with the development of computer image processing and UAV technology, various object detections have been widely studied and successfully applied in fields such as video surveillance, agricultural information, power line detection, archaeological research, and reconnaissance. The continuous development of UAV detection technology is playing an increasingly important role in airspace security, public security, and the protection of important facilities.

[0034] However, when the UAV performs infrared image detection in complex backgrounds such as forests, houses, utility poles, and thick cloud cover, affected by obstacles such as building facilities and background noise, the detection speed of the UAV target infrared image is slow and the detection accuracy is poor.

[0035] Therefore, this application proposes a drone detection method improved based on YOLOv5 to improve the speed and accuracy of drone infrared image detection using this method.

[0036] As Figure 1 shown, the method includes: S10. Obtain a drone infrared image dataset taken against a background containing occlusions.

[0037] Specifically, existing research has rarely detected small target drone object infrared images affected by building facilities and background noise in complex backgrounds, and the detection speed is slow. Therefore, to improve the speed and accuracy of drone infrared image detection, it is necessary to obtain drone infrared images in complex backgrounds, that is, to capture a multi-angle drone infrared image dataset under thick cloud cover with woods, houses, utility poles, high-voltage power towers, and the drone infrared image dataset also includes background noise.

[0038] S20. Preprocess and annotate the drone infrared image dataset to obtain a drone infrared image enhanced dataset, and the drone infrared image enhanced dataset includes a training set.

[0039] Specifically, perform data augmentation processing such as rotation, cropping, and adding noise on the obtained drone dataset and annotate it, and randomly divide it into a training set, a validation set, and a test set according to the ratio of 6:2:2.

[0040] S30. Build a drone detection model based on YOLOv5.

[0041] Specifically, build a drone detection model for complex backgrounds improved based on yolov5. YOLOv5 is the most stable version of the YOLO (You Only Look Once) series in engineering applications. The YOLOv5 structure includes the Backbone, Neck, and Head parts, and use the YOLOv5 network architecture to build a drone detection model for detecting infrared images.

[0042] S40. Use a fully convolutional masked autoencoder to train the ImageNet dataset to obtain a pre-trained model.

[0043] Specifically, perform self-supervised pre-training using a fully convolutional masked autoencoder, train in a fully convolutional manner, randomly mask the input signal, downsample the feature map at different stages, generate the mask at the last stage, and recursively upsample until the best resolution. The model predicts the missing part in the case of the remaining context to obtain the pre-trained model.

[0044] During the process of training the pre-trained model, the MSE loss between the reconstruction target and the ground truth is used as the loss function, and the loss function is only applied to the masked patches:

[0045] where y_true is the ground truth and y_pred is the predicted value.

[0046] The fully convolutional masked autoencoder adopts an asymmetric Encoder-Decoder architecture. In the masking stage, a ConvNet is used as the encoder, and a ConvNeXt Block is used as the decoder. In the encoding stage, the masked input data can be regarded as a two-dimensional sparse pixel matrix. Therefore, in the encoding stage, submanifold output definition sparse convolution is used for training, and operations are only performed on the matrix points of visible pixels. The convolution output is calculated only when the center of the convolution kernel covers a valid input index.

[0047] S50, layer-by-layer greedy pre-training and fine-tuning the drone detection model based on the pre-trained model and the training set to obtain an improved drone detection model.

[0048] Specifically, the pre-trained weights of the pre-trained model are added to the drone detection model, the Backbone part is frozen and thawed layer by layer. The dataset is randomly divided proportionally each time, and at the same time, the hyperparameters of the drone detection model are fine-tuned to obtain an improved drone detection model.

[0049] S60, using the improved drone detection model to detect the drone infrared image dataset to obtain the detection results.

[0050] Specifically, the improved drone detection model is used for drone target detection, that is, the drone infrared image dataset taken by the drone is detected to obtain the detection results.

[0051] For drone detection in complex backgrounds, a deep learning model based on YOLOv5 is trained. To meet the requirements of drone infrared image detection, the Backbone part, Neck part, and Head part of the model are modified. In the training stage, a fully convolutional masked autoencoder is used for pre-training, and the model is gradually fine-tuned through transfer learning. Finally, a model with improved accuracy, detection speed, and computational cost is obtained.

[0052] In one embodiment, to improve the accuracy of the drone infrared image detection model, the step of preprocessing and annotating the drone infrared image dataset to obtain the enhanced drone infrared image dataset includes: First, perform data augmentation on the UAV infrared image dataset and annotate it to obtain the enhanced UAV infrared image dataset. Specifically, the data augmentation process includes the following means: Mirror, flip, rotate, and scale the UAV infrared image dataset; Add random noise (such as Gaussian noise, random cropping, color transformation, etc.) to the UAV infrared image dataset; Fill the UAV infrared image dataset through linear or non-linear interpolation methods; Use a generative adversarial network to generate data similar to but slightly different from the original data; Data resampling, increasing the samples of minority classes or reducing the samples of majority classes.

[0053] Then, randomly divide the enhanced UAV infrared image dataset into a training set according to a certain proportion. Specifically, this method usually randomly divides the enhanced UAV infrared image dataset into a training set, a validation set, and a test set in a ratio of 6:2:2. The training set and the validation set are used for model training, and the test set is used for functional testing of the trained model.

[0054] The method of this application performs data augmentation processing such as rotation, cropping, and adding noise on the obtained UAV dataset, enhances and fills the data, so that the model trained according to the image dataset is more accurate.

[0055] In one embodiment, in order to improve the speed and accuracy of UAV infrared image detection, the step of constructing a UAV detection model based on YOLOv5 includes: First, construct an initial UAV detection model based on the improved YOLOv5. The initial UAV detection model includes a Backbone part, a Neck part, and a Head part. Specifically, the YOLOv5 network architecture includes an input layer, a Backbone layer, a Neck layer, and a Head layer. First, based on the initial UAV detection model of the improved YOLOv5, where the initial UAV detection model includes a Backbone part, a Neck part, and a Head part.

[0056] Then, replace the Backbone part with ConvNeXtV2 with efficient channel attention and dynamic convolution; replace the feature extraction structure of the Neck part with the C2f-faster structure with EMA; use a lightweight decoupled detection head in the Head part to obtain the UAV detection model. Specifically, the Neck part uses the C2f-faster structure to replace the original C3 structure of YOLOv5. The C2f-faster structure is as Figure 2As shown in the figure; the detection head part uses a decoupled detection head to predict the category and coordinates respectively to accelerate the convergence speed and improve the detection accuracy. The decoupled head includes a convolutional layer with a kernel size of 1×1 to reduce the channel dimension to 256, and then passes through two parallel branches. Each branch uses two convolutional layers with a size of 3×3, and the outputs are the number of categories at each position, 1 IOU at each position, and 4 regression values at each position. The decoupled detection structure is as Figure 3 shown.

[0057] The method of the present invention aims at the UAV targets in complex scenarios and at long distances. By adding the ECA attention mechanism to ConvNextV2, while reducing the computational amount, the accuracy for complex backgrounds is improved. At the same time, C2f-faster is used to optimize feature extraction, and a decoupled detection head is used to further optimize the real-time performance and accuracy.

[0058] In one embodiment, in order to improve the speed and accuracy of UAV infrared image detection, ConvNeXtV2 includes ConvNeXt Block. ConvNeXtV2 with efficient channel attention and dynamic convolution includes layer normalization, multi-stage ConvNeXt Block sampling, feature weight adjustment under the ECA attention mechanism, and regularization. Specifically, as Figure 4 shown, ConvNeXt v2 includes the following structure: First, layer normalization: the input data is subjected to layer normalization and preprocessing; Second, multi-stage ConvNeXt Block sampling: the forward propagation is divided into 4 stages. Considering the model accuracy and computational amount comprehensively, each stage is stacked by [d0, d1, d2, d3]=[3, 3, 9, 3] ConvNeXt Blocks respectively. After each stage, downsampling is performed and input to the next stage, and the output channel numbers are [96, 192, 384, 768] respectively.

[0059] Among them, as Figure 5 , the multi-stage ConvNeXt Block sampling includes the following steps: Step 1: Perform efficient convolution, decompose the standard convolution into depth convolution and pointwise convolution, and use a depth convolution layer (DWConv) with a 7*7 convolution kernel to reduce the computational cost.

[0060] Step 2: Process the features through a multi-layer perceptron MLP.

[0061] The multi-layer perceptron MLP includes the following steps: Step 2.1: Perform layer normalization LN, and expand the channels to 4*dim through a point convolution with k = 1; LN is implemented through the following formula:

[0062]

[0063]

[0064] Where xi represents the i-th element in the input tensor, μ is the mean of this layer, L is the number of elements to be normalized, that is, the total number of elements in each channel, σ2 is the variance, ϵ is a small value to prevent division by zero, and γ and β are trainable scaling and offset parameters.

[0065] Step 2.2: Pass through the GELU activation function;

[0066] Where represents the sigmoid function.

[0067] Step 2.3: Then perform global response normalization;

[0068] Gx is the L2 norm of the input, Nx is the value after normalizing Gx, y is the output result, and γ and β are trainable scaling and offset parameters.

[0069] By global response normalization, the feature collapse problem of the fully convolutional masked autoencoder pre-training model can be solved, and the feature cosine distance of the network can be stabilized:

[0070] Where C is the number of channels, and Xi and Xj are the features of channels i and j respectively.

[0071] Step 2.4: After a point convolution with k = 1, the number of channels is restored from 4*dim to dim.

[0072] ConvNeXtV2 with the ECA attention mechanism also includes the following structure: Third, feature weight adjustment under the ECA attention mechanism: ECA is an efficient channel attention mechanism for deep neural networks. After global average pooling, it compresses the spatial dimension and captures the dependencies between channels through one-dimensional convolution, avoiding complex dimensionality reduction and dimensionality increase processes. It considers each channel and its k nearest neighbors, performs local cross-channel interaction information, and realizes the characteristics of high efficiency and lightweight.

[0073] ECA uses the frequency band matrix Wk to learn channel information:

[0074] Among them, Wk involves k*C parameters, and Wk avoids complete independence among different groups. For the weights w1,1, w1,2....w1,k corresponding to yi, only the information interaction between yi and its k nearest neighbors is considered, and the calculation formula is as follows:

[0075] By sharing the same learning parameters, the information interaction between channels is achieved through 1D convolution with a convolution kernel size of k:

[0076] Among them, C1D represents 1D convolution, and this method is called the ECA module, which only involves k parameter information.

[0077] When the number of groups is fixed, the high-dimensional (low-dimensional) channels are proportional to the long-distance (short-distance) convolution, and the coverage range of the cross-channel information interaction (i.e., the kernel size k of the 1D convolution) is proportional to the channel dimension C. There is a mapping φ between i and C:

[0078] Given the channel dimension C, the convolution kernel size k can be determined adaptively:

[0079] That is

[0080] where γ and β are hyperparameters, and taking the absolute value and rounding down to the nearest odd number ensures that the kernel size is odd.

[0081] After obtaining the kernel size k, the ECA module applies 1D convolution to the input features to learn the importance of each channel relative to other channels. The output is the 1D convolution of the input with a kernel size of k.

[0082]

[0083] Fourth, regularization: Regularization is performed through drop_path, and the multi-branch structure in the deep learning model is randomly deleted.

[0084] The method of this application replaces the Backbone part with ConvNeXtV2 with the ECA attention mechanism added. By introducing the channel attention mechanism, the model can better focus on important features and suppress unimportant features, thereby improving the model detection accuracy.

[0085] In one embodiment, to improve the speed and accuracy of UAV infrared image detection, the pre-trained model includes pre-trained weights, and the UAV detection model includes YOLOv5 pre-trained weights. The step of greedily pre-training and fine-tuning the UAV detection model layer by layer according to the pre-trained model and the training set to obtain the improved UAV detection model can be specifically executed as follows: First, add the pre-trained weights and YOLOv5 pre-trained weights to the UAV detection model to obtain the first UAV detection model. Specifically, give the ConvNeXt pre-trained weights and YOLOv5 pre-trained weights to the Backbone part and the Head part respectively for pre-training to obtain the first UAV detection model.

[0086] Then, use the training set and the validation set to unfreeze and train the Backbone part of the first UAV detection model layer by layer to obtain the second UAV detection model. Specifically, when the Backbone part is frozen, use the training set to unfreeze and train each layer of the Backbone part layer by layer, and each layer of the data set is randomly divided according to a ratio each time. During the training process, use the validation set to verify to prevent overfitting of the model training.

[0087] Next, use the Bayesian optimizer to adjust the hyperparameters of the second UAV detection model to obtain the improved UAV detection model. Specifically, use the Bayesian optimizer to adjust the hyperparameters with the hyperparameters as the spatial scope.

[0088] The method of this application is an effective deep neural network training strategy through layer-by-layer greedy pre-training. By training layer by layer, it reduces the training complexity, alleviates the gradient problem, and learns the hierarchical features of the data, improving the training efficiency of the UAV detection model while ensuring the training quality.

[0089] In one embodiment, to improve the speed and accuracy of UAV infrared image detection, after adjusting the UAV detection model according to the pre-trained model and the training set to obtain the improved UAV detection model, the following steps can also be executed: First, input the test set into the improved UAV detection model to obtain the test result. Specifically, input the test set into the improved UAV detection model to obtain the test result, and the test result is output as a positive example or a negative example.

[0090] Then, evaluate the improved UAV detection model according to the test result.

[0091] In one embodiment, to improve the speed and accuracy of UAV infrared image detection, evaluating the improved UAV detection model according to the test result can be specifically executed as follows: First, obtain the true result of the test set, and calculate the evaluation index according to the test result and the true result. Specifically, the evaluation indexes are P, R, AP, and GFLOPs.

[0092] Where P represents precision; R represents recall; The calculation method is as follows: TP (True Positive): True positive example. That is, the predicted result is a positive example and the actual result is also a positive example.

[0093] FP (False Positive): False positive example. That is, the predicted result is a positive example while the actual result is a negative example.

[0094] TN (True Negative): True negative example. That is, the predicted result is a negative example and the actual result is also a negative example.

[0095] FN (False Negative): False negative example. That is, the predicted result is a negative example while the actual result is a positive example.

[0096]

[0097]

[0098] The value of AP (Average Precision) is the area under the corresponding Precision-Recall curve, which reflects the accuracy of the model; GFLOPs represents the computational complexity of the model, reflecting the occupancy of computing resources and the operation speed of the model.

[0099] Then, it is determined whether the evaluation indicators meet the requirements. If the evaluation indicators meet the requirements, it is determined that the improved UAV detection model can be put into use; otherwise, it is determined that the improved UAV detection model cannot be put into use. Specifically, it is determined whether the model can be put into use according to the preset evaluation indicator values, avoiding errors caused by overfitting of the model.

[0100] Meanwhile, this application also compares different models through ablation experiments, and the comparison results are as Figure 6 shown. The figure compares the performance of different model structures from dimensions such as P (precision), R (recall), AP (area under the Precision-Recall curve), and GFLOPs (computational complexity of the model). It can be seen from the figure that the performance of the UAV detection model in the method of this application is the best.

[0101] This application uses the test set to verify the obtained improved UAV detection model to ensure the accuracy of the improved UAV detection model.

[0102] Based on the above method, an embodiment of this application also discloses a UAV detection system improved based on YOLOv5. As Figure 7 , the system includes the following modules: The dataset acquisition module 701 is used to acquire a UAV infrared image dataset captured under complex backgrounds and occlusions; The dataset preprocessing module 702 is used to preprocess and annotate the UAV infrared image dataset to obtain an enhanced UAV infrared image dataset, and the enhanced UAV infrared image dataset includes a training set; The detection model construction module 703 is used to construct a UAV detection model based on YOLOv5; The pre-trained model construction module 704 is used to train the ImageNet dataset using a fully convolutional mask autoencoder to obtain a pre-trained model; The detection model training module 705 is used to greedily pre-train and fine-tune the UAV detection model layer by layer according to the pre-trained model and the training set to obtain an improved UAV detection model; The actual detection module 706 is used to detect the UAV infrared image dataset using the improved UAV detection model to obtain detection results.

[0103] In one embodiment, the dataset preprocessing module 702 is specifically used to perform data augmentation processing and annotation on the UAV infrared image dataset to obtain an enhanced UAV infrared image dataset; randomly divide the enhanced UAV infrared image dataset into a training set according to a ratio.

[0104] In one embodiment, the detection model construction module 703 is specifically used to construct an initial UAV detection model based on the improved YOLOv5. The initial UAV detection model includes a Backbone part, a Neck part, and a Head part; replace the Backbone part with ConvNeXtV2 with efficient channel attention and dynamic convolution; replace the feature extraction structure of the Neck part with a C2f-faster structure; use a lightweight decoupled detection head in the Head part to obtain a UAV detection model.

[0105] In one embodiment, the pre-trained model construction module 704 is specifically used to add pre-trained weights and YOLOv5 pre-trained weights to the UAV detection model to obtain a first UAV detection model; use the training set and the validation set to unfreeze and train the Backbone part, the Neck part, and the Head part of the first UAV detection model layer by layer to obtain a second UAV detection model; adjust the hyperparameters of the second UAV detection model using a Bayesian optimizer to obtain an improved UAV detection model.

[0106] The embodiments of the present application also disclose a computer device.

[0107] Specifically, the computer device includes a memory and a processor, and a computer program capable of being loaded and executed by the processor is stored on the memory, which is the above-mentioned method for improving UAV detection based on YOLOv5.

[0108] The embodiments of the present application also disclose a computer-readable storage medium.

[0109] Specifically, the computer-readable storage medium stores a computer program that can be loaded and executed by a processor, such as a computer program for an improved drone detection method based on YOLOv5 as described above. The computer-readable storage medium may include, for example: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0110] This specific embodiment is only an interpretation of the present invention and does not limit the present invention. After reading this specification, those skilled in the art can make modifications to this embodiment that do not contribute creatively as needed, but as long as they are within the scope of the claims of the present invention, they are protected by the patent law.

Claims

1. A drone detection method based on improved YOLOv5, characterized in that: The method is based on the ImageNet dataset and comprises: Obtain a dataset of drone infrared images taken under complex backgrounds and occlusions; Preprocessing and labeling the UAV infrared image dataset to obtain a UAV infrared image enhancement dataset, wherein the UAV infrared image enhancement dataset includes a training set; Build a drone detection model based on YOLOv5; Using a fully convolutional masked autoencoder to train the ImageNet dataset to obtain a pre-trained model; Greedily pre-training and fine-tuning the drone detection model layer by layer according to the pre-trained model and the training set to obtain an improved drone detection model; The improved UAV detection model is used to detect the UAV infrared image dataset to obtain a detection result.

2. The method according to claim 1, characterized in that: The preprocessing and labeling of the UAV infrared image dataset to obtain the UAV infrared image enhanced dataset includes: Performing data enhancement processing on the UAV infrared image dataset and annotating the dataset to obtain an enhanced UAV infrared image dataset; The UAV infrared image enhancement dataset is randomly divided into a training set in proportion.

3. The method according to claim 1, characterized in that: The construction of the drone detection model based on YOLOv5 includes: Construct an initial drone detection model based on improved YOLOv5. The initial drone detection model includes Backbone, Neck and Head parts. Replace the Backbone part with ConvNeXtV2 which adds efficient channel attention and dynamic convolution; The Neck part feature extraction structure is replaced by the C2f-faster structure; A lightweight decoupling detection head is used in the Head part to obtain a drone detection model.

4. The method according to claim 1, characterized in that: The ConvNeXtV2 includes ConvNeXt Block, and the ConvNeXtV2 with efficient channel attention and dynamic convolution includes layer normalization, multi-stage ConvNeXt Block sampling, feature weight adjustment and regularization under the ECA attention mechanism.

5. The method according to claim 1, characterized in that: The drone infrared image enhancement data set is further randomly divided into a validation set in proportion, the pre-trained model includes pre-trained weights, and the drone detection model includes YOLOv5 pre-trained weights; The method of greedily pre-training and fine-tuning the drone detection model layer by layer according to the pre-trained model and the training set to obtain an improved drone detection model includes: Adding the pre-trained weights and the YOLOv5 pre-trained weights to the drone detection model to obtain a first drone detection model; Unfreeze and train the Backbone part of the first drone detection model layer by layer using the training set and the validation set to obtain a second drone detection model; The improved drone detection model is obtained by adjusting the hyperparameters of the second drone detection model using a Bayesian optimizer.

6. The method according to claim 1, characterized in that: The UAV infrared image enhancement dataset is also randomly divided into a test set in proportion; After greedily pre-training and fine-tuning the drone detection model layer by layer according to the pre-trained model and the training set to obtain an improved drone detection model, the method further includes: Inputting the test set into the improved drone detection model to obtain a test result; The improved drone detection model is evaluated based on the test results.

7. The method according to claim 6, characterized in that: The evaluating the improved drone detection model according to the test results comprises: Obtaining the true result of the test set; Calculate the evaluation index according to the test results and the true results; Determine whether the evaluation indicators meet the requirements; If the evaluation index meets the requirements, it is determined that the improved drone detection model can be put into use; Otherwise, it is determined that the improved drone detection model cannot be put into use.

8. An improved drone detection system based on YOLOv5, characterized in that: The system comprises: The data set acquisition module (701) is used to acquire a data set of infrared images of drones taken under complex backgrounds and occlusions; A data set preprocessing module (702) is used to preprocess and annotate the drone infrared image data set to obtain a drone infrared image enhancement data set, wherein the drone infrared image enhancement data set includes a training set; A detection model building module (703) is used to build a drone detection model based on YOLOv5; A pre-trained model construction module (704) is used to train the ImageNet dataset using a fully convolutional masked autoencoder to obtain a pre-trained model; A detection model training module (705), configured to greedily pre-train and fine-tune the drone detection model layer by layer according to the pre-trained model and the training set to obtain an improved drone detection model; The actual detection module (706) is used to detect the drone infrared image data set using the improved drone detection model to obtain a detection result.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program which can be loaded by the processor and executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Infrared target detection method and device for self-supervised learning, equipment and storage medium

    CN120953574A