Improved YOLOv5 unmanned aerial vehicle aerial image detection method and system

By introducing feature refinement module and lightweight decoupling head module in the YOLOv5 model, the problems of poor detection of small targets and large computing resources in drone images are solved, and more efficient target detection and lower computing costs are achieved.

CN120164129APending Publication Date: 2025-06-17YANCHENG INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510214125.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

YOLOv5 has poor detection effect when facing small and dense targets and low resolution in drone images. Due to its large network structure and high computing requirements, it has a long training time and high computing resources, which is suitable for environments with limited resources.

Method used

Based on the YOLOv5 model, feature refinement module and lightweight decoupling head module are introduced to improve the target detection accuracy and reduce the calculation cost. The feature refinement module enhances feature extraction through attention mechanism, while the lightweight decoupling head module improves detection efficiency by reducing the amount of parameters and using the GSConv module.

Benefits of technology

The improved YOLOv5 model significantly improves the accuracy of small target detection in drone aerial images and reduces the consumption of computing resources, making it suitable for environments with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164129A_ABST
    Figure CN120164129A_ABST
Patent Text Reader

Abstract

The invention provides an improved YOLOv5 unmanned aerial vehicle aerial image detection method and system, and the method comprises the steps: dividing a Vi sDrone unmanned aerial vehicle aerial image data set into independent subsets according to a data proportion of 8: 1: 1; preprocessing the unmanned aerial vehicle aerial images in the independent subsets to obtain preprocessed image data; a YOLOv5 model is improved through a feature refining module and a lightweight decoupling head module, and an unmanned aerial vehicle aerial image detection model is trained based on the preprocessed image data and the improved YOLOv5 model; obtaining an aerial image of a to-be-detected unmanned aerial vehicle; and inputting a to-be-detected unmanned aerial vehicle aerial image into the unmanned aerial vehicle aerial image detection model to obtain a detection result. According to the improved YOLOv5 unmanned aerial vehicle aerial image detection method and system, a feature refining module is provided on the basis of a YOLOv5 model, and meanwhile, a lighter decoupling head is used, so that the target detection precision is improved, and the calculation cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly to a method and system for detecting UAV aerial images by improving YOLOv5. Background Art

[0002] YOLOv5 is the fifth version of the YOLO (You Only Look Once) series of object detection algorithms. The core idea of the YOLO algorithm is to transform the object detection task into a single neural network, and by predicting the bounding boxes and class information of objects simultaneously in a single forward pass, real-time object detection is achieved. YOLOv5 is an improved version based on YOLOv2, YOLOv3, and YOLOv4. It uses a Feature Pyramid Network to extract multi-scale features in order to better capture the feature information of objects at different scales. At the same time, YOLOv5 designs a CSP structure in the backbone network, which improves the effect of feature extraction. Compared with previous versions, YOLOv5 has improved both in terms of accuracy and speed. It can achieve a relatively fast inference speed while maintaining a high accuracy rate, and is suitable for real-time object detection scenarios. In addition, YOLOv5 has also optimized the detection effect of small objects and has good multi-scale feature fusion capabilities.

[0003] However, when faced with the situation of small and dense objects and low resolution in UAV images, YOLOv5s still has certain detection bottlenecks. The original YOLOv5 model cannot accurately detect objects with a resolution lower than 8×8. In addition, due to the large network structure and high computational requirements of YOLOv5, the training time is long and more computational resources are needed, which limits its application in some environments with limited resources.

[0004] In view of this, it is urgent to improve the method and system for detecting UAV aerial images by YOLOv5 to at least solve the above deficiencies. Summary of the Invention

[0005] One of the purposes of the present invention is to provide a method and system for detecting UAV aerial images by improving YOLOv5. On the basis of the YOLOv5 model, a feature refinement module is proposed, and at the same time, a more lightweight decoupled head is used, which improves the object detection accuracy and reduces the computational cost.

[0006] The method for detecting UAV aerial images by improving YOLOv5 provided by the embodiments of the present invention includes:

[0007] Step 1: Divide the VisDrone UAV aerial image dataset into independent subsets according to a data ratio of 8:1:1. The independent subsets are respectively: a training set, a validation set, and a test set;

[0008] Step 2: Preprocess the UAV aerial images in the independent subset to obtain preprocessed image data;

[0009] Step 3: Improve the YOLOv5 model through the feature refinement module and the lightweight decoupling head module, and train the UAV aerial image detection model based on the preprocessed image data and the improved YOLOv5 model;

[0010] Step 4: Obtain the UAV aerial image to be detected;

[0011] Step 5: Input the UAV aerial image to be detected into the UAV aerial image detection model to obtain the detection result.

[0012] Preferably, in Step 2: Preprocess the UAV aerial images in the independent subset to obtain preprocessed image data, including:

[0013] Align the UAV aerial images in the independent subset into RGB images of size 640*640 to obtain preprocessed image data.

[0014] Preferably, the feature refinement module includes:

[0015] Obtain the shallow input feature map X and the deep input feature map Y output by the backbone network, splice them and use convolution to enhance the learning ability to output the feature map F;

[0016] Assign channel attention weights to the feature map F in the channel dimension;

[0017] Assign spatial attention weights to the feature map F in the spatial dimension;

[0018] Multiply the feature map F by the channel attention weight and the spatial attention weight respectively and add them element by element.

[0019] Preferably, the lightweight decoupling head module is constructed as follows:

[0020] Add residual edges to two parallel branches of the YOLOX decoupling head and replace the standard convolution with the GSConv module to obtain the lightweight decoupling head module.

[0021] Preferably, in Step 4: Obtain the UAV aerial image to be detected, including:

[0022] Obtain the expected flight route and the real-time flight route of the first target UAV;

[0023] Calculate the route similarity between the real-time flight route and the expected flight route in real time;

[0024] If the route similarity is less than or equal to the preset route similarity threshold, determine the first timestamp;

[0025] Retrospect the control instruction sequence of the first target UAV within the preset first target duration based on the first timestamp;

[0026] Obtain the control instruction pair where the control source of adjacent control instructions in the control instruction sequence switches;

[0027] If the acquisition is successful, obtain the second timestamp of the control instruction triggered earlier and the third timestamp of the control instruction triggered later in the control instruction pair;

[0028] If the time length between the first timestamp and the second timestamp is less than or equal to the preset second target duration, obtain the source image for interference exclusion basis according to the second timestamp;

[0029] Obtain the interference exclusion basis based on the source image for interference exclusion and exclude the target interference;

[0030] When the target interference is excluded, obtain the aerial photography image of the UAV to be detected.

[0031] Preferably, obtaining the source image for interference exclusion basis according to the second timestamp includes:

[0032] Obtain the perception parameters of the first target UAV recorded at the second timestamp;

[0033] Identify the obstacle position based on the perception parameters;

[0034] If the obstacle position is successfully identified, obtain the obstacle position image captured by the UAV group based on the second timestamp and use it as the source image for interference exclusion basis;

[0035] If the obstacle position identification fails, use the first target UAV as the second target UAV;

[0036] Based on the position of the counter UAV of the second target UAV recorded at the second timestamp, determine the search area according to the interference range of the countermeasure device;

[0037] Based on the second timestamp, obtain the area image of the search area captured by the UAV group and use it as the source image for interference exclusion basis.

[0038] Preferably, determining the search area based on the position of the counter UAV of the second target UAV recorded at the second timestamp according to the interference range of the countermeasure device includes:

[0039] Obtain a spherical area with the position of the counter UAV as the center of the sphere and the maximum interference distance as the radius;

[0040] Determine the first intersection area between the spherical area and the ground area and use the first intersection area as the search area.

[0041] Preferably, based on the position of the counter UAV of the second target UAV recorded according to the second timestamp, and according to the interference range of the countermeasure device, a search area is determined, further including:

[0042] Narrow the search area according to the first intersection area;

[0043] Among them, narrowing the search area according to the first intersection area includes:

[0044] Obtain the second intersection area obtained by the mutual coverage of the first intersection areas;

[0045] Based on a preset regional association feature extraction template, and according to the second intersection area, extract a set of regional association feature values; the regional association feature values include: the number of regions of the first intersection area to which the second intersection area belongs, and the second timestamp when determining the first intersection area to which the second intersection area belongs;

[0046] According to the set of regional association feature values and a preset search sub-area determination rule base, determine the search sub-area corresponding to each second intersection area;

[0047] Integrate the search sub-areas to obtain the narrowed search area.

[0048] Preferably, obtaining an interference elimination basis and eliminating target interference based on an interference elimination basis source image includes:

[0049] Obtain the interference type corresponding to the interference elimination basis source image; the interference types include: obstacle interference and countermeasure device interference;

[0050] Obtain the historical interference elimination strategy for the interference type;

[0051] Train an interference elimination strategy determination model according to the historical interference elimination strategy;

[0052] Determine the standard basis item type according to the historical interference elimination strategy; the standard basis item types include: obstacle three-dimensional data, countermeasure device type, and countermeasure device distance;

[0053] Obtain the basis extraction template corresponding to the standard basis item type;

[0054] Based on the basis extraction template and according to the interference elimination basis source image, extract the interference elimination basis;

[0055] Input the interference elimination basis into the interference elimination strategy determination model to obtain the target interference elimination strategy;

[0056] Eliminate the target interference based on the target interference elimination strategy.

[0057] The improved UAV aerial image detection system based on YOLOv5 provided by the embodiments of the present invention includes:

[0058] The data division subsystem is used to divide the VisDrone UAV aerial image dataset into independent subsets according to a data ratio of 8:1:1. The independent subsets are: the training set, the validation set, and the test set;

[0059] The preprocessing subsystem is used to preprocess the UAV aerial images in the independent subsets to obtain preprocessed image data;

[0060] The training subsystem is used to improve the YOLOv5 model through the feature refinement module and the lightweight decoupling head module, and train the UAV aerial image detection model based on the preprocessed image data and the improved YOLOv5 model;

[0061] The image to be detected acquisition subsystem is used to acquire the UAV aerial image to be detected;

[0062] The detection result acquisition subsystem is used to input the UAV aerial image to be detected into the UAV aerial image detection model to obtain the detection result.

[0063] The beneficial effects of the present invention are as follows:

[0064] Based on the YOLOv5 model, the present invention proposes a feature refinement module. At the same time, a more lightweight decoupling head is used, which improves the target detection accuracy and reduces the computational cost.

[0065] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the structures specifically pointed out in this application document.

[0066] The technical solutions of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings

[0067] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings:

[0068] Figure 1 is a schematic diagram of the method for detecting UAV aerial images that improves YOLOv5 in the embodiments of the present invention;

[0069] Figure 2 is a schematic diagram of the feature refinement module in the embodiments of the present invention;

[0070] Figure 3 is a schematic diagram of the lightweight decoupling head module in the embodiments of the present invention;

[0071] Figure 4Schematic diagram of the UAV aerial image detection system for improving YOLOv5 in the embodiments of the present invention. Detailed implementation manners

[0072] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0073] The embodiments of the present invention provide a method for detecting UAV aerial images by improving YOLOv5, as Figure 1 shown, including:

[0074] Step 1: Divide the VisDrone UAV aerial image dataset into independent subsets according to the data ratio of 8:1:1. The independent subsets are: training set, validation set, and test set;

[0075] Step 2: Preprocess the UAV aerial images in the independent subsets to obtain preprocessed image data;

[0076] Among them, Step 2: Preprocess the UAV aerial images in the independent subsets to obtain preprocessed image data, including:

[0077] Align the UAV aerial images in the independent subsets into RGB pictures with a size of 640*640 to obtain preprocessed image data;

[0078] Step 3: Improve the YOLOv5 model through a feature refinement module and a lightweight decoupled head module, and train a UAV aerial image detection model based on the preprocessed image data and the improved YOLOv5 model;

[0079] Among them, the feature refinement module includes:

[0080] Obtain the shallow input feature map X and the deep input feature map Y output by the backbone network, splice them, and use convolution to enhance the learning ability to output the feature map F;

[0081] Assign channel attention weights to the feature map F in the channel dimension;

[0082] Assign spatial attention weights to the feature map F in the spatial dimension;

[0083] Multiply the feature map F by the channel attention weight and the spatial attention weight respectively, and then add them element by element;

[0084] Among them, the construction method of the lightweight decoupled head module is as follows:

[0085] Add residual edges to two parallel branches of the YOLOX decoupled head, and replace the standard convolution with the GSConv module to obtain the lightweight decoupled head module;

[0086] Step 4: Obtain the aerial images of the drone to be detected; the aerial images of the drone to be detected are the aerial images of the drone that need to be subjected to target detection;

[0087] Step 5: Input the aerial images of the drone to be detected into the drone aerial image detection model to obtain the detection results.

[0088] The working principle and beneficial effects of the above technical solution are as follows:

[0089] The present invention trains a drone aerial image detection model based on an improved YOLOv5 model. The specific process is as follows:

[0090] First, divide the VisDrone drone aerial image dataset into a training set, a validation set, and a test set according to a ratio of 8:1:1, and then align the image data in each subset into RGB images of size 640*640;

[0091] During training, input the preprocessed image data corresponding to the training set into the backbone network; further filter out useless features through the feature refinement module in the Neck layer. According to the three-layer output in the network, continue to output three feature maps of different sizes in the head layer through the backbone network. The schematic diagram of the feature refinement module is as Figure 2 shown. Assign attention weights to the fused feature map in the spatial dimension and the channel dimension respectively to further refine the target feature information and suppress the conflicts brought by the fusion of features of different scales;

[0092] Among them, the calculation formula of the channel attention weight is as follows:

[0093]

[0094] By performing global average pooling and global maximum pooling operations on the feature map F, the information in the spatial dimension is compressed, so as to pay more attention to the channel information, such as feature information such as the color and texture of the target. Then, perform element-wise addition on the output fed into the shared network, and finally generate the channel attention weight after the Sigmoid operation;

[0095] Among them, the output formula of the spatial attention weight is as follows:

[0096] S w =σ(f 7×7 [AvgPool(F):MaxPool(F)])

[0097] In the above formula, f 7×7It is a 7×7 convolution operation. First, global average pooling and global max pooling based on channels are performed on the feature map. By compressing channel information, the image pays more attention to the target-related regions. Then, a dimensionality reduction operation using a 7×7 convolution is carried out. Finally, the Sigmoid activation is applied to the output feature map to obtain the spatial attention weights, which pay more attention to the small target object regions, thereby reducing the possibility of false detection and missed detection;

[0098] Finally, the output of the feature refinement module is:

[0099] F = f 1×1 ([X:Y])

[0100] Z = F·C w +F·S w

[0101] X and Y are the shallow input feature map and the deep input feature map respectively. After simple concatenation, a 1×1 convolution is used to enhance the learning ability to output the feature map F; F passes through the C w module and S w module to obtain the corresponding weights. The distribution of attention can be reflected by the parameters of the weight matrix, enabling the model to accurately obtain more effective features; Finally, the feature map F is multiplied by the channel weight C w and the spatial weight S w respectively, and then added element by element; The AFRM module emphasizes focusing on more meaningful features in the channel dimension and the spatial dimension respectively, reducing information redundancy in the fusion process;

[0102] In addition, a new lightweight decoupled head is constructed. The traditional YOLOX decoupled head contains a 1×1 convolution dimensionality reduction operation and two parallel branches, each branch having two 3×3 convolution modules for classification and regression tasks respectively. Although it greatly improves the convergence speed and performance of the model, the huge number of parameters has been difficult to meet the real-time requirements. Therefore, the present invention further improves on the basis of the YOLOX decoupled head. The lightweight decoupled head module (LD-Head) is as Figure 3As shown, since each additional convolution brings a multiplicative increase in the number of parameters, resulting in a decrease in the model's detection speed, the LD-Head structure of the present invention simplifies the additional 3×3 convolution modules in the two branches. At the same time, a residual connection is added to the convolution module of each branch, and the GSConv module is introduced to replace the standard convolution, further reducing the network parameters; GSConv is a convolution module that combines standard convolution and depthwise separable convolution. First, a dimensionality reduction operation of standard convolution is performed on the input feature map, then the same operation is performed using grouped convolution, and the results of the two convolutions are concatenated. Finally, through the Shuffle operation, the feature information of the two convolutions is mixed with each other; using the lightweight convolution method of GSConv can significantly reduce the computational cost while having a model learning ability not inferior to that of the standard convolution. LD-Head reduces the network optimization difficulty by utilizing the residual idea and adopts a more lightweight convolution method to improve the detection performance without affecting the real-time performance of the model.

[0103] After the training stops, the validation set is used to assist in adjusting the model hyperparameters to obtain the finally trained model to be evaluated. Then, the test set is used to finally evaluate the performance of the model, and the model to be evaluated that meets the preset evaluation criteria is used as the UAV aerial image detection model.

[0104] Finally, the obtained UAV aerial image to be detected is input into the UAV aerial image detection model to obtain the detection result.

[0105] The present invention proposes a feature refinement module based on the YOLOv5 model. At the same time, a more lightweight decoupled head is used to improve the object detection accuracy and reduce the computational cost.

[0106] In one embodiment, step 4: obtaining the UAV aerial image to be detected includes:

[0107] Obtaining the expected flight route and the real-time flight route of the first target UAV; the first target UAV is: a single UAV in the UAV group, and the UAV group is all UAVs used for aerial photography tasks; the expected flight route is: the pre-set aerial flight route of the first target UAV; the real-time flight route is: the actual flight route of the first target UAV; due to the triggering of the active obstacle avoidance system of the first target UAV and other interferences (such as: low-altitude aircraft countermeasure devices), its actual route and the preset route may not be the same.

[0108] Calculating the route similarity between the real-time flight route and the expected flight route in real time; wherein, the route similarity is: the route trajectory similarity.

[0109] If the route similarity is less than or equal to a preset route similarity threshold, determine the first timestamp; where the preset route similarity threshold is set in advance by a human; the first timestamp is: a reliable time mark recorded when the route similarity is less than or equal to the preset route similarity threshold;

[0110] Based on the first timestamp, trace back the control instruction sequence of the first target UAV within a preset first target duration; the preset first target duration is set in advance by a human, for example: 5 minutes; the control instruction sequence is: the instruction sequence for controlling the first target UAV, sorted in the order of the time when the instructions are issued;

[0111] Obtain the control instruction pair where the control source of adjacent control instructions in the control instruction sequence switches; the control source is: the source party that issues the control instruction, such as: the ground control center, the active avoidance system built in the UAV, and unknown source parties, etc.; a switch means that the source party changes;

[0112] If the acquisition is successful, obtain the second timestamp of the control instruction triggered first and the third timestamp of the control instruction triggered later in the control instruction pair;

[0113] If the time length between the first timestamp and the second timestamp is less than or equal to a preset second target duration, then according to the second timestamp, obtain the source image for interference exclusion basis; the preset second target duration is set in advance by a human, for example: 10 seconds;

[0114] The statement that if the time length between the first timestamp and the second timestamp is less than or equal to a preset second target duration, then according to the second timestamp, obtain the source image for interference exclusion basis, includes:

[0115] Obtain the perception parameters of the first target UAV recorded at the second timestamp; the perception parameters include: the environmental data obtained by the first target UAV through sensors (radar, camera, infrared sensor, etc.);

[0116] Based on the perception parameters, identify the obstacle position; the obstacle position is: the specific position information of the obstacle identified through the perception parameters;

[0117] If the obstacle position is successfully identified, based on the second timestamp, obtain the obstacle position image captured by the UAV group and use it as the source image for interference exclusion basis;

[0118] If the obstacle position identification fails, use the first target UAV as the second target UAV;

[0119] Based on the position of the counter-drone of the second target drone recorded at the second timestamp, determine the search area according to the interference range of the countermeasure device; the position of the counter-drone is the position of the second target drone recorded at the second timestamp; the interference range of the countermeasure device is: the spatial range of the largest effective interference drone among the known countermeasure devices.

[0120] Among them, based on the position of the counter-drone of the second target drone recorded at the second timestamp, determine the search area according to the interference range of the countermeasure device, including:

[0121] Obtain a spherical area with the position of the counter-drone as the center of the sphere and the maximum interference distance as the radius;

[0122] Determine the first intersection area between the spherical area and the ground area, and use the first intersection area as the search area;

[0123] Based on the second timestamp, obtain the area image of the search area captured by the drone swarm and use it as the source image for interference exclusion basis;

[0124] Obtain the interference exclusion basis based on the source image for interference exclusion and exclude the target interference; the interference exclusion basis is: obstacle information and countermeasure device information; obtaining the interference exclusion basis based on the source image for interference exclusion and excluding the target interference means that when the interference type is countermeasure device interference, the drone travels along the expected flight route overcoming the interference of the countermeasure device; when the interference type is obstacle interference, after the drone actively avoids, it then travels along the expected flight route from the route point of the nearest expected flight route.

[0125] When the target interference is excluded, obtain the aerial image of the drone to be detected.

[0126] The working principle and beneficial effects of the above technical solution are:

[0127] The image used for target detection as the detection basis greatly affects the subsequent detection results. Therefore, it is necessary to specifically analyze the acquisition process of the aerial image of the drone to be detected in the area to be aerial-photographed. Generally, the staff will design the flight route of the area to be aerial-photographed, and the first target drone can fly and take pictures according to the designed route. However, the actual route of the first target drone will change according to the unexpected situations it encounters. Therefore, it is necessary to specifically consider different unexpected situations and introduce the route similarity. When the route similarity is less than or equal to the preset route similarity threshold, at this time, the degree of route deviation is already relatively large, indicating that an unexpected situation has occurred. Record the current first timestamp and trace back the control instruction sequence of the first target drone based on the first timestamp to improve the tracing efficiency.

[0128] Obtain control instruction pairs in the control instruction sequence where the control source of adjacent control instructions switches. If the control source switches, locate the control instruction pair where the control source switches. The second timestamp of the control instruction triggered earlier and the third timestamp of the control instruction triggered later in the control instruction pair are the key time points of interference. At this moment, the perception data of other drones in the drone swarm except the first target drone, in addition to being used as subsequent drone aerial images to be detected, can also assist in determining the interference information of the first target drone, improving the degree of information reuse;

[0129] If the interference type is obstacle interference, the first target drone can directly perceive the obstacle position. After determining the obstacle position, obtain the image of the obstacle position (obstacle position image) captured by the drone swarm at the second timestamp as the source image for interference elimination basis; if the avoidance is not due to identifying an obstacle, then at this time the second target drone may be interfered by other devices. Determine the search area based on the interference range of the known countermeasure device and the position of the countermeasure drone of the second target drone; based on the second timestamp, obtain the area image of the search area captured by the drone swarm as the source image for interference elimination basis; determine the source image for interference elimination basis based on different perception situations of the first target drone at the second timestamp, improving the suitability of subsequent interference elimination and further improving the rationality of subsequent obtaining of the detection basis (i.e., the drone aerial image to be detected).

[0130] In one embodiment, based on the position of the countermeasure drone of the second target drone recorded at the second timestamp, and according to the interference range of the countermeasure device, determining the search area further includes:

[0131] Narrow the search area according to the first intersection area;

[0132] Among them, narrowing the search area according to the first intersection area includes:

[0133] Obtain the second intersection area obtained by the mutual coverage of the first intersection areas;

[0134] Based on a preset region association feature extraction template, extract a set of region association feature values according to the second intersection area; the region association feature values include: the number of regions of the first intersection area to which the second intersection area belongs, and the second timestamp when determining the first intersection area to which the second intersection area belongs; the region association feature extraction template is used to extract the number of regions of the first intersection area to which the second intersection area belongs and the second timestamp when determining the first intersection area to which the second intersection area belongs by comparing the first intersection area and the second intersection area;

[0135] Determine the search sub-region corresponding to each second intersection region according to the region association eigenvalue set and the preset search sub-region determination rule library; the preset search sub-region determination rule library stores multiple one-to-one corresponding region association eigenvalue sets and search sub-region determination rules. For example, set a region number threshold (for example: 2). If the region number is greater than the preset region number threshold, use the corresponding second intersection region as the search sub-region; if the region number is less than the region number threshold, use the first intersection region to which the corresponding second intersection region belongs as the third intersection region; determine whether the second timestamps when determining the third intersection region are the same; if they are the same, use the corresponding second intersection region as the search sub-region; if they are not the same, use the corresponding third intersection region as the search sub-region.

[0136] Integrate the search sub-regions to obtain a reduced search region.

[0137] The working principle and beneficial effects of the above technical solution are as follows:

[0138] The search region range is too large, which reduces the search efficiency of the countermeasure device. Therefore, the present invention further reduces the first intersection region. Specifically, introduce a region association feature extraction template to extract the number of regions of the first intersection region to which the second intersection region belongs and the second timestamp when determining the first intersection region to which the second intersection region belongs; match the region association eigenvalue set in the region association eigenvalue set and the search sub-region determination rule library to determine the search sub-region corresponding to the second intersection region, reducing the data volume of the source image for subsequent interference elimination basis and improving the extraction efficiency of subsequent interference elimination basis.

[0139] In one embodiment, step 4: Obtain the aerial image of the drone to be detected, and further include:

[0140] Obtain the subsequent featureization rules for the aerial image of the drone to be detected; wherein, the subsequent featureization rules are: how to perform subsequent feature engineering tasks based on the aerial image of the drone to be detected, and the subsequent feature engineering tasks are: featureize the aerial image of the drone to be detected to obtain the tasks of feature extraction for the input of the subsequent drone aerial image detection model, such as: detecting the texture of the detection target and the shape of the detection target, etc.

[0141] Determine the featureization elements according to the subsequent featureization rules; wherein, the featureization elements are: operations related to the featureization behavior, such as: what kind of preprocessing to perform and what kind of featureization technology to use.

[0142] Obtain the element - pointing feature types of the characterized elements. The element - pointing feature types include: identification features and potential features. Among them, the identification feature is: a feature describing the region of interest, such as: the shape feature and the edge feature of the region of interest, etc. The potential feature is: a feature obtained by quantitatively analyzing and capturing subtle changes in the image, such as: first - order statistical features (obtained by statistically analyzing the pixel intensity values within the ROI, including mean, variance, maximum, minimum, median, skewness, and kurtosis, etc.), texture features (reflecting the spatial relationship and texture information between pixels in the image, such as the gray - level co - occurrence matrix), and high - order features (model - based features, such as: wavelet - transform - based features and local binary patterns, etc.);

[0143] If the element - pointing feature type is an identification feature, perform data cleaning, data conversion, outlier processing, and data standardization on the corresponding drone aerial image to be detected;

[0144] If the element - pointing feature type is a potential feature, obtain the first image setting parameters of the data to be amplified. Among them, the data to be amplified is: the drone aerial image to be detected that is prepared for amplification. The first image setting parameters are: the setting parameters of the shooting device of the source drone of the data to be amplified;

[0145] Obtain the second image setting parameters that meet the potential feature extraction standard. Among them, meeting the potential feature extraction standard is: the image parameter setting standard that could extract potential features historically. The second image setting parameters are: the image parameter settings corresponding to the image parameter setting standard that could extract potential features historically;

[0146] Calculate the parameter coverage degree of the first image setting parameters with respect to the second image setting parameters, and generate an expansion requirement vector according to the parameter coverage degree. Among them, the parameter coverage degree is: the coverage degree of the parameter range of the first image setting parameters covering the parameter range of the second image setting parameters. The expansion requirement vector is: a representation vector of the expansion requirement, generated according to the missing parameter range part of the first image setting parameters with respect to the second image setting parameters;

[0147] Obtain pre - amplified data according to the expansion requirement vector. Among them, the pre - amplified data is: the amplified image data obtained according to the missing image setting parameters;

[0148] Determine the missing angle according to the identification features of the amplified images in the pre - amplified data. Among them, according to the identification features, determine the supplementary shooting angle required for the pre - amplified data as the missing angle;

[0149] Rotate the amplified image according to the missing angle to obtain a rotated image;

[0150] Extract potential features using the rotated image.

[0151] The working principle and beneficial effects of the above technical solution are as follows:

[0152] The aerial images of the UAV to be detected need to be characterized subsequently before they can be used for model output. However, the types of features obtained in the characterization task are different, and the appropriate data preprocessing strategies for each feature for the aerial images of the UAV to be detected are also different. Therefore, the present invention obtains the subsequent characterization rules for the aerial images of the UAV to be detected, and determines the characterization elements according to the subsequent characterization rules. According to the characterization elements, the data preprocessing strategy before feature extraction of the aerial images of the UAV to be detected is determined;

[0153] Specifically, the element types of the characterization elements include recognition features and potential features. Since the types of the element types of the characterization elements are different, the data preprocessing strategies are different. Conducting conventional preprocessing operations directly lacks pertinence. Therefore, if the element type of the characterization element is a recognition feature, since the recognition feature is qualitative and can be directly recognized from the image, its data preprocessing strategy is data cleaning, data conversion, outlier processing, and data standardization; if the element type of the characterization element is a potential feature, the feature is usually not directly visible and needs to be extracted through complex calculations and analyses. Therefore, its data preprocessing strategy is data augmentation. By adaptively processing according to the different types of the element types of the characterization elements, the rationality of formulating the data preprocessing strategy is improved;

[0154] Since there may be a problem of insufficient sample size in the image data of potential features, data augmentation techniques (such as mirroring, rotation, and scaling, etc.) can be used to increase sample diversity. However, blind augmentation will lead to the problem of too large training data. Therefore, the parameter coverage degree is determined according to the first image setting parameters of the data to be augmented and the second image setting parameters of the aerial images of the UAV to be detected, an expansion requirement vector is generated based on the parameter coverage degree, and pre-augmented data is obtained according to the expansion requirement vector. According to the recognition features of the augmented images in the pre-augmented data, the missing angle is determined, the augmented images are rotated according to the missing angle to obtain rotated images, and the potential features are extracted by using the rotated images, which improves the extraction accuracy of the potential features and the subsequent image detection results are also more accurate.

[0155] In one embodiment, obtaining the interference exclusion basis from the source image based on the interference exclusion basis and excluding the target interference includes:

[0156] Obtaining the interference type corresponding to the source image of the interference exclusion basis; the interference types include: obstacle interference and countermeasure device interference;

[0157] Obtain the historical interference elimination strategy for the interference type; the historical interference elimination strategy includes: the corresponding relationship between the obstacle distance, the three-dimensional data of the obstacle, the historical expected flight route, and the avoidance plan when the drone recognized an obstacle and triggered the active avoidance system in history, and, when the drone was interfered by a countermeasure device in history, the corresponding relationship between the historical countermeasure device type, the countermeasure device distance, and the drone anti-countermeasure strategy;

[0158] Train an interference elimination strategy determination model according to the historical interference elimination strategy;

[0159] Determine the standard basis item type according to the historical interference elimination strategy; the standard basis item type includes: the three-dimensional data of the obstacle, the countermeasure device type, and the countermeasure device distance;

[0160] Obtain the basis extraction template corresponding to the standard basis item type; the basis extraction template is the basis extraction rule corresponding to the standard basis item type, for example: the rule for extracting the three-dimensional data of the obstacle by fusing the interference elimination basis source images of drones in different directions; another example: the standard process of extracting the eigenvalue set of the interference elimination basis source image and matching it with the eigenvalue set corresponding to the preset countermeasure device type;

[0161] Extract the interference elimination basis based on the basis extraction template and according to the interference elimination basis source image;

[0162] Input the interference elimination basis into the interference elimination strategy determination model to obtain the target interference elimination strategy;

[0163] Eliminate the target interference based on the target interference elimination strategy.

[0164] The working principle and beneficial effects of the above technical solution are as follows:

[0165] The elimination logic for different interference types is different. For example: when the interference type is obstacle interference, the interference elimination strategy is to avoid obstacles based on the recognized three-dimensional data of the obstacle; when the interference type is countermeasure device interference, the interference elimination strategy is to resist countermeasures according to the countermeasure device type and the countermeasure device distance, and the interference type can be known in advance according to the recognition result of the recognized obstacle position. Therefore, according to the different interference types, select the corresponding historical interference elimination strategy to train the interference elimination strategy determination model; after training, extract the interference elimination basis according to the basis extraction template corresponding to the standard basis item type required to input the interference elimination strategy determination model and input it into the model to obtain the output target interference elimination strategy. The system can adaptively extract the interference elimination basis of the corresponding interference type according to the interference elimination basis source image and input it into the corresponding strategy determination model, improving the timeliness of interference elimination and further improving the rationality of obtaining the subsequent aerial photography images of the drone to be detected.

[0166] The embodiment of the present invention provides a drone aerial image detection system for improving YOLOv5, as Figure 4 shown, including:

[0167] A data division subsystem 1, configured to divide the VisDrone drone aerial image dataset into independent subsets according to a data ratio of 8:1:1. The independent subsets are respectively: a training set, a validation set, and a test set;

[0168] A preprocessing subsystem 2, configured to preprocess the drone aerial images in the independent subsets to obtain preprocessed image data;

[0169] A training subsystem 3, configured to improve the YOLOv5 model through a feature refinement module and a lightweight decoupled head module, and train a drone aerial image detection model based on the preprocessed image data and the improved YOLOv5 model;

[0170] A to-be-detected image acquisition subsystem 4, configured to acquire a to-be-detected drone aerial image;

[0171] A detection result acquisition subsystem 5, configured to input the to-be-detected drone aerial image into the drone aerial image detection model to obtain a detection result.

[0172] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. Improved YOLOv5 drone aerial image detection method, characterized in that: include: Step 1: Divide the VisDrone drone aerial image dataset into independent subsets according to the data ratio of 8:1:

1. The independent subsets are: training set, validation set and test set. Step 2: Preprocess the drone aerial images in the independent subset to obtain preprocessed image data; Step 3: Improve the YOLOv5 model through the feature refinement module and the lightweight decoupling head module, and train the drone aerial image detection model based on the preprocessed image data and the improved YOLOv5 model; Step 4: Obtain the aerial image of the drone to be detected; Step 5: Input the drone aerial image to be detected into the drone aerial image detection model to obtain the detection result.

2. The improved YOLOv5 unmanned aerial image detection method according to claim 1, characterized in that: Step 2: Preprocess the drone aerial images in the independent subset to obtain preprocessed image data, including: The drone aerial images in the independent subset are aligned into 640*640 RGB images to obtain preprocessed image data.

3. The improved YOLOv5 unmanned aerial image detection method according to claim 1, characterized in that: Feature refinement module, including: Get the shallow input feature map X and deep input feature map Y output by the backbone network, concatenate them and use convolution to enhance the learning ability to output feature map F; Assign channel attention weight to the feature map F in the channel dimension; Assign spatial attention weights to the feature map F in the spatial dimension; Multiply the channel attention weight and spatial attention weight of the feature map F respectively and add them element by element.

4. The improved YOLOv5 unmanned aerial image detection method according to claim 1, characterized in that: The lightweight decoupling head module is constructed as follows: We add residual edges to the two parallel branches of the YOLOX decoupling head and replace the standard convolution with the GSConv module to obtain a lightweight decoupling head module.

5. The improved YOLOv5 unmanned aerial image detection method according to claim 1, characterized in that: Step 4: Obtain the aerial image of the drone to be detected, including: Obtaining the expected flight route and real-time flight route of the first target UAV; Calculate the route similarity between the real-time flight route and the expected flight route in real time; If the route similarity is less than or equal to a preset route similarity threshold, determining a first timestamp; Backtracking a control instruction sequence of a first target UAV within a preset first target duration based on a first timestamp; Acquire a control instruction pair in which control sources of adjacent control instructions in a control instruction sequence are switched; If the acquisition is successful, the second timestamp of the control instruction triggered earlier and the third timestamp of the control instruction triggered later in the control instruction pair are acquired; If the time length between the first timestamp and the second timestamp is less than or equal to a preset second target time length, obtaining the interference elimination basis source image according to the second timestamp; Obtain interference elimination basis based on the source image of interference elimination basis and eliminate target interference; When the target interference is eliminated, the aerial image of the UAV to be detected is obtained.

6. The improved YOLOv5 unmanned aerial image detection method according to claim 5, characterized in that: According to the second timestamp, an interference elimination basis source image is obtained, including: Obtaining perception parameters of the first target UAV recorded at the second timestamp; Based on the perception parameters, the obstacle location is identified; If the obstacle position is successfully identified, the obstacle position image taken by the drone group is obtained based on the second timestamp and used as the source image for interference elimination; If the obstacle position recognition fails, the first target UAV is used as the second target UAV; Determine a search area based on the countermeasure UAV position of the second target UAV recorded at the second timestamp and according to the interference range of the countermeasure device; Based on the second timestamp, an area image of the search area taken by the drone group is obtained and used as a source image for interference elimination.

7. The improved YOLOv5 unmanned aerial image detection method according to claim 6, characterized in that: Based on the countermeasure UAV position of the second target UAV recorded at the second timestamp, and according to the interference range of the countermeasure device, a search area is determined, including: Get a spherical area with the counter drone position as the center and the maximum interference distance as the radius; A first intersection area between the spherical area and the ground area is determined, and the first intersection area is used as a search area.

8. The improved YOLOv5 unmanned aerial image detection method according to claim 7, characterized in that: Based on the countermeasure UAV position of the second target UAV recorded at the second timestamp, determining the search area according to the interference range of the countermeasure device, further comprising: According to the first intersection area, narrow the search area; Wherein, narrowing the search area according to the first intersection area includes: Obtain a second intersection area obtained by overlapping the first intersection areas; Based on a preset region association feature extraction template and according to the second intersection region, extracting a region association feature value set; the region association feature value includes: the region number of the first intersection region to which the second intersection region belongs, and a second timestamp when determining the first intersection region to which the second intersection region belongs; Determine the search sub-region corresponding to each second intersection region according to the region-associated feature value set and a preset search sub-region determination rule base; The search sub-areas are integrated to obtain a reduced search area.

9. The improved YOLOv5 unmanned aerial image detection method according to claim 5, characterized in that: Obtain interference elimination basis based on the source image of interference elimination basis and eliminate target interference, including: Obtain interference elimination based on the interference type corresponding to the source image; interference types include: obstacle interference and countermeasure device interference; Get historical interference elimination strategies of interference types; According to the historical interference elimination strategy, the interference elimination strategy determination model is trained; According to the historical interference elimination strategy, determine the standard basis item type; the standard basis item type includes: obstacle three-dimensional data, countermeasure device type and countermeasure device distance; Get the basis extraction template corresponding to the standard basis item type; Extract interference elimination basis based on the basis extraction template and the interference elimination basis source image; Determine the model of interference elimination based on the input interference elimination strategy to obtain the target interference elimination strategy; Eliminate target interference based on target interference elimination strategy.

10. Improved YOLOv5 drone aerial image detection system, characterized by: include: The data partitioning subsystem is used to divide the VisDrone drone aerial image dataset into independent subsets according to the data ratio of 8:1:

1. The independent subsets are: training set, validation set and test set. A preprocessing subsystem is used to preprocess the drone aerial images in the independent subset to obtain preprocessed image data; The training subsystem is used to improve the YOLOv5 model through the feature refinement module and the lightweight decoupling head module, and train the drone aerial image detection model based on the preprocessed image data and the improved YOLOv5 model; The image acquisition subsystem to be detected is used to obtain the aerial image of the drone to be detected; The detection result acquisition subsystem is used to input the drone aerial image to be detected into the drone aerial image detection model to obtain the detection result.