A Low-Altitude Security Target Detection Method and System Based on Deep Learning

Through the improved AD-YOLOv5s model, combined with ghost module and CBAM attention, the feature fusion network and detection head are optimized, and the problem of insufficient detection accuracy of micro UAVs in low-altitude security is solved, and low-cost, high-reality low-altitude security target detection is achieved.

CN114792390BActive Publication Date: 2025-07-22HOHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210186976.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-07-22
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

The existing low-altitude security solutions are insufficient in detecting micro drones, and the equipment is large in size and high in deployment costs, making real-time end-to-end output impossible, especially in complex electromagnetic environments.

Method used

The AD-YOLOv5s model is built, and feature extraction is extracted by introducing ghost module and CBAM attention module, improving feature fusion network and detection head structure, combined with TensorRT optimization acceleration, which is suitable for embedded devices.

Benefits of technology

It realizes high-precision detection of micro drones, reduces equipment costs, and has real-time capabilities and the ability to adapt to complex electromagnetic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114792390B_ABST
    Figure CN114792390B_ABST
Patent Text Reader

Abstract

The present invention discloses a low-altitude security target detection method and system based on deep learning. The present invention constructs an AD-YOLOv5s model for low-altitude security target detection, and optimizes and accelerates the trained model to be applicable to embedded devices. The AD-YOLOv5s model constructed by the present invention is improved on the basis of the YOLOv5s model. In the feature extraction network, a ghost module is introduced to construct a ghost-bottleneck CSP structure, which is used to replace the original bottleneck CSP structure. At the same time, a CBAM attention module is introduced to improve the detection accuracy. In the feature fusion network, an upsampling operation is added to generate a larger-size feature map, as well as a corresponding PAN structure, to suit the task of detecting small targets in low-altitude security. Compared with the prior art, the present invention has the advantages of high detection accuracy, good real-time performance, and low deployment cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of target detection and deep learning technology, and more specifically, to a low-altitude security target detection method and system, which can be applied to the field of low-altitude security. Background Art

[0002] Most of the existing low-altitude security solutions currently use radio detection technology to detect drones. Radio uses frequency bands to detect the presence of drones, and then uses optoelectronic systems to track and detect targets. For drones that fly silently, radio detection technology cannot detect them, and radar is usually required for additional detection.

[0003] However, with the development of artificial intelligence technology, the existing low-altitude security solutions can no longer meet the current growing low-altitude security needs. The current low-altitude security requirements urgently require an end-to-end fast and real-time detection solution, but the existing detection equipment is bloated and has harsh deployment conditions, especially in the complex electromagnetic environment in the city. The detection accuracy is greatly suppressed, and the operation is complicated, and real-time end-to-end output cannot be achieved. Therefore, low-altitude security target detection based on artificial intelligence technology came into being.

[0004] For small, low-altitude civilian drones, the more commonly used detection methods are still radio, audio, radar, and traditional images. The radio method can quickly locate the drone, but if the airborne station and the ground station change stations at the same time, that is, frequency hopping, the detection accuracy will be reduced, and its deployment cost is high. The audio method can effectively detect drones with high noise, but it is not suitable for low-noise drones. Radar has shown excellent performance in detecting large-sized drones, but radar cannot meet the detection requirements of tiny drones, and its deployment cost is high. Traditional image detection methods rely on manual features based on prior knowledge and experience, which are difficult and inefficient, so they may not be applicable in terms of speed and accuracy. Summary of the invention

[0005] Purpose of the invention: The purpose of the present invention is to provide a low-altitude security target detection solution based on deep learning, which has the advantages of high detection accuracy, good real-time performance, and low deployment cost.

[0006] Technical solution: To achieve the above-mentioned invention object, the present invention provides a low-altitude security target detection method based on deep learning, comprising the following steps:

[0007] Build an AD-YOLOv5s model for low-altitude security target detection and use the drone detection data set for training; optimize and accelerate the trained network model; use the optimized model to perform target detection on the collected image data;

[0008] Among them, the AD-YOLOv5s model is improved based on the YOLOv5s model, including: in the feature extraction network, a ghost module is introduced, combining the ideas of the bottleneck structure and the CSP structure to construct a new module, the ghost-bottleneckCSP structure, which replaces the original bottleneckCSP structure, and a CBAM attention module is introduced before each ghost-bottleneckCSP structure to complete the weighted extraction of the feature map using channel attention and spatial attention; in the feature fusion network, the performance of the model is first improved by means of feature enhancement. The specific operation is to perform an additional upsampling operation after two upsampling operations to generate a feature map of the fourth size, and a new PAN structure from the fourth size to the third size feature map is added, and the PAN structure from the second size to the first size is deleted; in the detection head structure, the detection head with the feature map size of the first size is deleted, and a detection head of the fourth size is added; where the sizes of the first to fourth sizes increase in sequence.

[0009] Preferably, the ghost-bottleneckCSP structure divides the input into two branches. One branch passes through a standard GBL module, then through a ghost-bottleneck structure, and batch normalization operations are used to reduce internal covariate shift, and a concat operation is performed with the other branch passing through a conventional convolution; the standard GBL module is a ghost module + batch normalization operation + Leaky Relu activation function; the ghost-bottleneck structure is mainly composed of two stacked ghost modules. The first ghost module is used as an expansion layer to increase the number of channels; the second is used to reduce the number of channels to match the direct connection path; batch normalization operations and Leaky Relu activation functions are added after the first ghost module, and only batch normalization is added to the second module; after passing through two ghost modules, a residual structure is used to perform feature superposition with the input.

[0010] Preferably, the CBAM attention module enriches the extracted high-level features in the channel dimension and the spatial dimension respectively by taking global average pooling operations and global max pooling operations, and after obtaining the spatial and channel weights respectively, they are weighted to the initial features to complete the dual attention adjustment of the features.

[0011] Preferably, in the feature fusion network, depthwise separable convolutions are introduced to replace the original convolution operations.

[0012] Preferably, the drone detection dataset for training the model is improved based on the existing UAVdataset. Duplicate images are deleted, and images containing drone flights are collected and labeled by referring to public data, and data augmentation is performed to expand the dataset.

[0013] Preferably, the trained AD-YOLOv5s model is optimized and accelerated through the TensorRT optimizer.

[0014] Preferably, the optimized and accelerated model is deployed on an embedded device.

[0015] Based on the same inventive concept, a low-altitude security target detection system provided by the present invention includes: a model construction and training unit for constructing an AD-YOLOv5s model for low-altitude security target detection and training it using the drone detection dataset; an optimization processing unit for optimizing and accelerating the trained network model; and a target detection unit for performing target detection on the collected image data using the optimized model. The AD-YOLOv5s model is improved based on the YOLOv5s model, including: in the feature extraction network, a ghost module is introduced, combining the ideas of the bottleneck structure and the CSP structure to construct a new ghost-bottleneckCSP structure, which replaces the original bottleneckCSP structure, and a CBAM attention module is introduced before each ghost-bottleneckCSP structure to complete weighted extraction of the feature map using channel attention and spatial attention; in the feature fusion network, the performance of the model is first improved by means of feature enhancement. The specific operation is to perform an additional upsampling operation after two upsampling operations to generate a feature map of the fourth size, and a new PAN structure from the fourth size to the third size feature map is added, and the PAN structure from the second size to the first size is deleted; in the detection head structure, the detection head with the feature map size of the first size is deleted, and a detection head of the fourth size is added; the sizes of the first to fourth sizes increase in sequence.

[0016] Based on the same inventive concept, the present invention provides a low-altitude security target detection device based on deep learning. A trained and optimized AD-YOLOv5s model for low-altitude security target detection is stored in the device, which is used to perform target detection on the collected image data. The AD-YOLOv5s model is improved on the basis of the YOLOv5s model, including: in the feature extraction network, a ghost module is introduced, combining the ideas of the bottleneck structure and the CSP structure to construct a new module, the ghost-bottleneckCSP structure, and it is used to replace the original bottleneckCSP structure. And a CBAM attention module is introduced before each ghost-bottleneckCSP structure to complete the weighted extraction of the feature map by using channel attention and spatial attention; in the feature fusion network, the performance of the model is first improved by means of feature enhancement. The specific operation is to perform one more upsampling operation after two upsampling operations to generate a feature map of the fourth size, and a new PAN structure from the fourth size to the third size feature map is added, and the PAN structure from the second size to the first size is deleted; in the detection head structure, the detection head with the feature map size of the first size is deleted, and a detection head of the fourth size is added; where the sizes of the first to fourth sizes increase in sequence.

[0017] Based on the same inventive concept, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, the steps of the above-mentioned low-altitude security target detection method based on deep learning are implemented.

[0018] Beneficial effects: Compared with the prior art, the present invention has the following technical effects: 1. Compared with the existing low-altitude security solutions, the present invention can avoid the frequency hopping interference of the radio solution, and does not need to consider the noise level of the unmanned aerial vehicle, as well as the incompetence of the radar for micro unmanned aerial vehicles. 2. Compared with the existing image detection solutions, the present invention introduces deep learning, which can automatically extract features and achieve end-to-end output. At the same time, in view of the characteristics of the limited available resources of embedded devices, lightweight improvements are made by using the ghost module and the adjustment of the neck network layer and the detection head structure; in addition, in view of the size of the unmanned aerial vehicle, the CBAM attention mechanism is introduced, so that more attention is paid to the position of the target to be detected during feature extraction, and at the same time, a micro target detection head is added to improve the performance of micro target detection. Description of the Drawings

[0019] Figure 1 It is a schematic diagram of the network model structure of the improved AD-YOLOv5s in the embodiment of the present invention.

[0020] Figure 2It is a schematic diagram of ghost convolution.

[0021] Figure 3 It is a schematic diagram of the ghost-bottleneck and ghost-bottleneckCSP structures constructed in the embodiment of the present invention.

[0022] Figure 4 It is a schematic diagram of the improved feature extraction network in the embodiment of the present invention.

[0023] Figure 5 It is a statistical result chart of target size classification in the embodiment of the present invention.

[0024] Figure 6 It is a schematic diagram of the network structure of the feature fusion layer.

[0025] Figure 7 It is a schematic diagram of the improved network structure of the feature fusion layer in the implementation of the present invention.

[0026] Figure 8 It is a schematic diagram of the improved detection head structure in the implementation of the present invention. Detailed implementation manners

[0027] The following further describes the solution of the present invention in conjunction with the accompanying drawings and specific embodiments.

[0028] The embodiment of the present invention discloses a low-altitude security target detection method based on deep learning. First, a UAV detection data set is collected, then a model for low-altitude security target detection is constructed and trained using the UAV detection data set. Then, the trained network model is optimized and accelerated to adapt to embedded devices. Finally, after the optimized model is deployed on the development board, the acquired image data can be used for UAV target detection. The following details the specific implementation process.

[0029] First, by referring to the existing UAV detection data sets, it is found that the UAV dataset provided by Google has the following problems: it contains a large number of repeated and similar pictures, which easily causes the trained model to overfit. To solve this problem, on the basis of the UAV dataset, the present invention collects and annotates images containing UAV flights. The main way to obtain the data set is to refer to the publicly available network data sets and write a crawler script to crawl. The images obtained by the crawler are unannotated. This part of the data is annotated using the labelimg tool for the acquired images. The main annotation information includes the category, size, and the position of the target in the image. At the same time, the copy and paste data augmentation method is used to expand the data set, and a new low-altitude security data set AD dataset containing 7500 pictures is obtained.

[0030] Secondly, the YOLOv5s model is selected as the basic network for improvement. The improved YOLOv5 object detection model is called AD-YOLOv5s, and its network structure is mainly divided into three parts, namely the feature extraction network (backbone), the feature fusion network (neck), and the detection head (prediction head). Among them, the feature extraction network is responsible for extracting features from the input image. The feature fusion network is between the feature extraction network and the detection head, aiming to better utilize the features extracted by the backbone. The detection head is the network that obtains the output content of the network and makes predictions using the previously extracted features. The network structure diagram is as shown in Figure 1 shown.

[0031] For the existing feature extraction network in the present invention, there are problems of large network depth and large model parameter occupation, and lightweight improvement is carried out. Considering that the existing feature extraction network uses a large number of convolutional modules, occupying a large amount of computing power, as shown in Figure 2 shown, the present invention uses the method of ghost convolution for convolution operation. First, a small number of feature maps are generated by conventional convolution with less computational effort, and then through linear operations, fewer feature maps are further used to generate new similar feature maps. Finally, the information in the two groups of feature maps is combined and output as all the feature information, achieving the purpose of reducing the computational effort. The following is the comparison of the convolutional computational effort of the two methods.

[0032] Assume that the input feature map is H×W×C in , the size of the output feature map is H′×W′, the number of output channels is m×r, and the size of the convolution kernel is k×k. The computational effort required for conventional convolution is:

[0033] (m×r)×H′×W′×C in ×k×k

[0034] If the method of ghost convolution is adopted, it is necessary to ensure that the convolution kernel size, stride, and padding remain unchanged, and in the first step of operation, m feature maps are generated, and each feature map generates r - 1 new feature maps through mapping, then a total of m×(r - 1) feature maps are generated in the second part.

[0035] Among them, the computational effort of the first part is:

[0036] m×H′×W′×C in ×k×k

[0037] The computational effort of the second part is:

[0038] m×(r - 1)×H′×W′×k×k

[0039] Then the total computational effort of the method of ghost convolution is:

[0040] m×H′×W′×C in ×k×k + m×(r - 1)×H′×W′×k×k

[0041] The computational complexity ratio between the conventional convolution and the Ghost convolution methods is:

[0042]

[0043] Generally, C in >> r, then

[0044]

[0045] Through the above calculation, it can be obtained that assuming that the conventional convolution generates 36 feature maps, using the Ghost convolution method, 6 feature maps are generated in the first step, and then each feature map is mapped to generate 5 similar feature maps. Then, it can be concluded from the above formula that the computational complexity of the model is reduced by 6 times. Using the Ghost convolution method can effectively solve the situation that the generation of redundant feature maps occupies a large amount of computing resources.

[0046] The feature extraction network in YOLOv5 is stacked by four layers of conventional convolutions and three Cross Stage Partial Network with Convolution (BottleneckCSP) networks. Its network depth is relatively deep. If the Ghost module is directly used to replace the original BottleneckCSP structure and conventional convolution, firstly, it will bring a large amount of model parameter computational complexity and increase the complexity of the model; secondly, it will lead to repeated gradient information and even the occurrence of gradient disappearance problems, which is not conducive to the effective extraction of features.

[0047] To solve the above problems, based on the Ghost module, the present invention draws on the structural design of the basic residual block of Resnet and the CSP structure in the Cross Stage Partial Network (CSPNet) to construct the Ghost-bottleneck and Ghost-bottleneckCSP structures, as Figure 3 shown.

[0048] Use the Ghost module to build the Ghost-bottleneck structure, as Figure 3As shown in (a). The ghost-bottleneck mainly consists of two stacked ghost modules. The first ghost module serves as an expansion layer to increase the number of channels, and the second is used to reduce the number of channels to match the direct (shortcut) path. Then, the input and output of these two ghost modules are connected using a shortcut. After the first ghost module, batch normalization and the Leaky Relu activation function are added, while only batch normalization is added to the second module. After passing through the two ghost modules, the residual structure is used to superimpose features with the input, which can enhance the gradient value of backpropagation between layers, avoid the vanishing gradient caused by the increase in model depth, and thus extract more fine-grained features without worrying about network degradation.

[0049] After constructing the ghost-bottleneck structure, in order to further reduce the computational bottleneck and solve the problem of vanishing gradient information, the CSP structure is introduced to form the ghost-bottleneck CSP structure, as Figure 3 shown in (b). The input is divided into two branches. One branch passes through a standard GBL module, that is, a ghost module + batch normalization operation + Leaky Relu activation function, and then passes through the above-mentioned ghost-bottleneck structure. After using batch normalization to reduce internal covariate shift and accelerating the network training process, a concat operation is performed with the other branch that passes through a conventional convolution. The design of this CSP branch is beneficial to reducing the computational bottleneck and memory consumption, and improving the problem of vanishing gradient, so that more abundant feature information can be extracted.

[0050] Based on the above analysis, the present invention uses the constructed ghost-bottleneck CSP structure to replace the original bottleneck CSP structure. At the same time, considering that the ghost module generates similar feature maps to meet the amount of feature information, which will lead to a decrease in the accuracy of the model, a convolutional block attention module (CBAM) attention mechanism is introduced into the improved feature extraction network.

[0051] Combining the above changes, the improved ghost-bottleneck CSP structure is used to replace the original feature extraction network, and the CBAM attention mechanism is added before the ghost-bottleneck CSP structure to obtain the improved feature extraction network, as Figure 4 shown.

[0052] For the dataset of low-altitude security target detection, first, according to the receptive field mapping relationship of the convolutional neural network Where: X represents the side length of the input image; Y represents the side length of the feature map generated after downsampling; N represents the number of pixel points in the original image corresponding to a point on the feature map. When the input image size is 640*640, the target size classification table of the dataset to be detected is formulated as follows.

[0053] Table 1 Target Size Classification Table

[0054]

[0055] It can be seen from this table that for targets with a size larger than 8*8 pixels, there are corresponding detection layers. For targets smaller than 8*8, the present invention defines them as tiny targets. It can be seen that the original detection layer does not set a detection layer for tiny targets. By classifying and counting the target sizes of the AD dataset mentioned above, we get Figure 5 .

[0056] It can be seen that the AD dataset contains a large number of tiny targets, while the number of large targets can be ignored by comparison.

[0057] Based on the above information, the present invention improves the target size classification table as follows.

[0058] Table 2 Improved Target Size Classification Table

[0059]

[0060] Based on the analysis of the low-altitude security target detection scenario, the present invention improves the feature fusion layer by means of feature enhancement.

[0061] The network structure of the feature fusion layer of YOLOv5 adopts the FPN+PAN structure, which adds a bottom-up feature pyramid behind the FPN layer, including two PAN structures, as Figure 6 shown.

[0062] The principle of the feature fusion layer is as follows. After the feature map with an input size of 160*160 is extracted layer by layer through the feature extraction layer, its size is compressed to 20*20, losing the shallow feature information and retaining the semantic feature information. Then, in the FPN layer, two upsampling operations are used to perform a concat operation with the feature maps of corresponding sizes in the feature extraction layer, enabling the strong semantic feature information to be transmitted from top to bottom while fusing the corresponding shallow feature information. After the FPN layer, the feature map is compressed again using two layers of PAN structure and concatenated with the feature maps of corresponding sizes in the FPN layer. Through such a structure setting, the FPN layer conveys strong semantic features from top to bottom, while the PAN structure conveys strong localization features from bottom to top. The two work together to perform feature aggregation on different detection layers from different feature extraction layers, further improving the feature aggregation ability.

[0063] According to the above analysis, the feature maps used in the detection head layer are the feature maps with sizes of 80*80 obtained after 8-fold downsampling, 40*40 obtained after 16-fold downsampling, and 20*20 obtained after 32-fold downsampling. The sizes of these three feature maps correspond to the target size classification table before optimization. As can be seen from the above, in order to match the optimized target size classification table, the existing feature fusion layer needs to be improved.

[0064] The network structure of the improved feature fusion layer is as Figure 7 shown. It can be seen from the figure that first, an additional upsampling operation is added after the two upsampling operations in the FPN structure, improving the FPN layer from the initial three-layer to a four-layer pyramid structure, and performing a concat operation with the feature maps of corresponding sizes in the feature extraction layer to obtain a feature map with a size of 160*160 for detecting tiny targets. Although the 20*20 feature map is not used in the prediction layer, the 20*20 feature map is retained in the FPN layer, enabling more detailed semantic feature information to be extracted and transmitted through the FPN structure to other feature maps, which is beneficial for subsequent recognition and detection of small targets. At the same time, two PAN structures are added after the 160*160 feature map to generate feature maps with sizes of 80*80 and 40*40 respectively. The FPN layer conveys strong semantic features from top to bottom, and the PAN structure conveys strong localization features from bottom to top to perform feature enhancement operations on the feature fusion layer. Considering that there is no detection requirement for large targets larger than 32*32, the PAN structure that generates the 20*20 feature map is deleted.

[0065] At the same time, in order to match the improvement of the feature fusion layer, we further improved the detection head layer, as Figure 8As shown in the figure. From the figure, we can learn that first, two detection heads responsible for detecting medium and small targets from feature maps of sizes 80*80 and 40*40 are retained, and the detection head responsible for detecting large targets and the related sampling convolution process, which are not required on the low-altitude security dataset proposed in the present invention, are deleted. At the same time, a detection head dedicated to detecting tiny targets is introduced using the feature map of size 160*160 mentioned above.

[0066] After feature enhancement improvement in the feature fusion layer, depthwise separable convolution is introduced to reduce the number of model parameters.

[0067] After the above improvements at the model level, the tensorRT acceleration technology is used to accelerate the model at the hardware level. Through the above acceleration method, the improved model is accelerated and deployed on the jetson nano development board to complete the process of data collection and processing.

[0068] The present invention introduces deep learning and constructs an AD-YOLOv5s model for target recognition of low-altitude security collected images. In the feature extraction network, the original bottleneckCSP structure is replaced with the ghost-bottleneckCSP structure, and a lightweight CBAM attention module is introduced. Channel attention and spatial attention are used to complete the weighted extraction of the feature map, improving the detection accuracy. In the feature fusion network, first, the performance of the model is improved by means of feature enhancement. After two upsampling operations, one more upsampling operation is performed to generate a larger-size feature map, and a corresponding PAN structure is added, while the PAN structure of the original small-size feature map is deleted. In the detection head structure, a corresponding lightweight three-detection-head structure is formed. In addition, in the feature fusion network, depthwise separable convolution is further introduced to replace the original convolution operation, achieving the purpose of reducing the model calculation amount. The present invention analyzes the target size of the collected drone detection dataset, improves the existing target size classification table, and makes it suitable for low-altitude security detection tasks. And the TensorRT technology is introduced to accelerate the improved model, making it better applicable to embedded devices with few available resources. Through the above means, the present invention can be well applied to low-altitude security target detection, with the advantages of high detection accuracy, good real-time performance, and low deployment cost.

[0069] The effects of the present invention will be described below in combination with comparative experiments. The research environment and initial settings of this experiment are shown in Tables 3 and 4.

[0070] Table 3 Research environment and initial settings table

[0071]

[0072] Table 4 Initial parameter settings table

[0073]

[0074] The training process uses the Adam optimizer for training, with the initial learning rate set to 0.01, the weight decay to 0.0001, the momentum to 0.9, and the batch size set to 16. Single-scale training is adopted in all experiments, and the image input size is 640 * 640 pixels. According to the characteristics of the model itself, the pre-trained model is yolov5s.pt. Each experiment runs for 200 epochs.

[0075] Experiment 1: Comparison experiment on object detection performance

[0076] As shown in Table 5, it is the data comparison between the improved YOLOv5s model based on feature enhancement proposed by the present invention and the original model. Among them, the model with only the CBAM module embedded is named CBAM - YOLOv5s, the model with the optimized feature fusion layer is named improved - YOLOv5s, and the model obtained by combining structural lightweight improvement and feature enhancement is AD - YOLOv5s.

[0077] Table 5 Comparison detection results of low - altitude security datasets

[0078]

[0079] As can be seen from the table, by introducing the CBAM attention mechanism, the mAP value, Recall, and mAP @0.5:0.95 value have all been slightly improved, which shows the feasibility of introducing the CBAM mechanism. At the same time, for the improved model improved - YOLOv5s of the feature fusion layer neck, its mAP @0.5:0.95 value has been improved to 84.5%, proving the improvement effect of the feature enhancement method on the detection effect of small targets. Finally, for the overall improved AD - YOLOv5s, its mAP value, Recall, and mAP @0.5:0.95 value have all been improved, especially the improvement of the Recall value, verifying the positive significance of the improvement of the EIoU loss function for the recall improvement of small targets.

[0080] Experiment 2: Comparison experiment on model lightweight

[0081] First, analyze the network structure as follows.

[0082] Table 6 YOLOv5 network framework

[0083]

[0084]

[0085] The improved AD-YOLOv5s network framework is as follows.

[0086] Table 7 AD-YOLOv5s Network Framework

[0087]

[0088]

[0089] By calculation, the number of parameters and GFLOPs of the two network models are obtained, as shown in Table 8.

[0090] Table 8 Number of Model Parameters, GFLOPs

[0091]

[0092] The total number of parameters of the improved network model is approximately 4.32 million, and the number of parameters of the original YOLOv5s network model is approximately 7.06 million. The number of parameters of the lightweight model is reduced by 38.8% compared to the original model. The GFLOPs of the AD-YOLOv5s model is 70.1% of the original model. Experiments show that: the number of parameters and computational complexity of the AD-YOLOv5s network model proposed in the present invention are significantly lower than those of the YOLOv5s network model.

[0093] Experiment 1 is a comparative test for the optimization of improving accuracy. This experiment will conduct a comparative test on the lightweight optimization of the model's speed. This experiment processes the model through the TensorRT module. The present invention converts the original YOLOv5s and the optimized model introduced with the ghost module and depthwise separable convolution into trt files (models obtained after being accelerated by TensorRT), and uniformly uses the trt suffix in the research to represent the accelerated model, and conducts tests. The obtained detection results are shown in the table.

[0094] Table 9 Test Comparison Results of the Low-Altitude Security Dataset after Lightweighting

[0095]

[0096] As can be seen from the table, by introducing TensorRT acceleration, the speed of each network model has been significantly improved. The YOLOv5s-trt with TensorRT acceleration has a frame rate increased to 20.5fps compared with the original YOLOv5s network, which is about twice the speed improvement before acceleration. At the same time, the mAP value only decreases by 1.3%, proving that the impact of introducing TensorRT on detection accuracy can be ignored. Compared with YOLOv5s-trt, the detection accuracy of Ghost-bottleneckCSP-YOLOv5s-trt only decreases by 2.8%, but the number of parameters decreases by 24.3%, which shows the feasibility of introducing the ghost module. Compared with YOLOv5s-trt, the number of parameters of DWConv-YOLOv5s-trt decreases by 7.3%, and the detection accuracy is almost the same without obvious decline, verifying the role of depthwise separable convolution in reducing the computational amount and maintaining the convolution effect.

[0097] The optimization method proposed in the present invention has significantly improved the YOLOv5s algorithm in embedded devices and already has the effect of real-time detection, which can be applied to actual project engineering. The above results show that the optimized model is more adaptable to the needs of actual scenarios, such as the situations over cities and airports, and has certain practical application value.

[0098] Based on the same inventive concept, an airborne security target detection system based on deep learning provided by an embodiment of the present invention includes: a model construction and training unit for constructing an AD-YOLOv5s model for airborne security target detection and training it using an unmanned aerial vehicle detection data set; an optimization processing unit for optimizing and accelerating the trained network model; and a target detection unit for performing target detection on the collected image data using the optimized model.

[0099] Based on the same inventive concept, an airborne security target detection device provided by the present invention stores a trained and optimized-accelerated AD-YOLOv5s model for airborne security target detection, which is used to perform target detection on the collected image data.

[0100] Based on the same inventive concept, an electronic device provided by the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, the steps of the above-mentioned airborne security target detection method based on deep learning are implemented.

Claims

1. A low-altitude security target detection method based on deep learning, characterized in that, The steps are as follows: Construct an AD-YOLOv5s model for low-altitude security target detection and train it using a UAV detection dataset; optimize and accelerate the trained network model; Use the optimized model to perform target detection on the collected image data; The AD-YOLOv5s model is improved based on the YOLOv5s model, including: in the feature extraction network, introduce the ghost module, combine the ideas of the bottleneck structure and the CSP structure to construct a new module, the ghost-bottleneckCSP structure, and use it to replace the original bottleneckCSP structure, and introduce the CBAM attention module before each ghost-bottleneckCSP structure to complete the weighted extraction of the feature map using channel attention and spatial attention; in the feature fusion network, first improve the performance of the model by means of feature enhancement. The specific operation is to perform an additional upsampling operation after two upsampling operations to generate a feature map of the fourth size, and add a new PAN structure from the fourth size to the third size feature map, and delete the PAN structure from the second size to the first size; in the detection head structure, delete the detection head with the feature map size of the first size and add a detection head of the fourth size; The sizes of the first to fourth sizes increase in sequence; The ghost-bottleneckCSP structure divides the input into two branches. One branch passes through a standard GBL module, then through the ghost-bottleneck structure, and uses batch normalization operations to reduce internal covariate shift, and performs a concat operation with the other branch passing through a conventional convolution; the standard GBL module is a ghost module + batch normalization operation + Leaky Relu activation function; the ghost-bottleneck structure is mainly composed of two stacked ghost modules. The first ghost module is used as an expansion layer to increase the number of channels; the second is used to reduce the number of channels to match the direct connection path; a batch normalization operation and a Leaky Relu activation function are added after the first ghost module, and only batch normalization is added to the second module; after passing through the two ghost modules, use the residual structure to perform feature superposition with the input.

2. The method for low-altitude security target detection based on deep learning according to claim 1, characterized in that The CBAM attention module enriches the extracted high-level features in the channel dimension and the spatial dimension respectively by taking global average pooling operations and global max pooling operations, and after obtaining the spatial and channel weights respectively, weights them to the initial features to complete the dual attention adjustment of the features.

3. The method for low-altitude security target detection based on deep learning according to claim 1, wherein In the feature fusion network, introduce depthwise separable convolution to replace the original convolution operation.

4. The method for low-altitude security target detection based on deep learning according to claim 1, characterized in that, The UAV detection dataset for training the model is improved on the existing UAV detection dataset UAV dataset. Duplicate pictures are deleted, and images containing UAV flights are collected and labeled by consulting public data, and data augmentation is performed to expand the dataset.

5. The low-altitude security target detection method based on deep learning according to claim 1, characterized in that The trained AD-YOLOv5s model is optimized and accelerated through the TensorRT optimizer.

6. The method for low-altitude security target detection based on deep learning according to claim 1, wherein, The optimized and accelerated model is deployed on an embedded device.

7. A low-altitude security target detection system based on deep learning, characterized in that, Including: A model construction and training unit that constructs an AD-YOLOv5s model for low-altitude security target detection and trains it using a drone detection dataset; An optimization processing unit for optimizing and accelerating the trained network model; And a target detection unit for performing target detection on the collected image data using the optimized model; Among them, the AD-YOLOv5s model is improved based on the YOLOv5s model, including: in the feature extraction network, a ghost module is introduced, combining the ideas of the bottleneck structure and the CSP structure to construct a new module, the ghost-bottleneckCSP structure, and using it to replace the original bottleneckCSP structure, and a CBAM attention module is introduced before each ghost-bottleneckCSP structure to complete the weighted extraction of the feature map using channel attention and spatial attention; in the feature fusion network, the performance of the model is first improved by means of feature enhancement. The specific operation is to perform one more upsampling operation after two upsampling operations to generate a feature map of the fourth size, and a new PAN structure from the fourth size to the third size feature map is added, and the PAN structure from the second size to the first size is deleted; in the detection head structure, the detection head with the feature map size of the first size is deleted, and a detection head of the fourth size is added; Among them, the sizes of the first to fourth sizes increase in sequence; the ghost-bottleneckCSP structure divides the input into two branches. One branch passes through a standard GBL module, then through a ghost-bottleneck structure, and uses batch normalization operations to reduce internal covariate shift, and performs a concat operation with the other branch that passes through a conventional convolution; the standard GBL module is a ghost module + batch normalization operation + Leaky Relu activation function; the ghost-bottleneck structure is mainly composed of two stacked ghost modules. The first ghost module is used as an expansion layer to increase the number of channels; the second is used to reduce the number of channels to match the direct connection path; a batch normalization operation and a Leaky Relu activation function are added after the first ghost module, and only batch normalization is added to the second module; after passing through the two ghost modules, a residual structure is used to perform feature superposition with the input.

8. An object detection device for low-altitude security based on deep learning, characterized in that, The device stores a trained and optimized and accelerated AD-YOLOv5s model for low-altitude security target detection, which is used to perform target detection on the collected image data; Among them, the AD-YOLOv5s model is improved based on the YOLOv5s model, including: in the feature extraction network, a ghost module is introduced, combining the ideas of the bottleneck structure and the CSP structure to construct a new module, the ghost-bottleneck CSP structure, and using it to replace the original bottleneck CSP structure, and a CBAM attention module is introduced before each ghost-bottleneck CSP structure to complete the weighted extraction of the feature map using channel attention and spatial attention; in the feature fusion network, first, the performance of the model is improved by means of feature enhancement. The specific operation is to perform an additional upsampling operation after two upsampling operations to generate a feature map of the fourth size, and a new PAN structure from the fourth size to the third size feature map is added, and the PAN structure from the second size to the first size is deleted; in the detection head structure, the detection head with the feature map size of the first size is deleted, and a detection head of the fourth size is added; Among them, the sizes of the first to fourth sizes increase in sequence; for the ghost-bottleneck CSP structure, the input is divided into two branches. One branch passes through a standard GBL module, then through a ghost-bottleneck structure, and batch normalization operation is used to reduce the internal covariate shift, and a concat operation is performed with the other branch passing through a conventional convolution; the standard GBL module is a ghost module + batch normalization operation + Leaky Relu activation function; the ghost-bottleneck structure is mainly composed of two stacked ghost modules. The first ghost module is used as an expansion layer to increase the number of channels; the second is used to reduce the number of channels to match the direct connection path; batch normalization operation and Leaky Relu activation function are added after the first ghost module, and only batch normalization is added to the second module; after passing through two ghost modules, the residual structure is used to perform feature superposition with the input.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements the steps of the deep learning-based low-altitude security target detection method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Target detection method combined with lightweight network

    CN113011365A

  • Lightweight aircraft detection method based on improved Yolov4-tiny

    CN113780211A