Anti-unmanned aerial vehicle detection method and system based on deep learning

By building a convolutional neural network that combines random noise and combining multiple voiceprint feature fusion and attention mechanisms, the problem of low detection accuracy in complex environments is solved in the existing technology, and the accuracy and stability of drone audio detection are achieved.

CN120220732AInactive Publication Date: 2025-06-27CHINA TELECOM UNMANNED TECHNOLOGY (JIANGSU) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510700241.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing drone audio detection technology has low detection accuracy in complex environments, making it difficult to effectively identify small drones and distinguish their voiceprint characteristics.

Method used

A deep learning-based method is adopted to construct a drone audio data set that integrates random noise, and drone audio detection is performed by fusion of multiple voiceprint features of Fbank features, MFCC features and GFCC features, combining a convolutional neural network that increases attention mechanism and residual structure.

Benefits of technology

By simulating noise interference in complex environments, the detection accuracy and stability of the model in practical applications can be improved, the recognition ability of drone audio is enhanced, and the error detection rate is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220732A_ABST
    Figure CN120220732A_ABST
Patent Text Reader

Abstract

The invention relates to an anti-unmanned aerial vehicle detection method and system based on deep learning. The method comprises the following steps: constructing an unmanned aerial vehicle audio data set fused with random noise; random noise in the unmanned aerial vehicle audio data set is fused based on the Fbank feature, the MFCC feature and the GFCC feature; and constructing a convolutional neural network based on an attention increasing mechanism and a residual structure, and detecting the audio of the unmanned aerial vehicle. And a random noise adding data enhancement mode is adopted. The working condition of the algorithm in an actual complex environment is simulated by adding various kinds of noise in a simulated real environment into training data. According to the mode, the model can learn and cope with diversified noise interference in the training stage, and the defect that the detection performance is reduced in a complex noise environment due to pure data set training in the prior art is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of anti-drone detection, and in particular, to an anti-drone detection method and system based on deep learning. Background Art

[0002] With the rapid development of modern technology, the application of drone technology has become increasingly widespread. However, with the sharp increase in the number of drones, a series of security issues have followed. Especially in some densely populated public areas and places with extremely high security requirements, the improper flight or even malicious intrusion of drones may trigger serious security accidents and potential threats. Therefore, it is particularly important to conduct real-time and effective monitoring of drones. Currently, common drone detection methods mainly rely on camera and radar technologies. Traditional camera-based detection methods identify drones by analyzing images. However, as drones continue to develop in the direction of miniaturization, their size is getting smaller and smaller. In some cases, cameras may be unable to capture clear enough image features to detect the presence of these small drones, especially when the drones are in complex environmental backgrounds, the accuracy and reliability of detection are greatly reduced.

[0003] In the prior art, drones can also be detected through drone audio detection technology. Generally, the collected drone sound signals are preprocessed first, and the spectral features of the sound signals are obtained through operations such as frame segmentation and spectral transformation; then the extracted features are input into a convolutional neural network to detect and classify the drone audio to determine whether it is the target drone and its category.

[0004] However, most of the existing drone audio detection works analyze using a pure data set without background interference. In an ideal pure data environment, the model can more accurately identify the audio features of drones. But in the actual complex environment, urban sirens, engine sounds, etc., and natural wind sounds, rain sounds, etc. will also be mixed in, resulting in lower detection accuracy of the existing drone audio detection technology in the actual application process. Summary of the Invention

[0005] 1. Problems to be Solved Based on this, it is necessary to provide an anti-drone detection method and system based on deep learning that can improve the detection accuracy of drone audio detection technology for the above technical problems.

[0006] 2. Technical Solutions In the first aspect, this application provides an anti-drone detection method based on deep learning. The method includes: Construct a drone audio data set fused with random noise; Fuse the random noise in the UAV audio dataset based on Fbank features, MFCC features, and GFCC features; Construct a convolutional neural network based on an enhanced attention mechanism and a residual structure to detect UAV audio.

[0007] In one embodiment, constructing a UAV audio dataset that fuses random noise includes: Determine the types and sources of noise; Divide the noise into training data and test data; Preprocess the UAV audio data; Enhance the training data of the noise; Add noise to the training data within a preset signal-to-noise ratio range and construct a UAV audio dataset that fuses random noise.

[0008] In one embodiment, enhancing the training data of the noise includes: Randomly select noise samples from the training data of the noise; Sample the noise samples and keep the sampling length of the noise samples consistent with the sampling length of the UAV audio; Perform random sampling of the signal-to-noise ratio within a preset signal-to-noise ratio range; Add noise to the UAV audio data based on the average power of the UAV audio signal and the average power of the sampled noise.

[0009] In one embodiment, constructing a convolutional neural network based on an enhanced attention mechanism and a residual structure includes: Define two repeated convolutional blocks, and each convolutional block includes a convolutional layer, a BN normalization layer, and a ReLU activation layer; The convolutional layer is used to receive the UAV audio and extract the features of the UAV audio dataset; The BN normalization layer performs batch normalization on the features of the UAV audio dataset output by the convolutional layer; The ReLU activation layer applies a non-linear activation function to the normalized data.

[0010] In one embodiment, the enhanced attention mechanism specifically includes: Obtain a three-dimensional input feature map; Perform global average pooling and global max pooling on the input feature map to obtain a compressed vector after pooling the input feature map; Input the compressed vector into a shared multi-layer perceptron and add the outputs of the perceptron element by element; Apply an activation function to the result of adding element by element to finally generate channel attention weights, and the channel attention weights are used to represent the importance of each channel; Multiply the channel attention weights element-wise by the feature map adjusted by the channel attention to obtain the feature map processed by the attention mechanism.

[0011] In one embodiment, adding the residual structure further includes: Set the convolutional layer and the BN normalization layer as the main path, and the main path is used for feature extraction and normalization of the UAV audio dataset; Add the output of the attention mechanism to the output of the BN normalization layer in the main path.

[0012] In a second aspect, the present application also provides an anti-UAV detection system based on deep learning. The system includes: An audio data construction module, configured to construct a UAV audio dataset fused with random noise; An audio feature fusion module, configured to fuse the random noise in the UAV audio dataset based on Fbank features, MFCC features, and GFCC features; A neural network construction module, configured to construct a convolutional neural network based on adding an attention mechanism and a residual structure and detect UAV audio.

[0013] In a third aspect, the present application also provides a computer system. The computer system includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: Construct a UAV audio dataset fused with random noise; Fuse the random noise in the UAV audio dataset based on Fbank features, MFCC features, and GFCC features; Construct a convolutional neural network based on adding an attention mechanism and a residual structure and detect UAV audio.

[0014] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: Construct a UAV audio dataset fused with random noise; Fuse the random noise in the UAV audio dataset based on Fbank features, MFCC features, and GFCC features; Construct a convolutional neural network based on adding an attention mechanism and a residual structure and detect UAV audio.

[0015] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented: Construct a UAV audio dataset fused with random noise; Fuse the random noise in the UAV audio dataset based on Fbank features, MFCC features, and GFCC features; Construct a convolutional neural network based on an enhanced attention mechanism and residual structure to detect UAV audio.

[0016] 3. Beneficial effects This application adopts the above method, and the present invention adopts a data augmentation method of randomly adding noise. By adding various noises simulating real environments to the training data, the working conditions of the algorithm in actual complex environments are simulated. This method enables the model to learn to cope with diverse noise interferences during the training stage, overcoming the defect of the existing technology that the detection performance decreases in complex noise environments due to training based on a pure dataset; The present invention adopts a method of fusing multiple voiceprint features based on Fbank features + MFCC features + GFCC features to comprehensively describe UAV audio features from multiple dimensions, making up for the deficiency of the existing technology that only uses a single feature, resulting in incomplete feature information, thereby more accurately identifying UAV audio signals and improving the accuracy of detection; The present invention proposes an audio detection method based on an enhanced attention mechanism and residual structure. The attention mechanism enables the model to automatically focus on the key audio feature parts when processing audio data and ignore irrelevant or interfering information. In complex noise environments, the attention mechanism helps the model more accurately capture the effective features of UAV audio and enhance the stability of the model under different noise intensities and types of interference. Description of the drawings

[0017] Figure 1 Is the basic neural network framework of the convolutional neural network in one embodiment; Figure 2 Is the structural diagram of the CBAM module in one embodiment; Figure 3 Is the neural network structural diagram with an enhanced attention mechanism and residual structure in one embodiment; Figure 4 Is the structural block diagram of the anti-UAV detection system based on deep learning in one embodiment; Figure 5 Is the internal structural diagram of the computer system in one embodiment. Detailed implementation manners

[0018] In order to make the objectives, technical solutions, and advantages of this application clearer, the following further describes this application in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0019] In one embodiment, the method includes the following steps: Step 202: Construct a UAV audio dataset incorporating random noise.

[0020] Among them, this application cites the open-source noise dataset ESC-50 as the background noise source, and selects a total of 20 noise sounds including 10 natural scene sounds (such as wind sound, rain sound, etc.) and 10 outdoor sounds (such as car sound, airplane sound, crowd sound, etc.) in ESC-50. Randomly select 5 types from the 20 noises, with a total of 200 noise data, which are used for UAV audio data augmentation during training; the remaining 15 noises, with a total of 600 data, are used to add noise to the test data to simulate the detection of UAV audio under unknown noise interference.

[0021] Step 204: Fuse the random noise in the UAV audio dataset based on Fbank features, MFCC features, and GFCC features.

[0022] Among them, considering the Fbank feature: it has a strong ability to capture the energy distribution of the entire audio frequency band and can well retain the spectral detail features of the original audio; the MFCC feature: is good at simulating the human ear's auditory characteristics, is more sensitive to low-frequency sounds, and has a relatively low sensitivity to high-frequency sounds; the GFCC feature: uses a different filter from the MFCC feature and has a better processing effect on high-frequency signals.

[0023] It is worth mentioning that this application adopts a weighted fusion-based method, and the formula is as follows: ; ; Among them, is the weight coefficient of the fusion, is the Fbank feature, is the MFCC feature, is the GFCC feature, is the fused feature.

[0024] Step 206: Construct a convolutional neural network based on adding an attention mechanism and a residual structure and detect UAV audio.

[0025] In the above anti-UAV detection method based on deep learning, the present invention adopts a data augmentation method of random noise addition. By adding various noises simulating the real environment to the training data, it simulates the working conditions of the algorithm in the actual complex environment. This method enables the model to learn to cope with diverse noise interferences during the training stage, overcoming the defect that the detection performance of the existing technology decreases in a complex noise environment due to training based on a pure dataset; The present invention adopts a method for fusing multiple voiceprint features based on Fbank features + MFCC features + GFCC features, comprehensively describing the drone audio features from multiple dimensions, making up for the deficiency of the existing technology that only uses a single feature, resulting in incomplete feature information, so as to more accurately identify the drone audio signal and improve the accuracy of detection; The present invention proposes an audio detection method based on adding an attention mechanism and a residual structure. The attention mechanism enables the model to automatically focus on the key audio feature parts and ignore irrelevant or interfering information when processing audio data. In a complex noise environment, the attention mechanism helps the model to more accurately capture the effective features of the drone audio and enhance the stability of the model under the interference of different noise intensities and types.

[0026] In one embodiment, constructing a drone audio dataset fused with random noise includes: Determining the types and sources of noise; dividing the noise into training data and test data; preprocessing the drone audio data; enhancing the training data of the noise; adding noise to the training data within a preset signal-to-noise ratio range and constructing a drone audio dataset fused with random noise.

[0027] Among them, considering that the sound information of the drone will be attenuated to a certain extent when propagating in the air, and the high-frequency part will be attenuated more than the low-frequency part, so this patent preprocesses the original drone audio signal x(m) as shown in the following formula: ; Among them, =0, 1, 2,..., n, where n represents the signal length, represents the amplitude of the sampling point signal, represents the weight, which is set to 0.95 in this application.

[0028] It is worth mentioning that the enhancement processing of the training data of the noise includes: Randomly selecting noise samples from the training data of the noise; sampling the noise samples and keeping the sampling length of the noise samples consistent with the sampling length of the drone audio; randomly sampling the signal-to-noise ratio within a preset signal-to-noise ratio range; adding noise to the drone audio data based on the average power of the drone audio signal and the average power of the sampled noise.

[0029] Among them, randomly select one from the noise training data as a noise sample; sample the noise sample, requiring the sampling length to be consistent with the drone audio length, to obtain x noise ; randomly sample the signal-to-noise ratio snr between [-5, 10]; add noise to the drone audio data, and the processing process is as shown in the formula: ; ; ; Among them, represents the average power of the original UAV audio signal, represents the average power of the noise x noise . represents the value of the original UAV audio signal at discrete time m, where m = 0, 1, 2,..., n, and n represents the signal length; represents the signal after adding noise, snr represents the signal-to-noise ratio, that is, the ratio of the signal power to the noise power.

[0030] In this embodiment, two kinds of noises are randomly selected from the test noise dataset for mixing, and the noise addition operation is performed between [-5, 10], so as to simulate the working sound of the UAV in noise environments with different complexities, and form a test set for evaluating the performance of the model in unknown noise environments.

[0031] It is worth mentioning that the framing and windowing of the UAV audio signal (after adding noise). Considering that the audio has a long time and cannot be directly fed into the neural network for training, so this patent frames the audio, the frame length is 30 ms, and the frame shift is 10 ms. Since framing may cause discontinuity (mutation) between the head and tail of each frame of audio and the front and back signals, so this patent samples the Hamming window for windowing processing to make the audio signal smoother. The Hamming window formula is as follows: ; Among them, is the sampling window length; is the window function.

[0032] In one embodiment, as Figure 1 shown, a basic neural network framework based on a convolutional neural network is built, including: Define two repeated convolutional blocks. The convolutional block includes a convolutional layer, a BN normalization layer, and a ReLU activation layer; the convolutional layer is used to receive the UAV audio and extract the features of the UAV audio dataset; the BN normalization layer performs batch normalization on the features of the UAV audio dataset output by the convolutional layer; the ReLU activation layer applies a non-linear activation function to the normalized data.

[0033] Among them, the basic neural network framework based on a convolutional neural network is composed of two 3*3 convolutional layers, two BN normalization layers, and two ReLU activation layers alternately in sequence, and performs feature extraction and transformation on the input data to output the processed result.

[0034] It is worth mentioning that building a basic framework based on a convolutional neural network mainly includes two repeated modules, and each module is successively composed of a convolutional layer, a BN normalization layer, and a ReLU activation layer. Building steps: Define the first convolutional block: Convolutional layer (Conv2d): Receives the input data and extracts features from it through a convolutional kernel. For example, a 3×3 convolutional kernel slides over the input data to calculate the weighted sum of the local area, thereby detecting low-level features such as edges and textures; BN normalization layer (BatchNorm): Performs batch normalization on the output of the convolutional layer. This helps to stabilize network training, accelerate convergence, and allows the use of a higher learning rate. It reduces internal covariate shift by adjusting the mean and variance of each batch of inputs; ReLU activation layer (ReLU): Applies a non-linear activation function to the normalized data. The ReLU (Rectified Linear Unit) function sets all negative values to zero and keeps positive values unchanged. This introduces non-linearity, enabling the network to learn more complex patterns. Otherwise, multiple linear transformations are equivalent to a single linear transformation.

[0035] Define the second convolutional block: Repeat the above steps to build a second sequence consisting of a convolutional layer, a BN normalization layer, and a ReLU activation layer. This block further extracts higher-level features on the basis of the features extracted by the previous block, or performs deeper transformations on the features.

[0036] Function of the convolutional layer: Extract local features of the input data. By sliding a learnable convolutional kernel over the input data, it performs a convolution operation to capture the spatial hierarchy and patterns. For example, edges, corners, etc. can be recognized in an image; Function of the BN normalization layer (BatchNorm): Stabilize and accelerate the training process of the neural network. By normalizing the mean and variance of each batch of inputs, it reduces internal covariate shift, making the network less sensitive to the choice of parameter initialization and learning rate, and having a certain regularization effect to prevent overfitting; Function of the ReLU activation layer (ReLU): Introduce non-linearity. It allows the network to learn and represent more complex non-linear relationships. Without a non-linear activation function, no matter how many layers of linear transformations are stacked, the effect is equivalent to a single linear transformation, thus limiting the expressive power of the network.

[0037] In one embodiment, as Figure 2 shown, adding an attention mechanism specifically includes: Obtain a three-dimensional input feature map; perform global average pooling and global max pooling on the input feature map to obtain a compressed vector after pooling of the input feature map; input the compressed vector into a shared multi-layer perceptron and add the outputs of the perceptron element-wise; apply an activation function to the result of the element-wise addition to finally generate channel attention weights, where the channel attention weights are used to represent the importance of each channel; multiply the channel attention weights element-wise by the feature map adjusted by channel attention to obtain a feature map processed by the attention mechanism.

[0038] Among them, the purpose of CBAM (Convolutional Block Attention Module) is to make the network pay more attention to key features, which is achieved by successively deploying a channel attention module and a spatial attention module.

[0039] The dimension of the input feature map is (C*H*W). It is respectively averaged along the width direction (WAvgPooling) to obtain a feature map of dimension (C*1*W), and averaged along the height direction (HAvgPooling) to obtain a feature map of dimension (C*H*1); after the two are concatenated, they are convolved in two dimensions (Concat&Conv2d), and then batch-normalized (BatchNorm) and linearly transformed (Linear); then they are respectively convolved in two dimensions (Conv2d) and passed through the Sigmoid activation function to obtain two weight feature maps; finally, the weight feature maps are multiplied by the original input feature map, and the output feature map still has a dimension of (C*H*W).

[0040] It is worth mentioning that for the preparation of the input feature map: receive an input feature map with a dimension of (C×H×W). For the construction of the channel attention module (Channel Attention Module, CAM): perform global average pooling (GlobalAvg Pooling) and global max pooling (Global Max Pooling) on the input feature map. These two pooling operations are respectively performed on the spatial dimension (H×W), compressing the feature map into a vector of (C×1×1), thereby aggregating spatial information. Send these two (C×1×1) vectors into a shared multi-layer perceptron (Shared MLP). This MLP usually includes a dimensionality reduction layer, a ReLU activation layer, and a dimensionality increase layer. Add the two outputs of the MLP element-wise. Apply the Sigmoid activation function to the result of the addition to generate channel attention weights of dimension (C×1×1). These weights represent the importance of each channel. Multiply these channel attention weights element-wise by the original input feature map (C×H×W) to obtain a feature map adjusted by channel attention.

[0041] Construction of Spatial Attention Module (SAM): For the feature map adjusted by channel attention, perform average pooling along the channel dimension (Avg Pooling along channel) and max pooling along the channel dimension (Max Pooling along channel). These two pooling operations compress the feature map into a single-channel feature map of (1×H×W) in the channel dimension. Concatenate these two (1×H×W) feature maps in the channel dimension to form a (2×H×W) feature map. Apply a two-dimensional convolutional layer (Conv2d) to the concatenated feature map (for example, using a 7×7 convolutional kernel) to reduce the number of channels from 2 to 1, obtaining a (1×H×W) feature map. Apply the Sigmoid activation function to the output of the convolutional layer to generate spatial attention weights of dimension (1×H×W). These weights represent the importance of each spatial position. Multiply these spatial attention weights element-wise by the feature map (C×H×W) adjusted by channel attention to obtain the final feature map processed by the CBAM attention mechanism, with the dimension still being (C×H×W).

[0042] In this embodiment, the CBAM module focuses on key features by sequentially deploying the channel attention module and the spatial attention module. It first performs global average pooling and global max pooling on the input feature map respectively, sends the results to a shared multi-layer perceptron for processing, generates channel attention weights through an activation function to screen important channels; then performs global average pooling and global max pooling on the feature map output by the channel attention module, concatenates the two in the channel dimension, and generates spatial attention weights through convolution and an activation function to further highlight key regions. Finally, through two weightings, it strengthens the network's attention to key features; suppresses unimportant features, improves the feature utilization efficiency, enhances the network's ability to extract effective information, and thus can effectively detect noisy drone audio.

[0043] In one embodiment, as Figure 3 shown, the residual structure is realized through skip connections, aiming to solve the problem of gradient disappearance in the training of deep neural networks and improve the model performance and stability; adding the residual structure also includes: Set the convolutional layer and the BN normalization layer as the main path, and the main path is used to extract features and normalize the drone audio dataset; add the output of the attention mechanism to the output of the BN normalization layer in the main path.

[0044] Among them, in the improved neural network structure, the residual structure is realized through skip connections. After the input data passes through the first convolutional layer and the BN normalization layer, it is directly added and fused with the output of the CBAM module; similarly, the features passing through the second Conv convolutional layer and the BN normalization layer are also added to the features after passing through the corresponding CBAM module through skip connections.

[0045] It is worth mentioning that the first residual connection: Main Path: The input data first undergoes feature extraction and normalization through the first convolutional layer (Conv Layer) and the first BN normalization layer (BN Layer). Skip Connection: At the same time, the output of the CBAM module (the input of this CBAM module is the original feature before the first convolutional layer, or the feature after the first convolutional layer but before the BN layer) is directly added (Add) to the output of the first BN normalization layer in the main path. Fusion: This element-wise addition operation enables the network to learn the "residual" between the input and the output after multiple transformations, rather than directly learning the complete mapping.

[0046] The second residual connection: Main Path: The features fused through the first residual connection continue to be processed through the second convolutional layer (Conv Layer) and the second BN normalization layer (BN Layer). Skip Connection: Similar to the first residual connection, the output of the CBAM module corresponding to the second convolutional block (the input of this CBAM module is the feature before the second convolutional layer) is directly added (Add) to the output of the second BN normalization layer in the main path. Final Output: Such fused features are the final processing results of this improved network structure.

[0047] Alleviating the vanishing gradient problem: Skip connections provide a "shortcut" for gradients, allowing gradients to directly backpropagate to the early layers of the network, thus effectively alleviating the vanishing gradient problem during the training of deep networks and enabling deeper networks to be effectively trained.

[0048] Improving the training stability and performance of the model: By learning residuals, the network is easier to optimize, the convergence speed is accelerated, and better performance can be achieved.

[0049] Helping the network better extract and utilize features: Residual connections enable the network to more easily learn the identity mapping, which means that if a certain convolutional layer fails to learn useful information, it can simply pass the original input, avoiding the "degradation" problem, and thus more effectively extracting and utilizing features.

[0050] In this embodiment, this method enables the network to learn the residuals between the input and the output after multiple-layer transformation, effectively alleviating the problem of gradient disappearance during the training of deep networks, improving the training stability and performance of the model, and helping the network to better extract and utilize features.

[0051] It is worth mentioning that the UAV audio detection neural network with an attention mechanism and a residual structure has the following advantages. (1) Enhance the feature extraction ability, which can more accurately capture the subtle and key feature information in UAV audio, and deeply explore its unique voiceprint characteristics. (2) Improve the anti-interference performance, effectively cope with complex environmental noise interference, greatly reduce the impact of noise on the detection results, and significantly reduce the false detection rate. (3) Strengthen the robustness of the model. In diverse unknown noise environments and different scenarios, it can adaptively adjust the focus of attention and stably and reliably achieve accurate detection of UAV audio.

[0052] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0053] Based on the same inventive concept, the embodiments of the present application also provide a deep learning-based anti-UAV detection system for implementing the above-mentioned deep learning-based anti-UAV detection method. The implementation solutions provided by this system to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the following deep learning-based anti-UAV detection system can refer to the limitations on the deep learning-based anti-UAV detection method in the above text, and will not be repeated here.

[0054] In one embodiment, as Figure 4 shown, a deep learning-based anti-UAV detection system is provided, including: an audio data construction module, an audio feature fusion module, and a neural network construction module, where: The audio data construction module is used to construct a UAV audio data set fused with random noise; The audio feature fusion module is used to fuse the random noise in the UAV audio data set based on Fbank features, MFCC features, and GFCC features; A neural network construction module for constructing a convolutional neural network based on an enhanced attention mechanism and a residual structure and detecting drone audio.

[0055] In one embodiment, the audio data construction module is further configured to: determine the type and source of noise; divide the noise into training data and test data; preprocess the drone audio data; enhance the training data of the noise; add noise to the training data within a preset signal-to-noise ratio range and construct a drone audio dataset integrated with random noise.

[0056] In one embodiment, the audio data construction module is further configured to: randomly select noise samples from the training data of the noise; sample the noise samples and keep the sampling length of the noise samples consistent with the sampling length of the drone audio; perform random sampling of the signal-to-noise ratio within a preset signal-to-noise ratio range; add noise to the drone audio data based on the average power of the drone audio signal and the average power of the sampled noise.

[0057] In one embodiment, the neural network construction module is further configured to: define two repeated convolutional blocks, where each convolutional block includes a convolutional layer, a BN normalization layer, and a ReLU activation layer; the convolutional layer is used to receive the drone audio and extract the features of the drone audio dataset; the BN normalization layer performs batch normalization on the features of the drone audio dataset output by the convolutional layer; the ReLU activation layer applies a non-linear activation function to the normalized data.

[0058] In one embodiment, the neural network construction module is further configured to: obtain a three-dimensional input feature map; perform global average pooling and global max pooling on the input feature map to obtain a compressed vector after pooling the input feature map; input the compressed vector into a shared multi-layer perceptron and add the outputs of the perceptron element-wise; apply an activation function to the result of the element-wise addition to finally generate channel attention weights, where the channel attention weights are used to represent the importance of each channel; multiply the channel attention weights element-wise by the feature map adjusted by the channel attention to obtain a feature map processed by the attention mechanism.

[0059] In one embodiment, the neural network construction module is further configured to: set the convolutional layer and the BN normalization layer as the main path, where the main path is used to extract features and perform normalization on the drone audio dataset; add the output of the attention mechanism to the output of the BN normalization layer in the main path.

[0060] Each module in the above anti-drone detection system based on deep learning can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer system in hardware form or be independent of it, or be stored in the memory in the computer system in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.

[0061] In one embodiment, a computer system is provided. The computer system may be a server, and its internal structure diagram may be as shown in Figure 5 . The computer system includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer system is used to provide computing and control capabilities. The memory of the computer system includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer system is used to store data. The network interface of the computer system is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an anti-drone detection method based on deep learning.

[0062] In one embodiment, a computer system is provided. The computer system may be a terminal, and its internal structure diagram may be as shown in Figure 5 . The computer system includes a processor, a memory, a communication interface, a display screen, and an input system connected through a system bus. Among them, the processor of the computer system is used to provide computing and control capabilities. The memory of the computer system includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer system is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an anti-drone detection method based on deep learning.

[0063] Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer system to which the solution of the present application is applied. The specific computer system may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0064] In one embodiment, a computer system is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.

[0065] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in the above method embodiments.

[0066] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps in the above method embodiments.

[0067] It should be noted that the user information (including but not limited to user system information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties.

[0068] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the method embodiments as described above. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0069] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0070] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. An anti-drone detection method based on deep learning, characterized in that, The method includes: Construct a UAV audio dataset integrated with random noise; Fuse the random noise in the UAV audio dataset based on Fbank features, MFCC features, and GFCC features; Construct a convolutional neural network based on an added attention mechanism and residual structure to detect UAV audio.

2. The anti-drone detection method based on deep learning according to claim 1, wherein The construction of the UAV audio dataset integrated with random noise includes: Determine the types and sources of noise; Divide the noise into training data and test data; Preprocess the UAV audio data; Enhance the training data of the noise; Add noise to the training data within a preset signal-to-noise ratio range and construct a UAV audio dataset integrated with random noise.

3. The anti-drone detection method based on deep learning according to claim 2, characterized in that, The enhancement of the training data of the noise includes: Randomly select noise samples from the training data of the noise; Sample the noise samples and keep the sampling length of the noise samples consistent with the sampling length of the UAV audio; Perform random sampling of the signal-to-noise ratio within a preset signal-to-noise ratio range; Add noise to the UAV audio data based on the average power of the UAV audio signal and the average power of the sampled noise.

4. The anti-drone detection method based on deep learning according to claim 1, characterized in that, The construction of the convolutional neural network based on an added attention mechanism and residual structure includes: Define two repeated convolutional blocks, and each convolutional block includes a convolutional layer, a BN normalization layer, and a ReLU activation layer; The convolutional layer is used to receive UAV audio and extract the features of the UAV audio dataset; The BN normalization layer performs batch normalization on the features of the UAV audio dataset output by the convolutional layer; The ReLU activation layer applies a non-linear activation function to the normalized data.

5. The anti-drone detection method based on deep learning according to claim 4, wherein The added attention mechanism specifically includes: Obtain a three-dimensional input feature map; Perform global average pooling and global max pooling on the input feature map to obtain a compressed vector after pooling the input feature map; Input the compressed vector into a shared multi-layer perceptron and add the outputs of the perceptron element-wise; Apply an activation function to the result of the element-wise addition to finally generate channel attention weights, and the channel attention weights are used to represent the importance of each channel; Multiply the channel attention weights element-wise by the feature map adjusted by channel attention to obtain a feature map processed by the attention mechanism.

6. The anti-drone detection method based on deep learning according to claim 5, wherein, The addition of the residual structure further includes: Set the convolutional layer and the BN normalization layer as the main path, and the main path is used to extract features and perform normalization on the UAV audio dataset; Add the output of the attention mechanism to the output of the BN normalization layer in the main path.

7. An anti-drone detection system based on deep learning, characterized in that, The system includes: An audio data construction module for constructing a UAV audio dataset integrated with random noise; An audio feature fusion module for fusing the random noise in the UAV audio dataset based on Fbank features, MFCC features, and GFCC features; A neural network construction module for constructing a convolutional neural network based on an added attention mechanism and residual structure to detect UAV audio.

8. A computer system, including a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Unmanned aerial vehicle detection method based on deep learning

    CN109753903A

  • ConvLSTM and CBAM attention mechanism fused non-contact heart rate measurement method

    CN115024706A

  • Low-slow small unmanned aerial vehicle detection method based on acoustic fusion features

    CN119626256A

  • Unmanned aerial vehicle sound recognition method and device

    CN119811428A

  • Unmanned aerial vehicle acoustic identification method and device based on frequency band feature extraction

    CN119889350A