A method and system for identifying signals of unmanned aerial vehicles

By employing a signal-to-noise ratio estimation algorithm combining time-frequency domain joint filtering and multi-scale frequency domain attention network, the accuracy and real-time performance issues of UAV signal detection and recognition in complex electromagnetic environments are addressed, achieving high-precision UAV model classification.

CN120849923BActive Publication Date: 2025-12-16SHANDONG UNIV

Patent Information

Application Number
CN202511358724.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-12-16
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

Existing UAV signal detection and identification technologies suffer from problems such as inaccurate signal-to-noise ratio estimation, insufficient anti-interference capability, and high computational complexity in complex electromagnetic environments, making it difficult to achieve real-time and accurate UAV model classification.

Method used

A signal-to-noise ratio estimation algorithm based on joint time-frequency domain filtering is adopted, which combines adaptive noise suppression, feature enhancement and deep learning regression. Signal recognition is performed through a multi-scale frequency domain attention network, including adaptive noise basis estimation, time-frequency feature enhancement, multi-feature fusion regression and UAV signal classification model.

Benefits of technology

Achieving high-precision signal-to-noise ratio estimation across the entire dynamic range improves the robustness and classification accuracy of UAV signal detection, meets the real-time processing needs of edge devices, and enables accurate identification of UAV models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849923B_ABST
    Figure CN120849923B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of unmanned aerial vehicle signal identification method and system, belong to unmanned aerial vehicle signal identification technical field.It includes: step 1: adaptive noise base estimation;The input unmanned aerial vehicle IQ signal, i.e.unmanned aerial vehicle signal is carried out short-time Fourier transform, and time-frequency matrix is obtained, generates unmanned aerial vehicle time-frequency diagram, estimates noise power spectrum, calculates the power spectral density of each frequency point;Step 2: time-frequency feature enhancement;5 order Butterworth band-pass filter is used to frequency domain filtering to signal;Soft threshold domain denoising of soft short-time Fourier transform result is used;Step 3: multi-feature fusion regression;Energy ratio feature, spectral peak sharpness feature are calculated, and deep learning feature is extracted, feature fusion and regression;Step 4: through the unmanned aerial vehicle signal classification model after training, unmanned aerial vehicle signal identification is realized.The present application method is high in precision in complex environment signal-to-noise ratio estimation, and unmanned aerial vehicle classification precision and generalization are excellent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and system for unmanned aerial vehicle (UAV) signal recognition, belonging to the field of UAV signal recognition technology. Background Technology

[0002] With the rapid popularization of consumer and industrial drones, their applications in aerial surveying, logistics, and other fields are becoming increasingly widespread. However, this has also led to security risks such as unauthorized flights intruding into airport airspace and eavesdropping on privacy. Accurate detection and identification of drone signals has become a key technological support for ensuring public safety. Currently, drone communication largely relies on open frequency bands such as 2.4GHz. This band has a complex electromagnetic environment, with issues such as WiFi signal interference and sudden impulse noise, causing severe performance degradation of traditional signal processing methods. For example, in scenarios with a signal-to-noise ratio (SNR) below -10dB, traditional energy detection-based algorithms can have a false negative rate of over 30%. Furthermore, the application of frequency-hopping communication technology renders identification methods relying on fixed frequency characteristics ineffective. Existing SNR estimation algorithms have significant limitations in complex environments: the energy ratio method has an error exceeding 5dB under non-Gaussian noise, and the cyclostationary method has poor adaptability to frequency-hopping signals and is computationally complex, making it difficult to meet the real-time processing needs of edge devices. Meanwhile, the classification of drone models relies too heavily on single-dimensional features, which are easily masked under strong interference, causing the classification accuracy of mainstream algorithms to plummet to below 60%. There is an urgent need to build a signal processing framework that can adapt to the full dynamic signal-to-noise ratio range and resist complex interference.

[0003] In UAV signal detection and recognition systems, accurate signal-to-noise ratio (SNR) estimation is a crucial prerequisite for achieving adaptive threshold adjustment and improving detection robustness. Traditional estimation algorithms have significant shortcomings in complex electromagnetic environments: the energy ratio method is susceptible to non-Gaussian noise interference, with estimation errors exceeding 5 dB in impulse noise scenarios; while methods based on cyclostationary characteristics offer strong noise resistance, they are poorly adaptable to frequency-hopping UAV signals, have high computational complexity, and struggle to meet the real-time processing requirements of edge devices. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention proposes a signal-to-noise ratio estimation algorithm based on joint time-frequency domain filtering. By integrating adaptive noise suppression, feature enhancement, and deep learning regression, it achieves high-precision estimation across the entire dynamic range (-20dB to 18dB).

[0005] This invention addresses the technical bottlenecks in UAV signal detection and identification by proposing a signal-to-noise ratio (SNR) estimation algorithm based on joint time-frequency domain filtering and a multi-scale frequency-domain attention classification network, possessing significant theoretical and practical value. Theoretically, by integrating adaptive noise suppression, feature enhancement, and deep learning regression, it overcomes the estimation accuracy limitations of traditional algorithms across the entire dynamic range (-20dB to 18dB), providing a new method for signal processing in low SNR scenarios. The introduction of the multi-scale frequency-domain attention mechanism solves the problem of insufficient robustness of single features under strong interference, enriching the theoretical system of fine-grained wireless signal identification. In terms of application, the proposed algorithm can be integrated into security systems in key areas such as airports and stadiums, enabling real-time UAV detection and model identification with a response time controlled within 1 second, meeting the deployment requirements of edge computing devices. The self-constructed DroneRF2025 dataset covers eight mainstream UAV models, providing a standardized testing benchmark for related research and promoting the industrialization of UAV surveillance technology.

[0006] The technical solution of this invention is as follows:

[0007] A method for identifying drone signals includes:

[0008] Step 1: Adaptive noise basis estimation; Perform short-time Fourier transform on the input UAV IQ signal (i.e., the UAV signal) to obtain the time-frequency matrix, generate the UAV time-frequency map, estimate the noise power spectrum, and calculate the power spectral density at each frequency point;

[0009] Step 2: Time-frequency feature enhancement; a 5th-order Butterworth bandpass filter is used to perform frequency domain filtering on the signal; soft thresholding is used to denoise the soft short-time Fourier transform results;

[0010] Step 3: Multi-feature fusion and regression; calculate energy ratio features and spectral peak sharpness features, extract deep learning features, and perform feature fusion and regression;

[0011] Step 4: Achieve drone signal recognition through the trained drone signal classification model.

[0012] According to a preferred embodiment of the present invention, adaptive noise basis estimation includes:

[0013] Step 1.1: Obtain the time-frequency matrix S(t,f);

[0014] After performing a short-time Fourier transform on the input IQ signal, a time-frequency matrix S(t,f) is obtained, where dimension t represents a time frame and dimension f represents a frequency point. Each element S(t,f) is a complex number, representing the time-frequency amplitude and phase information of the signal at the t-th frame and the f-th frequency point. In the STFT, the phase information is the time position mark of a specific frequency component in the corresponding time frame.

[0015] Step 1.2: Calculate the energy value at each time frequency point;

[0016] The energy value at each time frequency point is obtained by taking the square of the modulus of each element in the time-frequency matrix S(t,f).

[0017] Step 1.3: Calculate the mean over the time dimension to obtain the power spectral density;

[0018] For each fixed frequency point f, the power spectral density at that frequency point is obtained by averaging the energy values ​​of all time frames at that frequency. ;

[0019] Arrange the power spectral densities at different frequency points from low to high to form an ordered sequence, namely the power spectral density sequence.

[0020] 1.4: Sliding window minimum tracking of the power spectral density sequence:

[0021] Build a sliding window;

[0022] The minimum power spectral density within each sliding window is selected as the initial noise estimate at the center frequency of that sliding window.

[0023] By sliding the window sequentially and traversing all frequency points, a complete initial noise estimation sequence can be obtained. ;

[0024] Exponential smoothing filtering is performed using the exponential smoothing formula, as shown below:

[0025] ;

[0026] Where α = 0.05 is the smoothing coefficient. This indicates the smoothing result of the previous frequency point.

[0027] According to a preferred embodiment of the present invention, a 5th-order Butterworth bandpass filter is used to perform frequency domain filtering on the signal; including:

[0028] Step 2.1: Carrier frequency tracking: Use a carrier frequency tracking algorithm to detect the carrier frequency position of the signal in real time. ;

[0029] Step 2.2: Dynamic Filter Settings: Dynamically adjust the passband range of the filter according to the carrier frequency position. ;

[0030] Step 2.3: Filtering: Perform filtering operation on the time-frequency matrix S(t, f);

[0031] Step 2.4: Noise Suppression: Set the stopband attenuation of the filter to no less than 40dB.

[0032] According to a preferred embodiment of the present invention, soft thresholding denoising is performed using the soft short-time Fourier transform result; comprising:

[0033] Step 2.5: Adaptive Threshold Calculation; Calculate the adaptive threshold for each frequency point according to the following formula:

[0034] ;

[0035] in, The resulting noise substrate;

[0036] Step 2.6: Soft threshold filtering: When the amplitude at a given frequency point At that time, retain the signal components and subtract the threshold from them, that is: , This refers to an adaptive threshold; when When the frequency point is considered to be noise, it is set to zero.

[0037] According to a preferred embodiment of the present invention, calculating the energy ratio characteristic includes:

[0038] Step 3.1: Calculate the total energy of the denoised signal. As shown below:

[0039] ;

[0040] in, , where is the number of time frames; , where is the number of frequency points; This refers to the denoised signal;

[0041] Step 3.2: Calculate the total noise energy As shown below:

[0042] ;

[0043] in, As a noise floor, For time frames;

[0044] Step 3.3: Calculate the energy ratio characteristics As shown below:

[0045] .

[0046] According to a preferred embodiment of the present invention, calculating the peak sharpness characteristics includes:

[0047] Step 3.4: Extract the power spectrum of the denoised signal As shown below:

[0048] ;

[0049] Step 3.5: Find the main lobe peak of the power spectrum As shown below:

[0050] ;

[0051] Step 3.6: Calculate the 3dB bandwidth :

[0052] Find the power spectrum greater than or equal to All frequency points; the difference between the maximum and minimum values ​​among these frequency points is the 3dB bandwidth;

[0053] Step 3.7: Calculate the peak sharpness characteristics As shown below:

[0054] .

[0055] According to a preferred embodiment of the present invention, deep learning features are extracted, including:

[0056] A lightweight convolutional neural network architecture is constructed, which includes an input layer, a first convolutional layer, a second convolutional layer, and a global average pooling layer.

[0057] In the input layer, the input is a denoised time-frequency matrix S(t,f) of 256×256;

[0058] The first convolutional layer consists of 16 3×3 convolutional kernels with a stride of 1 and ReLU activation;

[0059] The second convolutional layer consists of 32 3×3 convolutional kernels with a stride of 1 and ReLU activation;

[0060] Global average pooling layer: converts the feature map into a 32-dimensional feature vector F.

[0061] According to a preferred embodiment of the present invention, feature fusion and regression includes:

[0062] Step 3.8: Feature Normalization:

[0063] For three types of features And deep learning features, namely 32-dimensional feature vectors F;

[0064] Normalize them separately to get F1, F2, and F3;

[0065] Step 3.9: Integration of attention mechanisms:

[0066] Attention mechanism fusion through attention networks;

[0067] Attention networks consist of a first-layer MLP and a second-layer MLP;

[0068] The input consists of the F1, F2, and F3 values ​​of the three features to be fused.

[0069] First, perform global average pooling on each feature to convert the high-dimensional features into low-dimensional vectors, which are then used as input to the attention network.

[0070] After passing through the first layer of MLP, which consists of 64 neurons and uses ReLU as the activation function, non-linearity is introduced to enhance the feature expression capability;

[0071] After passing through a second MLP layer, which consists of 3 neurons and uses the Softmax activation function, the sum of the 3 output weights is guaranteed to be 1. ;

[0072] Output 3 weights , respectively corresponding The importance of the value is [0,1]).

[0073] Step 3.10: Finally, calculate the signal-to-noise ratio. The estimate is shown in the following formula:

[0074] .

[0075] According to a preferred embodiment of the present invention, drone signal recognition is achieved through a trained drone signal classification model; including:

[0076] The UAV signal classification model is a multi-scale frequency domain attention network, including an input layer, a spectral enhancement module, a deep feature distillation module, a classification head, and an output layer.

[0077] The input layer receives the time-frequency spectrum generated by the short-time Fourier transform, which is the UAV time-frequency spectrum;

[0078] The spectral enhancement module, as the front-end processing unit of the UAV signal classification model, achieves multi-scale feature extraction through three parallel branches: the first branch is a 1×1 convolution branch; the second branch is a 3×3 depthwise separable convolution branch; and the third branch is a 5×5 dilated convolution branch. The features output by the three branches are fused through adaptive weights to form an initial enhanced feature map.

[0079] The deep feature distillation module consists of progressive processing stages, each stage including several FABlock units;

[0080] The FABlock unit includes a frequency band separation mechanism, dynamic attention calibration, and residual fusion path;

[0081] The frequency band separation mechanism refers to the process of decomposing the initial enhanced feature map into low-frequency approximate components and high-frequency detail components through two-dimensional wavelet transform.

[0082] Dynamic attention calibration refers to generating band-level attention masks by applying spatial attention modules and channel attention modules to low-frequency approximate components and high-frequency detail components, respectively.

[0083] The residual fusion path refers to: fusing the initial enhanced feature map through a 1×1 convolution, adding it to the main path's features (i.e., the frequency band-level attention mask), and then passing it through the BatchNorm and SiLU activation functions to achieve feature enhancement and dimensionality unification;

[0084] The classification head includes a global average pooling layer and a multilayer perceptron, where the hidden layers of the multilayer perceptron introduce Dropout and L2 regularization;

[0085] The output layer outputs the probability distribution of 8 common drone models through the Softmax activation function. The 8 common drone models include DJI Mavic2, Mavic3, Tello, Air3, Phantom3, and three non-DJI drones, defined as UAV1, UAV2, UAV3, UAV4, UAV5, UAV6, UAV7, and UAV8 respectively.

[0086] A further preferred method is to calculate the mask value using the following formula:

[0087]

[0088] in, This is the feature map of the f-th frequency band, where f = 1, 2, 3. For the Sigmoid function, The attention weight for the f-th frequency band, It is a feature map Perform global average pooling to extract global average information; It is a feature map The channel attention module is used to adaptively learn and weight the importance of channels.

[0089] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the above-described UAV signal recognition method.

[0090] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described UAV signal recognition method.

[0091] A drone signal identification system, comprising:

[0092] The adaptive noise basis estimation module is configured to: perform a short-time Fourier transform on the input UAV IQ signal (i.e., the UAV signal) to obtain the time-frequency matrix, generate the UAV time-frequency map, estimate the noise power spectrum, and calculate the power spectral density at each frequency point;

[0093] The time-frequency feature enhancement module is configured to: perform frequency domain filtering on the signal using a 5th-order Butterworth bandpass filter; and perform soft thresholding domain denoising using the soft short-time Fourier transform result.

[0094] The multi-feature fusion regression module is configured to: calculate energy ratio features and spectral peak sharpness features, extract deep learning features, and perform feature fusion and regression.

[0095] The drone signal recognition module is configured to recognize drone signals using a trained drone signal classification model.

[0096] Compared with the prior art, the present invention has the following advantages:

[0097] 1. High SNR Estimation Accuracy in Complex Environments: Traditional SNR estimation algorithms perform poorly in complex electromagnetic environments (such as non-Gaussian noise and frequency-hopping signal scenarios). The energy ratio method is susceptible to interference, and the cyclostationary method has poor adaptability and is computationally complex. This invention proposes a SNR estimation algorithm based on joint time-frequency domain filtering, constructing a three-order architecture of "noise suppression - feature extraction - intelligent regression." Through adaptive noise basis estimation (improved minimum statistics method combined with exponential smoothing filtering to accurately capture noise characteristics), time-frequency feature enhancement (Butterworth bandpass filtering dynamically adapts to the carrier frequency, soft threshold denoising specifically suppresses noise, improving SNR by 3-5 dB under low SNR conditions), and multi-feature fusion regression (fusion of energy ratio, spectral peak sharpness, and deep learning features, weighted using an attention mechanism), high-precision estimation is achieved across the entire dynamic range (-20 dB to 18 dB), breaking through the limitations of traditional algorithms and providing a more reliable SNR basis for UAV signal detection, improving detection robustness in complex scenarios.

[0098] 2. Superior Drone Classification Accuracy and Generalization: Traditional drone classification methods rely on single-dimensional features, which are prone to failure under strong interference, resulting in a significant drop in classification accuracy. The multi-scale frequency domain attention network (MSFA-Net) of this invention uses time-frequency spectra as input (with superior noise resistance compared to the original IQ sequence). The spectrogram enhancement module extracts features at multiple scales (1×1 convolution suppresses DC interference, 3×3 depthwise separable convolution captures local jumps, and 5×5 dilated convolution extracts large-scale correlations, adaptively fusing and strengthening features). The FABlock unit of the deep feature distillation module separates stable signal features from subtle differences through a dual path of "frequency domain channel separation - spatial attention calibration," suppressing noise-dominated frequency bands and accurately uncovering model-specific spectral details (such as distinguishing similar carrier frequencies and frequency hopping differences among DJI models). The classification head incorporates regularization to alleviate overfitting. Tested on the DroneRF2025 dataset, the classification accuracy for each model exceeds 95%, with a low misclassification rate (mostly <1%). It demonstrates strong generalization for 8 types of models, including DJI and non-DJI models, solving the problem of model identification in complex scenarios and contributing to accurate supervision. Attached Figure Description

[0099] Figure 1 Time-frequency image of DJI Mavic 3 drone;

[0100] Figure 2 This is a schematic diagram of a UAV signal classification method based on a multi-scale frequency domain attention network.

[0101] Figure 3 This is the confusion matrix of the model proposed in this invention on the DroneRF2025 dataset;

[0102] Figure 4 This is a schematic diagram of the attention network structure. Detailed Implementation

[0103] To provide a clearer understanding of the technical features of the present invention, specific embodiments are described below in conjunction with the accompanying drawings. The technical solutions of the present invention are clearly and completely described. Obviously, the described embodiments are merely some embodiments of the present invention, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0104] To facilitate understanding of the present invention, a more comprehensive description of the invention will be given below with reference to the accompanying drawings, and several embodiments of the invention will be provided. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of the invention will be more thorough and complete.

[0105] Example 1

[0106] A method for UAV signal recognition is disclosed. This invention employs a three-order architecture of "noise suppression feature extraction and intelligent regression." The core idea is to remove interference components through joint time-frequency domain processing, and then improve estimation accuracy by utilizing the complementarity of multi-dimensional features. This includes:

[0107] Step 1: Adaptive noise basis estimation; Perform a short-time Fourier transform on the input UAV IQ signal (i.e., the UAV signal with length N=4096) to obtain the time-frequency matrix S(t,f) (size 256×256), generating the UAV time-frequency map, as shown below. Figure 1 As shown, the noise power spectrum is estimated using an improved minimum statistics method, and the power spectral density at each frequency point is calculated;

[0108] Step 2: Time-frequency feature enhancement; a 5th-order Butterworth bandpass filter is used to perform frequency domain filtering on the signal; soft thresholding is used to denoise the soft short-time Fourier transform results;

[0109] Step 3: Multi-feature fusion and regression; calculate energy ratio features and spectral peak sharpness features, extract deep learning features, and perform feature fusion and regression;

[0110] Step 4: Achieve drone signal recognition through the trained drone signal classification model.

[0111] Example 2

[0112] The difference between the UAV signal recognition method described in Example 1 and the method described in Example 1 is as follows:

[0113] Adaptive noise basis estimation; including:

[0114] Step 1.1: Obtain the time-frequency matrix S(t,f);

[0115] After performing a short-time Fourier transform on the input IQ signal (length 4096), a time-frequency matrix S(t,f) (size 256×256) is obtained, where dimension t represents the time frame (256 frames in total) and dimension f represents the frequency point (256 frequency points in total). Each element S(t,f) is a complex number, representing the time-frequency amplitude and phase information of the signal at the f-th frequency point in the t-th frame. In the STFT, the phase information is the time position mark of a specific frequency component in the corresponding time frame; it reflects the start time, dynamic changes, and relative temporal relationship of the frequency component with other frequency components.

[0116] Step 1.2: Calculate the energy value at each time frequency point;

[0117] The energy value at each time frequency point is obtained by taking the square of the modulus of each element in the time-frequency matrix S(t,f); as shown below:

[0118] ;

[0119] in, and These represent the real and imaginary parts of the complex number, respectively; this step converts the complex time-frequency signal into a real energy value, reflecting the strength of the signal's energy at that time frame and frequency point.

[0120] Step 1.3: Calculate the mean over the time dimension to obtain the power spectral density;

[0121] For each fixed frequency point f, the power spectral density at that frequency point is obtained by averaging the energy values ​​of all time frames at that frequency. As shown below:

[0122] ;

[0123] in, This represents the total number of time frames.

[0124] The physical meaning of the result P(f) is: the average power of the signal at frequency f, which can intuitively reflect the distribution characteristics of signal energy on the frequency axis, providing a basis for subsequent noise basis estimation and signal-to-noise ratio calculation. Through this process, the power characteristics of each frequency point can be accurately extracted, preserving the frequency domain characteristics of the signal while suppressing the interference of instantaneous noise through time averaging, providing a reliable spectral basis for UAV signal analysis in complex environments.

[0125] Specifically, within the previously mentioned algorithm framework, the power spectral density is calculated at each frequency point f. This allows us to obtain power spectral density values ​​at different frequencies. Arranging the power spectral densities at different frequencies from low to high (or in another fixed order) forms an ordered sequence, known as the power spectral density sequence.

[0126] Step 1.4: Perform sliding window minimum tracking on the power spectral density sequence:

[0127] Set the window size to 32 frequency points, and construct a sliding window centered on frequency point f;

[0128] The minimum power spectral density within each sliding window is selected as the initial noise estimate at the center frequency of that sliding window.

[0129] The window is slid along the window sequentially (with a step size of 1) to traverse all frequency points, resulting in a complete initial noise estimation sequence. ;

[0130] Exponential smoothing filtering is performed using the exponential smoothing formula, as shown below:

[0131] ;

[0132] Where α=0.05 is the smoothing coefficient, which controls the weighting of new sampled values ​​and historical estimates; This indicates the smoothing result of the previous frequency point.

[0133] A 5th-order Butterworth bandpass filter is used to perform frequency domain filtering on the signal; including:

[0134] Step 2.1: Carrier frequency tracking: Use a carrier frequency tracking algorithm to detect the carrier frequency position of the signal in real time. ;

[0135] Step 2.2: Dynamic Filter Settings: Dynamically adjust the passband range of the filter according to the carrier frequency position. Ensure that the main frequency components of the signal are effectively captured.

[0136] Step 2.3: Filtering: Perform filtering operation on the time-frequency matrix S(t, f); suppress interference signals outside the filter passband.

[0137] Step 2.4: Noise Suppression: Set the stopband attenuation of the filter to no less than 40dB. Ensure that the filter can effectively suppress noise in the out-of-passband frequency band.

[0138] Soft thresholding denoising using soft short-time Fourier transform results; including:

[0139] Step 2.5: Adaptive Threshold Calculation; Calculate the adaptive threshold for each frequency point according to the following formula:

[0140] ;

[0141] in, The resulting noise substrate;

[0142] Step 2.6: Soft threshold filtering: When the amplitude at a given frequency point At that time, retain the signal components and subtract the threshold from them, that is: , This refers to an adaptive threshold; when When the frequency point is considered to be noise, it is set to zero.

[0143] Denoising effect: This denoising method can effectively preserve the main energy of the signal and specifically suppress noise-dominated frequency bands. Experiments show that, especially in the case of low signal-to-noise ratio (<-10dB), the signal-to-noise ratio can be improved by 3-5dB.

[0144] By using the above methods, time-frequency feature enhancement can effectively reduce the impact of noise on the signal while preserving useful information, thereby improving the accuracy of signal-to-noise ratio estimation.

[0145] Calculate the energy ratio characteristics; including:

[0146] Step 3.1: Calculate the total energy of the denoised signal. As shown below:

[0147] ;

[0148] in, , where is the number of time frames; , where is the number of frequency points; This refers to the denoised signal;

[0149] Step 3.2: Calculate the total noise energy As shown below:

[0150] ;

[0151] in, As a noise floor, For time frames;

[0152] Step 3.3: Calculate the energy ratio characteristics As shown below:

[0153] .

[0154] This formula, this characteristic, reflects the overall energy relationship between the signal and the noise.

[0155] Calculate the peak sharpness characteristics of the spectrum; including:

[0156] Step 3.4: Extract the power spectrum of the denoised signal As shown below:

[0157] ;

[0158] Step 3.5: Find the main lobe peak of the power spectrum As shown below:

[0159] ;

[0160] Step 3.6: Calculate the 3dB bandwidth :

[0161] Find the power spectrum greater than or equal to All frequency points; the difference between the maximum and minimum values ​​among these frequency points is the 3dB bandwidth;

[0162] Step 3.7: Calculate the peak sharpness characteristics As shown below:

[0163] .

[0164] As the formula shows, this feature reflects the degree of concentration in the signal spectrum.

[0165] Extracting deep learning features; including:

[0166] A lightweight convolutional neural network architecture is constructed, which includes an input layer, a first convolutional layer, a second convolutional layer, and a global average pooling layer.

[0167] In the input layer, the input is a denoised time-frequency matrix S(t,f) of 256×256;

[0168] The first convolutional layer consists of 16 3×3 convolutional kernels with a stride of 1 and ReLU activation;

[0169] The second convolutional layer consists of 32 3×3 convolutional kernels with a stride of 1 and ReLU activation;

[0170] Global average pooling layer: converts the feature map into a 32-dimensional feature vector F.

[0171] Feature fusion and regression; including:

[0172] Step 3.8: Feature Normalization:

[0173] For three types of features And deep learning features, namely 32-dimensional feature vectors F;

[0174] Normalize them separately to get F1, F2, and F3;

[0175] Step 3.9: Integration of attention mechanisms:

[0176] Attention mechanism fusion through attention networks;

[0177] Attention networks consist of a first-layer MLP and a second-layer MLP; such as Figure 4 As shown;

[0178] The input consists of the "feature descriptors" F1, F2, and F3 of the three features to be fused.

[0179] Typically, global average pooling is first performed on each feature to convert high-dimensional features into low-dimensional vectors, which are then used as input to the attention network.

[0180] After passing through the first layer of MLP, which consists of 64 neurons and uses ReLU as the activation function, non-linearity is introduced to enhance the feature expression capability;

[0181] After passing through a second MLP layer, which consists of 3 neurons (consistent with the number of features to be fused), the activation function is Softmax, ensuring that the sum of the 3 output weights is 1. ;

[0182] Output 3 weights , respectively corresponding The importance of , with a value range of [0,1]).

[0183] The attention mechanism fusion process involves: combining the three original features to be fused, including... Input attention network; preprocessed Extract feature descriptors (e.g., obtain low-dimensional vectors through global pooling), input them into the first and second MLP layers, and learn to output weights. The larger the weight, the higher the proportion of the corresponding feature in the fusion.

[0184] Step 3.10: Finally, calculate the signal-to-noise ratio. The estimate is shown in the following formula:

[0185] .

[0186] Where F1, F2, and F3 are the three normalized features (normalized to F1, F2, and F3 respectively), and satisfy a1+a2+a3=1.

[0187] Drone signal recognition is achieved through a trained drone signal classification model; including:

[0188] Traditional UAV signal classification methods rely excessively on single-dimensional features such as amplitude and phase of IQ time-domain signals. In environments with strong electromagnetic interference, these features are easily masked by sudden impulse noise and co-channel interference signals. For example, when there is -5dB of broadband noise, the signal-to-noise ratio loss of traditional time-domain features can reach more than 30%, causing the classification accuracy of mainstream machine learning algorithms (such as SVM and random forest) to plummet to below 60%, making it difficult to meet the model identification requirements in complex scenarios. To address this, this invention proposes a novel Multi-Scale Frequency-Domain Attention Network (MSFA-Net), which achieves high-precision classification of UAV models by fusing time-frequency feature enhancement, dynamic interference suppression, and hierarchical feature distillation mechanisms.

[0189] The overall network architecture is as follows Figure 2 As shown;

[0190] The UAV signal classification model is a multi-scale frequency domain attention network, including an input layer, a spectral enhancement module, a deep feature distillation module, a classification head, and an output layer.

[0191] The input layer receives the time-frequency spectrum generated by the short-time Fourier transform, which is the UAV time-frequency spectrum (256×256). This spectrum preserves the energy distribution characteristics of the signal in the time-frequency two-dimensional plane and is more resistant to random noise interference than the original IQ sequence.

[0192] The spectral enhancement module, as the front-end processing unit of the UAV signal classification model, achieves multi-scale feature extraction through three parallel branches: the first branch is a 1×1 convolution branch; the second branch is a 3×3 depthwise separable convolution branch; and the third branch is a 5×5 dilated convolution branch (dilation rate = 2). The features output by the three branches are fused through adaptive weights to form an initial enhanced feature map.

[0193] The deep feature distillation module consists of progressive processing stages, each stage including several FABlock (Frequency-Attention Block) units. Unlike traditional residual blocks, FABlock innovatively introduces a dual-path structure of "frequency domain channel separation - spatial attention calibration": in the FABlock unit, the main path achieves channel dimension compression through 1×1 convolution, and the branch paths use wavelet transform to decompose into 3 frequency band components. , , The importance weights (ranging from 0 to 1) of each frequency band are calculated through a squeezing excitation mechanism, and noise-dominant frequency bands are suppressed through a gating mechanism. Feature map downsampling is achieved between stages using convolutions with a stride of 2. Finally, a 128-dimensional high-dimensional feature vector is output in the fifth stage. The FABlock unit is the core innovation of MSFA-Net, and its structure mainly includes three functional components:

[0194] The FABlock unit includes a frequency band separation mechanism, dynamic attention calibration, and residual fusion path;

[0195] The frequency band separation mechanism refers to the decomposition of the initial enhanced feature map into low-frequency approximation components and high-frequency detail components through two-dimensional wavelet transform. Specifically, the two-dimensional wavelet transform first performs horizontal high-pass and low-pass filtering and downsampling on each row of the initial enhanced feature map, and then performs vertical high-pass and low-pass filtering and downsampling on each column of the result. Finally, it decomposes into one low-frequency approximation component that retains the core smoothing information and three high-frequency detail components that reflect details in the vertical, horizontal and diagonal directions, respectively. Among them, the low-frequency component retains the stable characteristics of the signal (such as carrier frequency and modulation method), while the high-frequency component contains subtle differences unique to the model (such as frequency hopping rate and filtering characteristics).

[0196] Dynamic attention calibration refers to applying a spatial attention module (implemented through 7×7 convolution) and a channel attention module (implemented through the publicly available Squeeze-and-Excitation Networks (SE) module) to low-frequency approximate components and high-frequency detail components respectively to generate a frequency band-level attention mask; the spatial attention module is implemented through 7×7 convolution, and the SE module is a mature existing module;

[0197] The residual fusion path refers to: fusing the initial enhanced feature map through a 1×1 convolution, adding it to the main path's features (i.e., the frequency band-level attention mask), and then passing it through the BatchNorm and SiLU activation functions to achieve feature enhancement and dimensionality unification;

[0198] The classification head includes a global average pooling layer and a multilayer perceptron (MLP). The hidden layers of the MLP introduce Dropout (ratio 0.3) and L2 regularization (coefficient 1e-5), which effectively alleviates the overfitting problem.

[0199] The output layer outputs the probability distribution of 8 common drone models through the Softmax activation function. The 8 common drone models include DJI Mavic2, Mavic3, Tello, Air3, Phantom3, and three non-DJI drones, defined as UAV1, UAV2, UAV3, UAV4, UAV5, UAV6, UAV7, and UAV8 respectively.

[0200] The mask value is calculated using the following formula:

[0201]

[0202] in, This is the feature map of the f-th frequency band, where f = 1, 2, 3. For the Sigmoid function, The attention weight for the f-th frequency band, It is a feature map Perform global average pooling to extract global average information; It is a feature map The channel attention module is used to adaptively learn and weight the importance of channels.

[0203] To demonstrate the superiority of the method of this invention, the proposed model was tested using self-collected DroneRF2025 drone data. The DroneRF2025 dataset involved in this invention includes eight drone models (Table 1 below), showing detailed information on the collection of this DroneRF2025 dataset. This dataset focuses on the radio frequency (RF) signal characteristics of drones, covering eight typical drone models, including DJI Mavic2, Mavic3, Tello, Air3, Phantom3, and UAV1-UAV3 (non-DJI models), comprehensively covering consumer and professional drone signal scenarios. The acquisition process strictly controlled experimental variables, using professional equipment (such as a high-performance microwave spectrum analyzer) and standardized procedures to ensure stable signal acquisition under parameters such as a center frequency of 2.44GHz, a measurement bandwidth of 100MHz, and a sampling frequency of 125MHz. This provides high-quality, reproducible dataset support for model testing, ensuring the reliability and comparability of experimental results.

[0204] Table 1. Detailed information on the DroneRF2025 dataset collected in this invention;

[0205]

[0206] The experimental data collection process is as follows:

[0207] Four researchers are required to collect experimental data.

[0208] The experimenter 1 uses a handheld drone controller 1 to control the drone 1 to fly. The flight speed is controlled at 3-10 m / s and the flight time is 5-10 seconds. At this time, the drone remote controller is 2-10 meters away from the antenna. During the flight, the control commands are constantly changed, such as forward, backward, left, right, up and down.

[0209] Experimenter 2 uses handheld drone controller 2 to control drone 2 to fly, and the flight requirements are the same as described above;

[0210] Experimenter 3 used a stopwatch to time the flight. One second after the drone started flying, the experimenter signaled experimenter 2 to save the drone's radio frequency signal and observe whether the flight was proceeding normally. At the same time, a handheld laser rangefinder (hereinafter referred to as the rangefinder) was used to measure the distance between the drone and the antenna in real time. If the drone was about to be out of range, the experimenter signaled experimenter 1 to adjust the drone's direction.

[0211] Experimenter 4 first adjusted the 4051B high-performance microwave spectrum analyzer (hereinafter referred to as the spectrum analyzer). After being instructed by experimenter 2, the experimenter immediately operated the spectrum analyzer to save the UAV radio frequency signal.

[0212] Experimental results are as follows Figure 3As shown in the matrix, the diagonal values ​​represent the percentage of correct classifications for each drone model, reflecting the core classification performance. The matrix reveals that the correct classification rates for each model are at a high level (95.02% for Mavic2, 97.76% for Mavic3, and 96.50% for Tello, etc.), with most models exceeding 95%. This indicates that, supported by the DroneRF2025 dataset, the proposed model can effectively capture the differences in radio frequency signal features among different drones, achieving high-accuracy classification and verifying the feasibility and superiority of the model in drone classification tasks. The off-diagonal values ​​represent the percentage of misclassifications, reflecting the model's confusion with models possessing similar features. The overall misclassification rate is low (mostly <1%), indicating that the model has good feature discrimination ability among different drone models. For example, only 0.44% of the Mavic 2 models were misclassified as Mavic 3. This reflects that although the two DJI models have similar signal features (such as similar communication protocols and frequency band utilization patterns), the model can still accurately extract subtle differences (such as carrier frequency fine-tuning and signal modulation details) through deep semantic learning frameworks (such as CCRM cross-channel feature enhancement and FFA hierarchical feature fusion), reducing the probability of confusion and highlighting the effectiveness of feature engineering and model architecture.

[0213] Example 3

[0214] A computer device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the UAV signal recognition method described in Embodiment 1 or 2.

[0215] Example 4

[0216] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the UAV signal recognition method described in Embodiment 1 or 2.

[0217] Example 5

[0218] A drone signal identification system, comprising:

[0219] The adaptive noise basis estimation module is configured to perform a short-time Fourier transform on the input UAV IQ signal (i.e., the UAV signal with length N=4096) to obtain the time-frequency matrix S(t,f) (size 256×256), and generate the UAV time-frequency map, as shown below. Figure 1 As shown, the noise power spectrum is estimated using an improved minimum statistics method, and the power spectral density at each frequency point is calculated;

[0220] The time-frequency feature enhancement module is configured to: perform frequency domain filtering on the signal using a 5th-order Butterworth bandpass filter; and perform soft thresholding domain denoising using the soft short-time Fourier transform result.

[0221] The multi-feature fusion regression module is configured to: calculate energy ratio features and spectral peak sharpness features, extract deep learning features, and perform feature fusion and regression.

[0222] The drone signal recognition module is configured to recognize drone signals using a trained drone signal classification model.

Claims

1. A method for identifying signals from unmanned aerial vehicles (UAVs), characterized in that, include: Step 1: Adaptive noise basis estimation; Perform a short-time Fourier transform on the input UAV IQ signal (i.e., the UAV signal) to obtain the time-frequency matrix, generate the UAV time-frequency diagram, estimate the noise power spectrum, and calculate the power spectral density at each frequency point. Step 2: Time-frequency feature enhancement; A fifth-order Butterworth bandpass filter is used to perform frequency domain filtering on the signal; Soft thresholding domain denoising is employed based on the soft short-time Fourier transform results; Step 3: Multi-feature fusion and regression; calculate energy ratio features and spectral peak sharpness features, extract deep learning features, and perform feature fusion and regression; Step 4: Achieve drone signal recognition using the trained drone signal classification model; including: The UAV signal classification model is a multi-scale frequency domain attention network, including an input layer, a spectral enhancement module, a deep feature distillation module, a classification head, and an output layer. The input layer receives the time-frequency spectrum generated by the short-time Fourier transform, which is the UAV time-frequency spectrum; The spectral enhancement module, as the front-end processing unit of the UAV signal classification model, achieves multi-scale feature extraction through three parallel branches: the first branch is a 1×1 convolution branch; the second branch is a 3×3 depthwise separable convolution branch; and the third branch is a 5×5 dilated convolution branch. The features output by the three branches are fused through adaptive weights to form an initial enhanced feature map. The deep feature distillation module consists of progressive processing stages, each stage including several FABlock units; The FABlock unit includes a frequency band separation mechanism, dynamic attention calibration, and residual fusion path; The frequency band separation mechanism refers to the process of decomposing the initial enhanced feature map into low-frequency approximate components and high-frequency detail components through two-dimensional wavelet transform. Dynamic attention calibration refers to generating band-level attention masks by applying spatial attention modules and channel attention modules to low-frequency approximate components and high-frequency detail components, respectively. The residual fusion path refers to: fusing the initial enhanced feature map through a 1×1 convolution, adding it to the main path's features (i.e., the frequency band-level attention mask), and then passing it through the BatchNorm and SiLU activation functions to achieve feature enhancement and dimensionality unification; The classification head includes a global average pooling layer and a multilayer perceptron, where the hidden layers of the multilayer perceptron introduce Dropout and L2 regularization; The output layer outputs the probability distribution of 8 common drone models through the Softmax activation function. The 8 common drone models are defined as UAV1, UAV2, UAV3, UAV4, UAV5, UAV6, UAV7 and UAV8. The mask value is calculated using the following formula: M f =σ(Conv7x7(GlobalAvgPool(F f ))+SE(F f )) Among them, F f Let f be the feature map of the f-th frequency band, where f = 1, 2, 3, σ is the Sigmoid function, and M... f For the attention weight of the f-th frequency band, GlobalAvgPool(F) f )) is for feature map F f Perform global average pooling to extract global average information; SE(F f ) is for feature map F f The channel attention module is used to adaptively learn and weight the importance of channels.

2. The UAV signal recognition method according to claim 1, characterized in that, Adaptive noise basis estimation; including: 1.1: Obtain the time-frequency matrix S(t,f); After performing a short-time Fourier transform on the input IQ signal, a time-frequency matrix S(t,f) is obtained, where dimension t represents a time frame and dimension f represents a frequency point. Each element S(t,f) is a complex number, representing the time-frequency amplitude and phase information of the signal at the t-th frame and the f-th frequency point. In the STFT, the phase information is the time position mark of a specific frequency component in the corresponding time frame. 1.2: Calculate the energy value at each time frequency point; The energy value at each time frequency point is obtained by taking the square of the modulus of each element in the time-frequency matrix S(t,f). 1.3: The power spectral density is obtained by averaging over the time dimension; For each fixed frequency point f, the energy values ​​of all time frames at that frequency are averaged to obtain the power spectral density P(f) at that frequency point. Arrange the power spectral densities at different frequency points from low to high to form an ordered sequence, namely the power spectral density sequence. 1.4: Sliding window minimum tracking of the power spectral density sequence: Build a sliding window; The minimum power spectral density within each sliding window is selected as the initial noise estimate at the center frequency of that sliding window. By sliding the window sequentially and traversing all frequency points, the complete initial noise estimation sequence P is obtained. n (f); Exponential smoothing filtering is performed using the exponential smoothing formula, as shown below: Where α = 0.05 is the smoothing coefficient. This indicates the smoothing result of the previous frequency point.

3. The UAV signal recognition method according to claim 1, characterized in that, A 5th-order Butterworth bandpass filter is used to perform frequency domain filtering on the signal; including: 2.1: Carrier frequency tracking: The carrier frequency position f0 of the signal is detected in real time using a carrier frequency tracking algorithm; 2.2: Dynamic Filter Settings: Based on the carrier frequency position, dynamically adjust the passband range of the filter to [f0-200kHz, f0+200kHz]; 2.3: Filtering: Perform filtering operation on the time-frequency matrix S(t,f); 2.4: Noise Suppression: Set the stopband attenuation of the filter to be no less than 40dB.

4. The UAV signal identification method according to claim 1, characterized in that, Soft thresholding denoising using soft short-time Fourier transform results; including: 2.5: Adaptive Threshold Calculation; The adaptive threshold for each frequency point is calculated according to the following formula: in, The resulting noise substrate; 2.6: Soft Threshold Filtering: When the amplitude at a given frequency |S(t,f)| ≥ τ(f), the signal component is retained and the threshold is subtracted from it, i.e.: S denoise (t,f)=S(t,f)-τ(f), where τ(f) refers to the adaptive threshold; when ∣ When S(t,f)∣<τ(f), the frequency point is considered noise and is set to zero.

5. The UAV signal identification method according to claim 4, characterized in that, Calculate the energy ratio characteristics; including: 3.1: Calculate the total energy E of the denoised signal. sig As shown below: Where, N t =256, where N is the number of time frames; f =256, which is the number of frequency points; S denoise (t,f) refers to the denoised signal; 3.2: Calculate the total noise energy E noise As shown below: Among them, P n * (f) represents the noise floor, N t For time frames; 3.3: Calculation of the energy ratio characteristic R E As shown below:

6. The method for identifying unmanned aerial vehicle (UAV) signals according to claim 1, characterized in that, Calculate the peak sharpness characteristics of the spectrum; including: 3.4: Extracting the power spectrum P of the denoised signal sig (f), as shown below: 3.5: Finding the main lobe peak P of the power spectrum peak As shown below: 3.6: Calculate the 3dB bandwidth BW3dB: Find the power spectrum greater than or equal to P peak / 2 of all frequency points; take the difference between the maximum and minimum values ​​among these frequency points, which is the 3dB bandwidth; 3.7: Calculation of Spectral Peak Sharpness Feature R S As shown below:

7. The method for identifying unmanned aerial vehicle (UAV) signals according to claim 1, characterized in that, Extracting deep learning features; including: A lightweight convolutional neural network architecture is constructed, which includes an input layer, a first convolutional layer, a second convolutional layer, and a global average pooling layer. In the input layer, the input is a denoised time-frequency matrix S(t,f) of 256×256; The first convolutional layer consists of 16 3×3 convolutional kernels with a stride of 1 and ReLU activation; The second convolutional layer consists of 32 3×3 convolutional kernels with a stride of 1 and ReLU activation; Global average pooling layer: converts the feature map into a 32-dimensional feature vector F.

8. The UAV signal identification method according to claim 6, characterized in that, Feature fusion and regression; including: 3.8: Feature Normalization: For three types of features R E R S The deep learning features, i.e., the 32-dimensional feature vector F, are normalized to become F1, F2, and F3 respectively. 3.9: Attention Mechanism Fusion: Attention mechanism fusion through attention networks; Attention networks consist of a first-layer MLP and a second-layer MLP; The input consists of the F1, F2, and F3 values ​​of the three features to be fused. First, perform global average pooling on each feature to convert the high-dimensional features into low-dimensional vectors, which are then used as input to the attention network. After passing through the first layer of MLP, which consists of 64 neurons and uses ReLU as the activation function, non-linearity is introduced to enhance the feature expression capability; After passing through the second layer of MLP, which includes 3 neurons and uses the Softmax activation function, it is ensured that the sum of the 3 weights of the output is 1, i.e., a1+a2+a3=1; Output three weights a1, a2, a3, which correspond to the importance of F1, F2, and F3 respectively, with values ​​ranging from [0,1]. 3.10: Finally, the signal-to-noise ratio is calculated. The estimate is shown in the following formula:

9. A drone signal identification system, characterized in that, include: The adaptive noise basis estimation module is configured to: perform a short-time Fourier transform on the input UAV IQ signal (i.e., the UAV signal) to obtain the time-frequency matrix, generate the UAV time-frequency map, estimate the noise power spectrum, and calculate the power spectral density at each frequency point; The time-frequency feature enhancement module is configured to: perform frequency domain filtering on the signal using a 5th-order Butterworth bandpass filter; and perform soft thresholding domain denoising using the soft short-time Fourier transform result. The multi-feature fusion regression module is configured to: calculate energy ratio features and spectral peak sharpness features, extract deep learning features, and perform feature fusion and regression. The drone signal recognition module is configured to: recognize drone signals using a trained drone signal classification model; including: The UAV signal classification model is a multi-scale frequency domain attention network, including an input layer, a spectral enhancement module, a deep feature distillation module, a classification head, and an output layer. The input layer receives the time-frequency spectrum generated by the short-time Fourier transform, which is the UAV time-frequency spectrum; The spectral enhancement module, as the front-end processing unit of the UAV signal classification model, achieves multi-scale feature extraction through three parallel branches: the first branch is a 1×1 convolution branch; the second branch is a 3×3 depthwise separable convolution branch; and the third branch is a 5×5 dilated convolution branch. The features output by the three branches are fused through adaptive weights to form an initial enhanced feature map. The deep feature distillation module consists of progressive processing stages, each stage including several FABlock units; The FABlock unit includes a frequency band separation mechanism, dynamic attention calibration, and residual fusion path; The frequency band separation mechanism refers to the process of decomposing the initial enhanced feature map into low-frequency approximate components and high-frequency detail components through two-dimensional wavelet transform. Dynamic attention calibration refers to generating band-level attention masks by applying spatial attention modules and channel attention modules to low-frequency approximate components and high-frequency detail components, respectively. The residual fusion path refers to: fusing the initial enhanced feature map through a 1×1 convolution, adding it to the main path's features (i.e., the frequency band-level attention mask), and then passing it through the BatchNorm and SiLU activation functions to achieve feature enhancement and dimensionality unification; The classification head includes a global average pooling layer and a multilayer perceptron, where the hidden layers of the multilayer perceptron introduce Dropout and L2 regularization; The output layer outputs the probability distribution of 8 common drone models through the Softmax activation function. The 8 common drone models are defined as UAV1, UAV2, UAV3, UAV4, UAV5, UAV6, UAV7 and UAV8. The mask value is calculated using the following formula: M f =σ(Conv7x7(GlobalAvgPool(F f ))+SE(F f )) Among them, F f Let f be the feature map of the f-th frequency band, where f = 1, 2, 3, σ is the Sigmoid function, and M... f For the attention weight of the f-th frequency band, GlobalAvgPool(F) f )) is for feature map F f Perform global average pooling to extract global average information; SE(F f ) is for feature map F f The channel attention module is used to adaptively learn and weight the importance of channels.

Citation Information

Patent Citations

  • SSVEP identification method based on time, frequency and time-frequency domain analysis and deep learning

    CN115581467A

  • Fusion feature classification and identification method based on noise reduction underwater acoustic signals

    CN118351881A

Cited By

  • An unmanned aerial vehicle frequency system real-time identification method and system based on fusion of transfer learning

    CN122365197A