Detection and identification method and system based on radio frequency signal of unmanned aerial vehicle

By performing frequency conversion processing and short-time Fourier transform on the drone radio frequency signals, a time-frequency graph data set is generated, and feature extraction and classification is used for deep learning models, the difficulty of drone detection in complex environments and low signal-to-noise ratio scenarios is solved, and high-precision and efficient recognition effects are achieved.

CN120123749APending Publication Date: 2025-06-10XIAN UNIV OF POSTS & TELECOMM +1
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510153906.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In complex environments and low signal-to-noise ratio scenarios, drone detection based on radio frequency signals faces problems such as difficulty in signal identification, inaccurate signal feature extraction, reduced detection range, high false alarm and leak alarm rates, and excessive detection delay.

Method used

By collecting the radio frequency signals of the drone, performing frequency conversion processing and short-time Fourier transform, generating time-frequency graph data sets, and using deep learning models for feature extraction and classification, combining adaptive time windows and Gaussian white noise simulation, the recognition accuracy and robustness are improved.

Benefits of technology

It significantly improves the recognition effect of drones in complex environments, enhances feature extraction capabilities, reduces false alarms and missed alarm rates, and improves the real-time and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123749A_ABST
    Figure CN120123749A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of communication, and discloses a detection and identification method and system based on a radio frequency signal of an unmanned aerial vehicle, and the method considers how to utilize zero-intermediate-frequency radio frequency signal collection, short-time Fourier transform and an improved deep learning model in a complex urban electromagnetic interference environment, reduces the detection delay as much as possible, and improves the detection accuracy. And high-accuracy identification of the low-slow small unmanned aerial vehicle is maintained, and the method has a rapid expansion capability for a new target type. Through the detection and identification method provided by the invention, high-precision detection and classification of common civil unmanned aerial vehicles and potential novel unmanned aerial vehicles can still be realized under the conditions of urban environment multi-source interference, complex noise and variable flight heights of low-altitude small targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and particularly to a method and system for detecting and identifying an unmanned aerial vehicle (UAV) radio frequency signal. Background Art

[0002] In recent years, unmanned aerial vehicles (UAVs) have been rapidly developed and popularized in both civilian and military fields, and are widely used in multiple scenarios such as photogrammetry, logistics transportation, emergency rescue, and public safety. However, the popularization of UAVs has also brought a series of new safety and regulatory issues, such as flight safety around airports and low-altitude monitoring in sensitive urban areas. To address these challenges, UAV detection and identification technologies have become an important research direction and application requirement. The current mainstream UAV detection methods can be mainly classified into the following categories: The first category is vision image-based detection, which identifies aircraft targets in the airspace through computer vision algorithms. For example, a target detection model is used to detect UAVs in a video or image sequence. However, this method is sensitive to lighting conditions, occlusion, and shooting angles, and its performance significantly degrades at night or in extreme weather conditions; the second category is radar signal-based detection, which uses a radar system to transmit electromagnetic waves and receive echoes, and detects UAVs through Doppler frequency shift or target signal characteristics. Although radar technology has a relatively high detection range in some cases, radar systems are usually costly, and there is a certain probability of missed detection under low-altitude and clutter-rich terrain conditions; traditional methods based on optoelectronic or radar means have relatively high missed detection and false detection rates in occlusion environments, multi-source interference, or urban backgrounds with high-rise buildings; the third category is acoustic feature-based detection. When a UAV is flying, its propellers generate specific noise spectra, which can be captured by a sound sensor. However, in the actual environment, there is a large amount of noise interference, and the acoustic characteristics of different types of UAVs often vary greatly, resulting in unstable detection accuracy; the fourth category is radio frequency (RF) signal-based detection. When a UAV is flying, it needs to communicate with a remote controller or a ground station, and correspondingly generates unique RF communication signals. By receiving, demodulating, and analyzing these signals, different types or models of UAVs can be distinguished to a certain extent.

[0003] Compared with other methods, UAV detection based on RF signals has certain concealment and anti-illumination capabilities, does not rely on visible light conditions or expensive radar equipment, and is particularly suitable for some application scenarios, such as urban low-altitude monitoring or night detection.

[0004] However, the following problems exist in complex environments and low signal-to-noise ratio scenarios:

[0005] 1. Difficult signal recognition

[0006] In complex environments (such as city centers, industrial parks, etc.), there are a large number of RF interference sources, such as numerous communication base stations, Wi-Fi routers, Bluetooth devices, and other radio devices. The RF signals generated by these interference sources will be superimposed on the RF signals of drones in terms of frequency, amplitude, etc. In the case of low signal-to-noise ratio, the RF signals of drones may be submerged in noise and interference signals, making it difficult for detection devices to accurately identify the unique RF signals of drones from among numerous signals.

[0007] For example, in a bustling commercial district, there are many Wi-Fi devices in stores working simultaneously, and the signal frequency bands they emit are similar to those of some drone RF communication frequency bands. When a drone enters this area, the detection system receives a complex mixed signal containing multiple signals, and it is very difficult to distinguish the drone signal from it, just as difficult as distinguishing a soft whisper of a specific person in a noisy crowd.

[0008] 2. Inaccurate signal feature extraction

[0009] Due to the complex environment and low signal-to-noise ratio, the received drone RF signals will be distorted. This will affect the extraction of signal features. For example, features such as the modulation method, frequency, and phase of the signal may change due to noise and interference.

[0010] For different models of drones, the characteristics of their RF signals are the key basis for distinguishing them. With the continuous update of drone models and flight modes, traditional methods are difficult to adapt to the diverse communication characteristics of new targets in a timely manner. However, in this scenario, inaccurate signal feature extraction will lead to the inability to correctly distinguish the types and models of drones. For example, the RF signal of a certain model of drone normally uses a specific modulation method, but in a complex environment and low signal-to-noise ratio, the characteristics of this modulation method may be distorted, and the detection system may misjudge the model of the drone.

[0011] 3. Reduced detection range

[0012] Obstacles in complex environments (such as high-rise buildings) will have effects such as reflection, refraction, and absorption on RF signals. In the case of low signal-to-noise ratio, the signal itself is relatively weak, and after attenuation by obstacles, the signal strength is further reduced. This greatly reduces the range within which the detection device can effectively receive drone RF signals.

[0013] For example, in a mountainous environment, the blocking and absorption effects of the mountain on RF signals are obvious. If a drone is flying on the other side of the mountain, in the case of low signal-to-noise ratio, the detection system may not be able to receive the RF signal of the drone. Even in the case of slightly better signals, its detection range will be much smaller than in an open environment.

[0014] 4. Increase in false alarm and missed alarm rates

[0015] False alarms refer to the situation where the detection system erroneously determines interference signals as UAV signals, while missed alarms refer to the situation where actual UAV signals are not detected. In complex environments and at low signal-to-noise ratios, interference signals are easily misjudged as UAV signals, and at the same time, due to the fact that UAV signals may be masked by noise, missed alarms are also likely to occur.

[0016] For example, when certain interference signals in the environment are similar to some characteristics of UAV signals in terms of frequency and amplitude, the detection system may issue false alarms. And when UAV signals are too weak to be ignored as noise, missed alarms will occur. Both of these situations will seriously affect the reliability of the detection system.

[0017] 5. Excessively high detection latency

[0018] In terms of real-time performance, excessively high detection latency will significantly restrict the efficiency of security monitoring and emergency prevention and control.

[0019] In view of this requirement, the present invention focuses on how to better detect and identify UAV radio frequency signals, and provides an effective solution for complex environments and low signal-to-noise ratio (SNR) scenarios. Summary of the Invention

[0020] Embodiments of the present invention provide a method and system for detecting and identifying UAV radio frequency signals to solve the problem of how to achieve high-precision detection and classification of common civilian UAVs and potential new UAVs in the case of multi-source interference, complex noise, and variable flight altitudes of low-altitude small targets in urban environments. To have a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a preface to the subsequent detailed description.

[0021] According to the first aspect of the embodiments of the present invention, a method for detecting and identifying UAV radio frequency signals is provided, including the following steps:

[0022] S1: Collect the radio frequency signals of micro UAVs, convert the radio frequency signals into zero intermediate frequency signals through frequency conversion processing to obtain an initial radio frequency signal sample library;

[0023] S2: Considering the time window with an adaptive duration and the balance relationship between the recognition accuracy rate and the inference time, construct an objective function for the time window length, and combine with the elbow method to obtain a given set of candidate time window lengths. Perform moving sliding window interception on the original radio frequency signals in the initial radio frequency signal sample library to obtain short-time radio frequency signal sequence samples;

[0024] S3: Add Gaussian white noise to the short - time radio - frequency signal sequence samples, and perform short - time Fourier transform to generate a time - frequency map dataset with different durations and different signal - to - noise ratios for model training;

[0025] S4: According to the characteristics of the time - frequency maps with different durations and different signal - to - noise ratios, improve and complete the training of the deep - learning model for classifying UAV time - frequency maps, and obtain an intelligent recognition model for UAV time - frequency maps;

[0026] S5: Based on the intelligent recognition model, input the time - frequency map of the signal to be recognized into the trained intelligent recognition model, and output the classification result of the UAV.

[0027] Based on the above - mentioned solution, in step S1, the step of collecting the radio - frequency signal of the micro - UAV, converting the radio - frequency signal into a zero - intermediate - frequency signal through frequency conversion processing to obtain the initial radio - frequency signal sample library specifically includes:

[0028]

[0029] where \(f\) c is the center frequency of the radio - frequency signal, LPF represents a low - pass filter, \(t\) is the time variable, \(x\) BB (t) is the baseband signal obtained after the down - conversion operation, \(j\) is the imaginary unit, satisfying \(j\) 2 \(=-1\), the real part \(x\) I (t) is the in - phase component, and the imaginary part \(x\) Q (t) is the quadrature component;

[0030] In the specific hardware implementation, after mixing the input radio - frequency signal with \(\cos(2\pi f\) c \(t)\) and \(\sin(2\pi f\) c \(t)\) respectively, and then passing through the LPF respectively to obtain the following forms:

[0031]

[0032] where and represent the in - phase baseband signal and the quadrature baseband signal respectively, and then combine them into a complex form

[0033] Based on the above - mentioned solution, step S2: Considering the time window with an adaptive duration and the balance relationship between the recognition accuracy and the inference time, construct an objective function for the time - window length, and combine the elbow method to obtain a given candidate set of time - window lengths, and perform moving - sliding - window interception on the original radio - frequency signal in the initial radio - frequency signal sample library to obtain short - time radio - frequency signal sequence samples, specifically including:

[0034] J(w) = α·Acc(w) - β·Time(w) (3)

[0035] Wherein, Acc(w) represents the recognition accuracy obtained by segmenting the signal with a time window length of w, Time(w) represents the corresponding inference time, and α and β are weight coefficients;

[0036] Combined with the elbow method, analyze the trend curves of Acc(w) and Time(w) changing with w, find the inflection point position as the key observation range, and search within the given set of candidate time window lengths according to formula (3) to select the w that maximizes J(w) * , denoted as:

[0037]

[0038] Wherein, w ∈ W means that w takes values in the given set of candidate time window lengths W. After determining w * , perform moving sliding window interception on the radio frequency signal x BB (t) obtained by intermediate frequency processing; if the total duration of the signal is T and the sliding window moving step size is Δ, then several short-time radio frequency signal sequences x k (t) can be obtained. The specific segmentation method can be seen in formula (5):

[0039] x k (t) = x BB (t + (k - 1)Δ), k = 1, 2, …, K (5)

[0040] Wherein,

[0041] On the basis of the above scheme, step S3: Add Gaussian white noise to the short-time radio frequency signal sequence samples, and generate a time-frequency map dataset with different durations and different signal-to-noise ratios after performing short-time Fourier transform, specifically including:

[0042] Add Gaussian white noise n(t) to the short-time radio frequency signal sequence x k (t) to obtain a new signal x' k (t), as shown in formula (6):

[0043] x' k (t) = x k (t) + n(t) (6)

[0044] Wherein, n(t) can be set as Gaussian white noise with zero mean and variance of σ 2 , and σ 2 is selected according to the target signal-to-noise ratio. Then, perform short-time Fourier transform on the obtained signal x' k (t), as shown in formula (7):

[0045]

[0046] Among them, φ(·) is the window function, τ is the time frame index, f is the frequency index, and m is the discrete time index, which is used to represent the sampling values of the signal at different moments, x' k (m) is the new signal x' obtained after adding Gaussian white noise k (t) at the discrete time point m. By converting the obtained complex-valued time-frequency distribution into an amplitude spectrum, a two-dimensional time-frequency matrix can be obtained. Then, representing this matrix as a time-frequency diagram of size H×W, a time-frequency diagram dataset is constructed.

[0047] Based on the above scheme, step S4: According to the characteristics of the time-frequency diagrams under different durations and different signal-to-noise ratios, improve and complete the training of the deep learning model for classifying UAV time-frequency diagrams to obtain an intelligent recognition model for UAV time-frequency diagrams, specifically including:

[0048] S41: Divide the time-frequency diagrams generated under different durations and different signal-to-noise ratios in step S3 into a training set for iteratively updating the model weights and a validation set for monitoring the generalization ability of the model in the actual application scenario to avoid the decrease in accuracy caused by overfitting;

[0049] S42: Use the multi-level MaR Block to perform local and global feature extraction for the two-dimensional characteristics of the time-frequency diagram in the convolutional path of the residual structure; at the output of each layer of the residual, regard each row of the time-frequency diagram as a one-dimensional sequence and place it inside the Mamba module. Through linear projection, SiLU non-linear mapping, and the state space model SSM, extract the feature dependencies of each frequency point changing with time, and finally fuse it with the residual backbone;

[0050] S43: In the training stage, use the cross-entropy loss function as the main optimization objective, and update the model parameters through backpropagation and stochastic gradient descent (SGD). Among them, the cross-entropy loss function is shown in formula (8):

[0051]

[0052] Among them, L CE represents the cross-entropy loss value, which is used to quantify the gap between the prediction effect of the model on the current batch of data and the actual situation. N batch represents the batch size, c is the number of UAV categories, y i,c is the true label of the i-th training sample in terms of category, then is the predicted probability output by the model;

[0053] S44: During the training process, monitor indicators such as the accuracy and recall rate of the model through the validation set.

[0054] Based on the above solution, step S5: Based on the intelligent recognition model, input the time-frequency diagram of the signal to be recognized into the trained intelligent recognition model, and output the classification result of the UAV, specifically including:

[0055] S51: Perform the same preprocessing, mixing, and short-time Fourier transform processes on the target signal to be recognized as in the training stage to convert the original signal into a time-frequency diagram;

[0056] S52: Input the time-frequency diagram preprocessed in S51 into the first-layer convolution of the trained MaRNet model;

[0057] S53: Introduce a weighted strategy or other post-processing mechanism in the inference stage;

[0058] S54: Connect the final recognition result to the urban low-altitude monitoring platform or other warning systems for real-time monitoring of possible illegal or dangerous targets; once an unknown or abnormal category UAV is recognized, a warning signal can be immediately issued; if a target with insufficient confidence is detected, it can be incorporated into the data collection and annotation link for subsequent incremental training or online fine-tuning.

[0059] Based on the above solution, step S52: Input the time-frequency diagram preprocessed in S51 into the first-layer convolution of the trained MaRNet model, and the MaRNet model specifically includes:

[0060] First, through an initial module named "Stem", use 3×3 convolution, batch normalization (BN), ReLU, and max pooling (MaxPool) to extract and compress the initial spatial and frequency features of the image;

[0061] Subsequently, enter multiple MaR Blocks in sequence. Each MaR Block additionally integrates a Mamba module within the residual structure to capture local features in convolution while using a learnable state space modeling (SSM) mechanism to obtain deeper time-frequency domain dependence information;

[0062] Finally, flatten the accumulated high-level features at the end of the network and input them into the fully connected layer (FC). The output result is converted into a probability vector of length c through the Softmax activation function where each element represents the prediction probability for the corresponding UAV category.

[0063] According to the second aspect of the embodiments of the present invention, a detection and recognition system based on UAV radio frequency signals is provided, including:

[0064] A signal preprocessing module, which is used to collect the radio frequency signals of the micro unmanned aerial vehicle, convert the radio frequency signals into zero intermediate frequency signals through frequency conversion processing, and obtain an initial radio frequency signal sample library;

[0065] A moving sliding window intercepting module, which is used to consider a time window with an adaptive duration and the balance relationship between the recognition accuracy and the inference time, construct an objective function for the time window length, and combine the elbow method to obtain a given candidate set of time window lengths, and perform moving sliding window interception on the original radio frequency signal in the initial radio frequency signal sample library to obtain short-time radio frequency signal sequence samples;

[0066] A Gaussian white noise simulation module, which is used to add Gaussian white noise to the short-time radio frequency signal sequence samples, and generate a time-frequency map data set with different durations and different signal-to-noise ratios for model training after performing short-time Fourier transform;

[0067] A recognition model module, which is used to improve and complete the training of a deep learning model for classifying the time-frequency map of the unmanned aerial vehicle according to the characteristics of the time-frequency maps with different durations and different signal-to-noise ratios, and obtain an intelligent recognition model for the time-frequency map of the unmanned aerial vehicle;

[0068] An output module, which is used to input the time-frequency map of the signal to be recognized into the trained intelligent recognition model based on the intelligent recognition model, and output the classification result of the unmanned aerial vehicle.

[0069] According to the third aspect of the embodiments of the present invention, a computer device is provided, including a memory and a processor, the memory stores a computer program, and it is characterized in that when the processor executes the computer program, the steps of the method are implemented.

[0070] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0071] The present invention enhances the feature extraction ability in different channel states and signal-to-noise ratio environments by performing intermediate frequency conversion, short-time Fourier transform, and Gaussian white noise simulation on radio frequency signals, significantly improving the recognition effect of drones in complex environments. It adopts an adaptive time window and combines the elbow method to balance recognition performance and inference speed. After the original signal is intercepted by a moving sliding window, it can not only retain key information but also effectively control the computational amount and reduce the time delay. A drone recognition framework based on MaR Block is proposed, which can not only extract image-level features from the overall time-frequency diagram but also split each row (time series corresponding to different frequency points) into one-dimensional sequences to deeply capture the details and correlations of each frequency point changing with the time dimension. At the same time, through the linkage of the deep learning model in the training and inference stages, it supports data collection and subsequent learning for new types or unknown targets, overcoming the deficiency that existing methods cannot adapt to new aircraft models or new channel characteristics in a timely manner. In addition, the present invention can be seamlessly connected to the urban low-altitude safety control platform to achieve early warning and positioning of illegal or abnormal flight targets, and trigger an auxiliary detection mechanism in a timely manner when unknown targets or low-confidence results appear, improving the safety protection level of the urban airspace.

[0072] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. Brief Description of the Drawings

[0073] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.

[0074] Figure 1 It is a schematic diagram of constructing drone radio frequency data based on signal reception and preprocessing shown according to an exemplary embodiment;

[0075] Figure 2 It is a schematic diagram of segmenting drone radio frequency signals based on zero intermediate frequency processing and sliding window strategy shown according to an exemplary embodiment;

[0076] Figure 3 It is a schematic diagram of a process for obtaining the optimal time window based on short-time Fourier transform shown according to an exemplary embodiment;

[0077] Figure 4 It is a schematic diagram of a drone recognition network framework based on a multi-branch MaR Block shown according to an exemplary embodiment;

[0078] Figure 5 It is a schematic diagram of a detection and recognition system based on drone radio frequency signals shown according to an exemplary embodiment;

[0079] Figure 6It is a schematic structural diagram of a computer device shown according to an exemplary embodiment. Detailed implementation manners

[0080] The following description and drawings fully illustrate the specific implementation manners herein, enabling those skilled in the art to practice them. Parts and features of some embodiments may be included in or replaced by parts and features of other embodiments. The scope of the embodiments herein includes the entire scope of the claims and all available equivalents of the claims. Herein, terms such as "first", "second", etc. are only used to distinguish one element from another, without requiring or implying any actual relationship or order between these elements. In fact, the first element can also be called the second element, and vice versa. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a structure, device or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such structure, device or equipment. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the structure, device or equipment including the said element. The embodiments herein are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0081] Herein, the character " / " indicates that the front and rear objects are in an "or" relationship. For example, A / B means: A or B.

[0082] Herein, the term "and / or" is a description of the associated relationship of an object, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B these three relationships.

[0083] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0084] Embodiment 1

[0085] Figure 1 An embodiment of the detection and recognition method based on the radio frequency signal of an unmanned aerial vehicle of the present invention is shown.

[0086] S1: Collect the radio frequency signals of common civilian micro unmanned aerial vehicles, convert them into zero intermediate frequency signals through frequency conversion processing, and obtain an initial radio frequency signal sample library;

[0087] S2: Based on the time window with adaptive duration and the balance relationship between the recognition accuracy and the inference time, construct the objective function of the time window length, and combine it with the elbow method to obtain the candidate set of the given time window lengths. Intercept the original RF signal by moving the sliding window within the initial RF signal sample library to obtain the short-time RF signal sequence samples;

[0088] S3: Add Gaussian white noise to the short-time RF signal sequence samples, and perform short-time Fourier transform to generate the time-frequency map dataset with different durations and different signal-to-noise ratios for model training;

[0089] S4: According to the characteristics of the time-frequency maps with different durations and different signal-to-noise ratios, improve and complete the training of the deep learning model for classifying UAV time-frequency maps to obtain the intelligent recognition model of UAV time-frequency maps;

[0090] S5: Based on the intelligent recognition model, input the time-frequency map of the signal to be recognized into the trained intelligent recognition model, and output the classification result of the UAV.

[0091] As Figure 2 shown, as a specific implementation, for step S1, to avoid the interference of the external electromagnetic environment, the collected original signal x(t) can be preprocessed, such as operations like filtering and gain control. Subsequently, to more conveniently extract the characteristics of the UAV RF signal, its RF signal is converted into a zero-intermediate frequency (Zero-IF) signal.

[0092] The present invention uses the following formula to implement the down-conversion operation:

[0093]

[0094] where f c is the center frequency of the RF signal, such as 2.4 GHz or 5.8 GHz, LPF represents the low-pass filter, which is used to filter out the high-frequency components generated by multiplication and thus retain the baseband part, t is the time variable, x BB (t) is the baseband signal obtained after the down-conversion operation, j is the imaginary unit, satisfying j 2 =-1, which is used to represent the imaginary part of the complex signal, and the real part x I (t) is the in-phase component, and the imaginary part x Q (t) is the quadrature component. In the specific hardware implementation, the present invention can mix the input signal with cos(2πf c t) and sin(2πf c t) respectively, and then obtain the following form through the LPF respectively:

[0095]

[0096] Among them, and respectively represent the in-phase baseband signal and the quadrature baseband signal, and are then combined into a complex form Through this step, an initial radio frequency signal sample library containing 10 types of unmanned aerial vehicles can be obtained, laying a foundation for subsequent feature extraction and model training, and can also reduce the hardware complexity during deployment and improve the baseband analysis accuracy.

[0097] As a specific implementation, for step S2, the present invention defines an objective function for the balance relationship between the time window length and the recognition accuracy and inference time, as shown in formula (3):

[0098] J(w) = α·Acc(w) - β·Time(w) (3)

[0099] Among them, Acc(w) represents the recognition accuracy obtained by segmenting the signal with the time window length w, Time(w) represents the corresponding inference time, and α and β are weight coefficients. In order to select the best value that can balance high accuracy and short inference time among multiple candidate time window lengths, the present invention combines the elbow method to analyze the trend curves of Acc(w) and Time(w) with the change of w, finds the "inflection point" position as the key observation range, and searches within the given candidate set of time window lengths according to the formula to select the w that maximizes J(w) * , denoted as:

[0100]

[0101] After determining w * , the radio frequency signal x BB (t) obtained by intermediate frequency processing is intercepted by a moving sliding window. If the total duration of the signal is T and the sliding window moving step size is Δ, several short-time radio frequency signal sequences x k (t) can be obtained. The specific segmentation method can be seen in formula (5):

[0102] x k (t) = x BB (t + (k - 1)Δ), k = 1, 2,..., K (5)

[0103] Among them, Through the above steps, 10 groups of short-time radio frequency signal sequence samples can be obtained while taking into account the recognition accuracy and real-time performance. This process can effectively segment long-time continuous signals and ensure that representative local features can still be extracted in an interference environment.

[0104] As a specific implementation, for step S3, in order to enhance the robustness of the model in different channel states and interference environments, the present invention processes the short-time radio frequency signal sequence xk Add Gaussian white noise \(n(t)\) to \(x(t)\) to obtain a new signal \(x'(t)\) as shown in Equation (6): k (t), as shown in Equation (6):

[0105] x' k (t) = x k (t) + n(t) (6)

[0106] where \(n(t)\) can be set as Gaussian white noise with zero mean and variance \(\sigma^2\), and \(\sigma\) is selected according to the target signal-to-noise ratio. Next, perform a short-time Fourier transform on the obtained signal \(x'(t)\), specifically referring to Equation (7): 2 of Gaussian white noise, and \(\sigma\) 2 is selected according to the target signal-to-noise ratio. Then, perform a short-time Fourier transform on the obtained signal \(x'(t)\), specifically referring to Equation (7): k (t), specifically referring to Equation (7):

[0107]

[0108] where \(\varphi(\cdot)\) is the window function, \(\tau\) is the time-frame index, \(f\) is the frequency index, \(m\) is the discrete-time index used to represent the sampled values of the signal at different times, and \(x'(m)\) is the value of the new signal \(x'(t)\) obtained after adding Gaussian white noise at the discrete-time point \(m\). Convert the obtained complex-valued time-frequency distribution to an amplitude spectrum to obtain a two-dimensional time-frequency matrix, and then represent this matrix as a time-frequency diagram of size \(H\times W\). Finally, construct a time-frequency diagram dataset containing 10 types of drones. The time-frequency diagrams obtained in this step provide richer feature support for subsequent recognition. k (m) is the value of the new signal \(x'(t)\) obtained after adding Gaussian white noise, convert the obtained complex-valued time-frequency distribution to an amplitude spectrum, obtain a two-dimensional time-frequency matrix, and then represent this matrix as a time-frequency diagram of size \(H\times W\). Finally, construct a time-frequency diagram dataset containing 10 types of drones. The time-frequency diagrams obtained in this step provide richer feature support for subsequent recognition. k (t) at the discrete-time point \(m\). Convert the obtained complex-valued time-frequency distribution to an amplitude spectrum to obtain a two-dimensional time-frequency matrix, and then represent this matrix as a time-frequency diagram of size \(H\times W\). Finally, construct a time-frequency diagram dataset containing 10 types of drones. The time-frequency diagrams obtained in this step provide richer feature support for subsequent recognition.

[0109] As a specific implementation, for step S4, after generating the time-frequency diagrams with multiple signal-to-noise ratios, the present invention uses them for training an improved deep learning model in this step to achieve high-precision classification and recognition of target categories. Based on the classic ResNet framework, this model deeply integrates the Mamba module and obtains stronger identification ability for UAV RF signal features through deep fusion in both spatial and temporal dimensions. The specific process is as follows:

[0110] S41: The present invention first divides the time-frequency diagrams generated with different durations and different signal-to-noise ratios in the previous step S3 into a training set and a validation set. Among them, the training set is used to iteratively update the model weights, and the validation set is used to monitor the generalization ability of the model in actual application scenarios to avoid a decrease in accuracy caused by overfitting. If there are some specific samples in a multi-source interference environment such as a city, data augmentation or balancing can also be performed at this stage.

[0111] S42: In the conventional ResNet path of the present invention, the Mamba module is deeply embedded, and hierarchical transformation is performed on the network convolution path, residual connection, and non-linear activation, enabling the Mamba module to fully absorb the time series information from the row vectors of each frequency point while retaining the high-order spatial feature extraction ability in the ResNet block. Specifically: The present invention proposes a novel MaRBlock (incorporating the ideas of the ResNet block and the Mamba module), and multiple MaRBlocks are connected in series to form a "multi-level MaRBlock"; the multi-level MaRBlock is used to extract local and global features for the two-dimensional characteristics of the time-frequency diagram in the convolution path of the residual structure; at the output of each layer of the residual, each row of the time-frequency diagram is regarded as a one-dimensional sequence and placed inside the Mamba module, and through modules such as linear projection, SiLU non-linear mapping, and state space model (SSM), the feature dependencies of each frequency point changing over time are extracted, and finally fused with the residual backbone. This design not only focuses on capturing the image patterns in the two-dimensional space (frequency domain), but also can deeply explore the dynamic behavior of the one-dimensional time series, making the network more robust and expressive in the identification of UAV radio frequency signals under complex interference.

[0112] The deep learning model adopted by the present invention is as Figure 4 shown. For the input time-frequency diagram, the present invention regards it as two-dimensional image data, and its size can be flexibly adjusted according to application requirements, such as C×H×W. In practice, C may be 1 (grayscale image) or 3 (pseudo-color image).

[0113] S43: In the training stage, the present invention uses the cross-entropy loss function as the main optimization objective, and updates the model parameters through backpropagation and stochastic gradient descent (SGD) or other optimization algorithms (such as Adam). The cross-entropy loss function is shown in formula (8):

[0114]

[0115] where L CE represents the cross-entropy loss value, which is used to quantify the gap between the prediction effect of the model on the current batch of data and the actual situation, N batch represents the batch size, c is the number of UAV categories, y i,c is the true label of the i-th training sample in terms of category, and

[0116] is the predicted probability output by the model.

[0117] Provide a specific implementation. For step S5, after the above training is completed, the present invention conducts an inference and recognition process on the UAV radio frequency signals collected in real time or offline. The core lies in: not only using the spatial feature extraction of the ResNet backbone, but also obtaining the time-domain dependence information through the Mamba path after row-segmenting the time-frequency map, so as to achieve fast and high-precision determination of UAV types in the application scenario.

[0118] Specifically, ResNet (Residual Network): is a convolutional neural network architecture in deep learning. Its feature is the introduction of residual connections, which effectively solves the problem of gradient disappearance or gradient explosion that occurs as the network depth increases, and can better extract spatial features in data such as images. In this scenario, using the ResNet backbone means leveraging its powerful spatial feature extraction ability to obtain feature information about the shape, structure, texture, etc. of the UAV in the spatial aspect from the input data (time-frequency map). These spatial features are very important bases for identifying UAV types. For example, different types of UAVs may have different wing shapes, fuselage sizes, color distributions, etc. in appearance, and ResNet can help capture these detailed features.

[0119] Mamba is a module specifically used for processing time series data or obtaining time-domain dependence information. Passing through the Mamba path after row-segmenting the time-frequency map means that after performing a certain form of time-frequency analysis on the data (such as converting the UAV signal into a time-frequency map, which can simultaneously display the information of the signal in time and frequency) and performing row-segmentation processing, Mamba is used to mine the dependence relationships and patterns in the time dimension of the data. For example, during the flight of a UAV, its signals (such as radio frequency signals, flight trajectory data, etc.) may have certain regularities and correlations in time. Mamba can capture this time-domain information, such as the change trend and periodicity of the signal over time. These time-domain information can also provide important supplements for accurately determining UAV types.

[0120] The specific steps of step S5 include:

[0121] S51: The present invention adopts the same preprocessing, mixing frequency, and short-time Fourier transform processes for the target signal to be recognized as in the training stage to convert the original signal into a time-frequency map. This step can ensure that the size and resolution of the time-frequency map are consistent with the samples used during training, thus avoiding model adaptation problems caused by input dimension or resolution mismatches.

[0122] S52: Input the time-frequency diagram preprocessed by S51 into the first layer of convolution of the pre-trained MaRNet model (i.e., the intelligent recognition model). The model first passes through an initial module called "Stem" to extract and compress the initial spatial and frequency features of the image using 3×3 convolution, batch normalization (BN), ReLU, and max pooling (MaxPool). Subsequently, it enters multiple MaR Blocks in sequence. Each MaR Block additionally integrates a Mamba module within the residual structure to capture deeper time-frequency domain dependence information using a learnable state space modeling (SSM) mechanism while the convolution captures local features. Finally, the accumulated high-level features at the end of the network are flattened and input into the fully connected layer (FC), and the output result is transformed into a probability vector of length c through the Softmax activation function. Each element represents the prediction probability for the corresponding UAV category. In this way, the MaRNet model can simultaneously mine complex pattern features such as frequency hopping and noise at both local and global scales to achieve accurate classification of target signals.

[0123] S53: In the inference stage, the present invention can introduce a weighting strategy or other post-processing mechanisms according to application requirements. For example, if the probability distribution is too dispersed or the highest probability is lower than a preset threshold, an auxiliary detection process (such as multi-modal verification or additional sensor information fusion) can be triggered to reduce the risk of misidentification or missed detection. In addition, if combined with an incremental learning framework, signals of suspected new types or unknown targets can be recorded and labeled, and the recognition range of the model can be continuously expanded in subsequent training.

[0124] S54: The present invention can dock the final recognition result with an urban low-altitude monitoring platform or other warning systems for real-time monitoring of possible illegal or dangerous targets. Once an unknown or abnormal category UAV is recognized, a warning signal can be immediately sent to notify relevant security departments to implement intervention or tracking. If a target with insufficient confidence is detected, it can be incorporated into the data collection and labeling process, and subsequent incremental training or online fine-tuning can be carried out to achieve continuous optimization and update of the model.

[0125] The method of the present invention significantly improves the detection speed and accuracy of low, slow, and small targets by performing high-precision feature extraction and classification on UAV radio frequency signals in complex environments, providing important technical support for urban airspace security prevention and control and other UAV recognition scenarios that require high reliability and real-time performance.

[0126] Embodiment 2

[0127] Figure 1An embodiment of the method for detecting and identifying based on UAV radio frequency signals of the present invention is shown. Taking 10 types of UAVs as an example, the specific steps are as follows:

[0128] S1: In a microwave anechoic chamber, collect the radio frequency signals of common civilian UAVs in the 2.4 GHz and 5.8 GHz frequency bands, and perform processing to convert them into zero-intermediate frequency signals, obtaining an initial radio frequency signal sample library containing 10 types of UAVs; (In the experiment, the data used belongs to these two frequency bands. In actual applications, signals of other frequency bands, such as 800 MHz, can be collected and processed. Eventually, through processing, they will be converted into zero-intermediate frequency signals for further processing. After converting to zero-intermediate frequency signals, short-time Fourier transform is performed. After obtaining the time-frequency diagram, it can be sent into the model for training / inference.)

[0129] S2: Use a time window with an adaptive duration and combine it with the elbow method to determine the optimal time window length on the premise of balancing the detection performance and inference time as much as possible. Perform moving window interception on the original radio frequency signal in the initial radio frequency signal sample library to obtain 10 groups of short-time radio frequency signal sequence samples;

[0130] S3: Add Gaussian white noise to the short-time radio frequency signal sequence samples, and perform short-time Fourier transform to generate time-frequency diagrams with different durations and different signal-to-noise ratios, obtaining a dataset containing time-frequency diagrams of 10 types of UAVs for model training. Without adding noise, each type of UAV has 500 time-frequency diagram samples, totaling 5,000 samples. Further, if the signal-to-noise ratio ranges from -12 dB to 15 dB with a step size of 3 dB, a total of 50,000 samples can be generated;

[0131] S4: Improve and complete the training of the deep learning model for classifying UAV time-frequency diagrams according to the characteristics of the time-frequency diagrams, obtaining an intelligent recognition model for UAV time-frequency diagrams;

[0132] S5: Input the time-frequency diagram of the signal to be recognized into the intelligent recognition model and output the classification result of the UAV.

[0133] As Figure 2 shown, for step S1, to avoid interference from the external electromagnetic environment, the collected original signal x(t) can be preprocessed, such as operations like filtering and gain control. Subsequently, to more conveniently extract the characteristics of the UAV radio frequency signal, its radio frequency signal is converted into a zero-intermediate frequency (Zero-IF) signal. The present invention uses the following formula to implement the down-conversion operation:

[0134]

[0135] where, f cis the center frequency of the radio frequency signal, such as 2.4 GHz or 5.8 GHz. LPF represents a low-pass filter, which is used to filter out the high-frequency components generated by multiplication to retain the baseband part. The real part x I (t) is the in-phase component, and the imaginary part x Q (t) is the quadrature component. In the specific hardware implementation, the present invention can mix the input signal with cos(2πf c t) and sin(2πf c t) respectively, and then obtain the following forms through the LPF respectively:

[0136]

[0137] Among them, and represent the in-phase baseband signal and the quadrature baseband signal respectively, and are combined into a complex form Through this step, an initial radio frequency signal sample library containing 10 types of unmanned aerial vehicles can be obtained, laying a foundation for subsequent feature extraction and model training, and can also reduce the hardware complexity during deployment and improve the baseband analysis accuracy.

[0138] Such as Figure 3 shown, for step S2, the present invention defines an objective function for the balance relationship between the time window length and the recognition accuracy and the inference time, as shown in formula (3):

[0139] J(w) = α·Acc(w) - β·Time(w) (3)

[0140] Among them, Acc(w) represents the recognition accuracy obtained by segmenting the signal with the time window length w, Time(w) represents the corresponding inference time, and α and β are weight coefficients. In order to select the best value that can balance high accuracy and short inference time among multiple candidate window lengths, the present invention combines the elbow method to analyze the trend curves of Acc(w) and Time(w) changing with w, searches for the "inflection point" position as the key observation range, and searches within the given candidate set of time window lengths according to the formula to select the w that maximizes J(w * , denoted as:

[0141]

[0142] After determining w * , perform a moving sliding window interception on the radio frequency signal x BB (t) obtained by intermediate frequency processing. If the total duration of the signal is T and the sliding window moving step size is Δ, several short-time radio frequency signal sequences x k (t) can be obtained. The specific segmentation method can be seen in formula (5):

[0143] x k x(t) = x BB (t+(k - 1)Δ), k = 1, 2, …, K (5)

[0144] Wherein, Through the above steps, 10 groups of short - time radio frequency signal sequence samples can be obtained while taking into account the recognition accuracy and real - time performance. This process can effectively segment long - time continuous signals and ensure that representative local features can still be extracted in an interference environment.

[0145] For step S3, in order to enhance the robustness of the model in different channel states and interference environments, the present invention adds Gaussian white noise n(t) to the short - time radio frequency signal sequence x k (t) to obtain a new signal x' k (t), as shown in formula (6):

[0146] x' k (t) = x k (t)+n(t) (6)

[0147] Wherein, n(t) can be set as Gaussian white noise with zero mean and variance of σ 2 and σ 2 is selected according to the target signal - to - noise ratio. Then, the obtained signal x' k (t) is subjected to short - time Fourier transform, and specifically, reference can be made to formula (7):

[0148]

[0149] Wherein, φ(·) is a window function, τ is the time - frame index, and f is the frequency index. The obtained complex - valued time - frequency distribution is converted into an amplitude spectrum to obtain a two - dimensional time - frequency matrix, and then this matrix is represented as a time - frequency diagram of size H×W. Finally, a time - frequency diagram dataset containing 10 types of drones is constructed. The time - frequency diagram obtained in this step provides richer feature support for subsequent recognition.

[0150] For step S4, after generating time - frequency diagrams with different durations and signal - to - noise ratios, the present invention uses them for training an improved deep - learning model in this step to achieve high - precision classification and recognition of target categories. This model deeply integrates the Mamba module on the basis of the classic ResNet framework and obtains stronger identification ability for the radio frequency signal characteristics of drones through deep fusion in both spatial and temporal dimensions. The specific process is as follows:

[0151] S41. First, the present invention divides the time-frequency diagrams generated with different durations and different signal-to-noise ratios in the aforementioned S3 into a training set and a validation set. The training set is used to iteratively update the model weights, while the validation set is used to monitor the generalization ability of the model in actual application scenarios, avoiding the decrease in accuracy caused by overfitting. If there are some specific samples in a multi-source interference environment such as a city, data augmentation or balancing can also be performed at this stage.

[0152] S42. In the conventional ResNet path, the present invention deeply embeds the Mamba module. Hierarchical transformation is performed on the network convolution path, residual connection, and non-linear activation, enabling the Mamba module to fully absorb the time series information from the row vectors of each frequency point while retaining the high-order spatial feature extraction ability in the ResNet block. Specifically: Multiple-level MaRBlock is used to perform local and global feature extraction for the two-dimensional characteristics of the time-frequency diagram in the convolution path of the residual structure; at the output of each layer of the residual, each row of the time-frequency diagram is regarded as a one-dimensional sequence and placed inside the Mamba module. Through modules such as linear projection, SiLU non-linear mapping, and state space model (SSM), the feature dependencies of each frequency point changing over time are extracted, and finally, they are fused with the residual backbone. This design not only focuses on capturing the image patterns in the two-dimensional space (frequency domain) but also can deeply explore the dynamic behavior of the one-dimensional time series, making the network more robust and expressive in the identification of UAV radio frequency signals under complex interference.

[0153] The deep learning model adopted by the present invention is as Figure 4As shown. The input of the model is the short-time Fourier transform image of the preprocessed UAV radio frequency signal, with a size of B×C×H×W, where B represents the batch size, C is the number of channels, and H and W are the height and width of the image respectively. In practice, C may be 1 (grayscale image) or 3 (pseudo-color image). The model first undergoes preliminary processing through the Stem module, which includes a 3×3 convolutional layer (Conv3×3), followed by batch normalization (BN) and the ReLU activation function, and then through max pooling (MaxPool) operation, aiming to extract preliminary features and reduce the computational burden. Then, the signal is processed through the stacking of multiple MaR Blocks. Each layer of the MaR Block includes multiple convolutional operations and is combined with downsampling operations (if necessary). Each layer of the MaR Block uses convolutional kernels to extract features and enhances the non-linear representation through the ReLU activation function, finally generating an output highly correlated with the time-frequency features of the input signal. This hierarchical design adopts a stacking method of (3, 4, 6, 3) to ensure that multi-scale semantic information is maintained during the feature extraction process. After each layer of the MaR Block, the signal enters the Mamba module, which further enhances the performance of the model. The Mamba module performs dimensionality reduction or dimensionality increase operations through linear projection, then extracts non-linear local features through convolution and the SiLU activation function, and then conducts global spatio-temporal dependence modeling through the SSM (state space model). Finally, the output of the Mamba module is added to the residual connection to ensure that the transmission of feature information is not limited by the network depth. At the end of the network, adaptive average pooling (AdaptiveAvgPool) is used to unify the size of the feature map and flatten it (Flatten) into a one-dimensional vector, and finally through the fully connected layer (FC) and the Softmax function, the final classification result, that is, the recognition result of the UAV, is obtained.

[0154] S43. In the training stage, the present invention takes the cross-entropy loss function (Cross Entropy) as the main optimization objective, and updates the model parameters through backpropagation and stochastic gradient descent (SGD) or other optimization algorithms (such as Adam). The cross-entropy loss function is shown in formula (8):

[0155]

[0156] where N batch represents the batch size, c is the number of UAV categories, y i,c is the true label of the i-th training sample in terms of category, and is the predicted probability output by the model.

[0157] S44. During the training process, metrics such as accuracy and recall of the model are monitored through the validation set. When the training converges, the present invention can obtain a deep learning model with strong discrimination ability for multi-category drones' RF signals to be recognized.

[0158] For step S5, after the above training is completed, the present invention conducts an inference and recognition process on the real-time or offline collected drones' RF signals. The core lies in: not only extracting spatial features along the ResNet backbone, but also obtaining time-domain dependency information through the Mamba path after row-segmenting the time-frequency map, so as to achieve fast and high-precision determination of the drone type in the application scenario. Specifically, it includes:

[0159] S51. The present invention adopts the same preprocessing, mixing frequency, and short-time Fourier transform processes for the target signal to be recognized as in the training stage to convert the original signal into a time-frequency map. This step can ensure that the size and resolution of the time-frequency map are consistent with the samples used during training, thus avoiding model adaptation problems caused by mismatches in input dimensions or resolutions.

[0160] S52. Input the time-frequency map into the trained MaRNet model. The network will sequentially pass through Stem, multiple MaRBlocks, and send the row vector to the Mamba module at the residual connection. This deeply coupled structure enables the network to take into account both frequency-domain and time-domain information, thereby giving a more reliable distribution estimate. Finally, the prediction scores of each category are output through a fully connected layer, and the Softmax activation function is used to convert them into a probability distribution, and the output is a vector of dimension c where each element represents the probability of the corresponding category.

[0161] S53. In the inference stage, the present invention can introduce a weighting strategy or other post-processing mechanisms according to application requirements. For example, if the probability distribution is too dispersed or the highest probability is lower than the preset threshold, an auxiliary detection process (such as multi-modal verification or additional sensor information fusion) can be triggered to reduce the risk of misrecognition or missed detection. In addition, if combined with an incremental learning framework, for signals suspected of being new types or unknown targets, data recording and annotation can be performed, and the recognition range of the model can be continuously expanded in subsequent training.

[0162] S54. The present invention can interface the final recognition result with the urban low-altitude monitoring platform or other warning systems for real-time monitoring of potentially illegal or dangerous targets. Once an unknown or abnormal category drone is recognized, a warning signal can be immediately sent out to notify the relevant security departments to implement intervention or tracking. If a target with insufficient confidence is detected, it can be incorporated into the data collection and annotation link, and subsequent incremental training or online fine-tuning can be carried out to achieve continuous optimization and update of the model.

[0163] Such as Figure 5The present invention provides a detection and recognition system based on UAV radio frequency signals, including:

[0164] A signal preprocessing module 10, configured to collect radio frequency signals of a micro UAV, convert the radio frequency signals into zero intermediate frequency signals through frequency conversion processing, and obtain an initial radio frequency signal sample library;

[0165] A moving sliding window intercepting module 20, configured to consider a time window with an adaptive duration and the balance relationship between the recognition accuracy rate and the inference time, construct an objective function for the time window length, and combine with the elbow method to obtain a given set of candidate time window lengths, and perform moving sliding window interception on the original radio frequency signal in the initial radio frequency signal sample library to obtain short-time radio frequency signal sequence samples;

[0166] A Gaussian white noise simulation module 30, configured to add Gaussian white noise to the short-time radio frequency signal sequence samples, and generate a time-frequency map dataset with different durations and different signal-to-noise ratios for model training after performing short-time Fourier transform;

[0167] A recognition model module 40, configured to improve and complete the training of a deep learning model for classifying UAV time-frequency maps according to the characteristics of time-frequency maps with different durations and different signal-to-noise ratios, and obtain an intelligent recognition model for UAV time-frequency maps;

[0168] An output module 50, configured to input the time-frequency map of the signal to be recognized into the trained intelligent recognition model based on the intelligent recognition model, and output the classification result of the UAV.

[0169] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 6 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store static information and dynamic information data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0170] Those skilled in the art can understand that Figure 6 the structure shown in merely represents a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0171] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiment are implemented.

[0172] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiment are implemented.

[0173] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0174] The present invention is not limited to the structures that have been described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A detection and identification method based on UAV radio frequency signals, characterized in that: The steps include: S1: Collect the RF signal of the drone, convert the RF signal into a zero intermediate frequency signal through frequency conversion processing, and obtain an initial RF signal sample library; S2: Considering the time window with adaptive length and the balance between recognition accuracy and inference time, an objective function of the time window length is constructed. Combined with the elbow method, a candidate set of a given time window length is obtained. The original RF signal is intercepted by a moving sliding window in the initial RF signal sample library to obtain a short-time RF signal sequence sample. S3: Add Gaussian white noise to the short-time RF signal sequence samples, and perform short-time Fourier transform to generate time-frequency graph datasets with different durations and different signal-to-noise ratios for model training; S4: According to the characteristics of the time-frequency graphs under different durations and signal-to-noise ratios, the deep learning model for the classification of drone time-frequency graphs is improved and trained to obtain an intelligent recognition model for drone time-frequency graphs; S5: Based on the intelligent recognition model, the time-frequency diagram of the signal to be recognized is input into the trained intelligent recognition model, and the classification result of the drone is output.

2. The method for detecting and identifying UAV radio frequency signals according to claim 1, characterized in that: The step of collecting the radio frequency signal of the drone in step S1, converting the radio frequency signal into a zero intermediate frequency signal by frequency conversion processing, and obtaining an initial radio frequency signal sample library specifically includes: Among them, f c is the center frequency of the RF signal, LPF represents a low-pass filter, t is a time variable, x BB (t) is the baseband signal obtained after down-conversion operation, j is an imaginary unit, and j 2 =-1, real part x I (t) is the in-phase component, the imaginary part x Q (t) is the orthogonal component; In the specific hardware implementation, the input RF signal is respectively compared with cos(2πf c t) and sin(2πf c t) After mixing, the following forms are obtained by passing through LPF respectively: in, and Represent the in-phase baseband signal and the orthogonal baseband signal respectively, and then combine them into a complex form 3. The detection and identification method based on the radio frequency signal of the unmanned aerial vehicle according to claim 1 is characterized in that: Step S2: Considering the time window of adaptive duration and the balance between recognition accuracy and inference time, construct an objective function of the time window length, combine the elbow method, obtain a given set of candidate time window lengths, perform moving sliding window interception on the original RF signal in the initial RF signal sample library, and obtain a short-time RF signal sequence sample, specifically including: J(w)=α·Acc(w)-β·Time(w) (3) Wherein, Acc(w) represents the recognition accuracy obtained after segmenting the signal with a time window length w, Time(w) represents the corresponding reasoning time, and α and β are weight coefficients; Combined with the elbow method, the trend curve of Acc(w) and Time(w) changing with w is analyzed to find the inflection point as the key observation range. According to formula (3), the search is performed within the candidate set of the given time window length to select the w that maximizes J(w). * , recorded as: Among them, ω∈W means that ω is taken from a given time window length candidate set W. When determining w * After that, the RF signal x is processed by the intermediate frequency BB (t) Perform moving sliding window capture; if the total signal duration is T, and the sliding window moving step is Δ, then several short-time RF signal sequences x can be obtained. k (t), the specific segmentation method can be found in formula (5): x k (t)=x BB (t+(k-1)Δ),k=1,2,…,K (5) in, 4. The method for detecting and identifying UAV radio frequency signals according to claim 1, characterized in that: Step S3: Add Gaussian white noise to the short-time RF signal sequence samples, and perform short-time Fourier transform to generate time-frequency graph data sets with different durations and different signal-to-noise ratios for model training, specifically including: In the short time RF signal sequence x k (t) and add Gaussian white noise n(t) to get the new signal x' k (t), as shown in formula (6): x' k (t)=x k (t)+n(t) (6) Where n(t) is set to zero mean and variance σ 2 Gaussian white noise, σ 2 Select according to the target signal-to-noise ratio, and then, the obtained signal x' k (t) performs short-time Fourier transform, as shown in formula (7): Among them, φ(·) is the window function, τ is the time frame index, f is the frequency index, m is the discrete time index, which is used to represent the sampling value of the signal at different times, and x' k (m) is the new signal x' obtained after adding Gaussian white noise k The value of (t) at the discrete time point m is converted into an amplitude spectrum to obtain a two-dimensional time-frequency matrix, which is then represented as a time-frequency graph of size H×W to construct a time-frequency graph dataset.

5. The method for detecting and identifying UAV radio frequency signals according to claim 1, characterized in that: Step S4: According to the characteristics of the time-frequency graphs under different durations and different signal-to-noise ratios, the deep learning model for the classification of drone time-frequency graphs is improved and trained to obtain an intelligent recognition model for drone time-frequency graphs, which specifically includes: S41: dividing the time-frequency graphs generated in step S3 with different durations and different signal-to-noise ratios into a training set for iteratively updating the model weights and a validation set for monitoring the generalization ability of the model in actual application scenarios to avoid a decrease in accuracy due to overfitting; S42: Use multi-level MaR Block to extract local and global features from the two-dimensional characteristics of the time-frequency graph in the convolution path of the residual structure; at the output of each layer of the residual, treat each row of the time-frequency graph as a one-dimensional sequence and place it inside the Mamba module. Through linear projection, SiLU nonlinear mapping and state space model SSM, extract the feature dependency of each frequency point over time, and finally merge it with the residual trunk; S43: In the training phase, the cross entropy loss function is used as the main optimization target, and the model parameters are updated through back propagation and stochastic gradient descent (SGD), where the cross entropy loss function is shown in formula (8): Among them, L CE Represents the cross entropy loss value, which is used to quantify the gap between the prediction effect of the model on the current batch of data and the actual situation. N batch represents the batch size, c is the number of drone categories, and y i,c is the true label of the i-th training sample in the category, is the predicted probability of the model output; S44: During the training process, the precision and recall of the model are monitored through the validation set.

6. The method for detecting and identifying UAV radio frequency signals according to claim 1, characterized in that: Step S5: Based on the intelligent recognition model, the time-frequency diagram of the signal to be recognized is input into the trained intelligent recognition model, and the classification result of the drone is output, which specifically includes: S51: The target signal to be identified is subjected to the same preprocessing, mixing and short-time Fourier transform process as in the training stage to convert the original signal into a time-frequency diagram; S52: Input the time-frequency graph preprocessed by S51 into the first convolution layer of the trained MaRNet model; S53: In the reasoning stage, a weighted strategy is introduced; S54: The final recognition result is connected to the urban low-altitude monitoring platform for real-time monitoring of possible illegal or dangerous targets. Once an unknown or abnormal category of drone is identified, a warning signal is immediately issued. If a target with insufficient confidence is detected, it is included in the data collection and labeling process, and incremental training or online fine-tuning is performed later.

7. The method for detecting and identifying UAV radio frequency signals according to claim 6, characterized in that: Step S52: Input the time-frequency graph preprocessed by S51 into the first convolution layer of the trained MaRNet model, wherein the MaRNet model specifically includes: First, through the initial module named "Stem", 3×3 convolution, batch normalization, ReLU and maximum pooling (MaxPool) are used to extract and compress the spatial and frequency initial features of the image; Then, it enters multiple MaR Blocks in sequence. Each MaR Block integrates an additional Mamba module in the residual structure to capture local features through convolution while using a learnable state-space modeling mechanism to obtain deeper time-frequency domain dependency information. Finally, the network end flattens the accumulated high-level features and inputs them into the fully connected layer, and the output is converted into a probability vector of length c through the Softmax activation function. Each element represents the predicted probability of the corresponding drone category.

8. The detection and identification system based on UAV radio frequency signals according to claim 1 is characterized in that: include: The signal preprocessing module is used to collect the radio frequency signal of the micro-UAV, convert the radio frequency signal into a zero intermediate frequency signal through frequency conversion processing, and obtain an initial radio frequency signal sample library; The moving sliding window interception module is used to consider the time window of adaptive duration and the balance between recognition accuracy and inference time, construct the objective function of the time window length, and combine the elbow method to obtain a given set of candidate time window lengths. The moving sliding window interception of the original RF signal in the initial RF signal sample library is performed to obtain a short-time RF signal sequence sample; Gaussian white noise simulation module, used to add Gaussian white noise to short-time RF signal sequence samples, and generate time-frequency graph data sets with different durations and different signal-to-noise ratios for model training after short-time Fourier transform; The recognition model module is used to improve and train the deep learning model for drone time-frequency image classification based on the characteristics of time-frequency images under different durations and signal-to-noise ratios, and obtain an intelligent recognition model for drone time-frequency images; The output module is used to input the time-frequency diagram of the signal to be identified into the trained intelligent recognition model based on the intelligent recognition model, and output the classification result of the drone.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Specific unmanned aerial vehicle model rapid identification method and system based on radio frequency fingerprint database

    CN120408226A

  • Unmanned aerial vehicle identification method and system, electronic equipment and storage medium

    CN121010908A

  • Radio spectrum monitoring method and system of low-altitude defense system

    CN121069017A

  • Full-band unmanned aerial vehicle identification control system and method based on deep learning

    CN121122083A

  • A deep learning-based full-band unmanned aerial vehicle identification regulation system and method

    CN121122083B