Radio frequency fingerprint identification method based on image transmission signal of unmanned aerial vehicle

By employing a dual-modal DRF-Net radio frequency fingerprinting scheme, which combines feature extraction and fusion of time-frequency graphs and cyclic spectrum graphs, the problem of low recognition rate of UAVs in low signal-to-noise ratio environments is solved, achieving high-precision UAV model identification.

CN121542889APending Publication Date: 2026-02-17HANGZHOU DIANZI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511651831.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing drone radio frequency fingerprinting technology suffers from low recognition rate and poor robustness in low signal-to-noise ratio environments, making it difficult to effectively address recognition challenges in complex electromagnetic environments.

Method used

A dual-modal DRF-Net radio frequency fingerprinting scheme based on time-frequency maps and cyclic spectrum is adopted. Through a dual-modal feature fusion architecture, the time-frequency features and cyclic stationary features of the signal are extracted respectively. CNN with deformable convolution and Vision Transformer are used for feature extraction. The multi-modal feature fusion module performs adaptive weighting and outputs the probability distribution of the UAV model.

Benefits of technology

The method significantly improves the recognition accuracy of UAVs in low signal-to-noise ratio and complex electromagnetic environments. Experimental results show that the recognition accuracy has increased from 49% at 0dB to 98% at 10dB, especially in the low signal-to-noise ratio region, where it is significantly improved compared to the single-mode method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121542889A_ABST
    Figure CN121542889A_ABST
Patent Text Reader

Abstract

A radio frequency fingerprint identification method based on an image transmission signal of an unmanned aerial vehicle comprises the following steps: firstly identifying the image transmission signal of the unmanned aerial vehicle, and then respectively carrying out short-time Fourier transform and cyclic spectrum analysis to respectively obtain a time-frequency diagram and a cyclic spectrum diagram; respectively inputting the time-frequency graph and the cyclic spectrum graph into a time-frequency feature extraction stream based on CNN and deformation convolution and a cyclic spectrum feature extraction stream based on Vision Transform, and extracting a local time-frequency feature and a global cyclic stationary feature of the signal; fusing the local time-frequency features and the global cyclostationary features through a multi-modal feature fusion module to obtain fused features; and inputting the fused features into a classification head, and outputting probability distribution of each unmanned aerial vehicle model. And the limitation of a traditional single feature recognition method is effectively overcome. According to the technology, the problem of low recognition rate of the low-altitude unmanned aerial vehicle in a complex electromagnetic environment and under the condition of low signal-to-noise ratio can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of radio frequency fingerprint identification, and relates to a radio frequency fingerprint identification method based on a UAV image transmission signal. BACKGROUND

[0002] In recent years, the UAV industry and application have developed rapidly due to the irreplaceable role of UAVs in various industries. However, the "black flight" of UAVs and the frequent occurrence of events involving dangerous goods carried by UAVs pose a serious threat to social safety. Therefore, the detection and identification of UAVs have become particularly urgent and necessary.

[0003] The research on the detection and identification of UAVs is carried out in different ways, such as radar, optical images and videos, audio signals, thermal images, and radio frequency signals. Radio frequency fingerprint technology can achieve monitoring and identification of distant UAVs, while traditional visual and sound identification requires close observation and listening and is easily affected by weather and noise. Radio frequency fingerprint technology has stronger concealment, the identification process can be carried out in real time, has faster response speed, ensures the safety and efficiency of tracking and monitoring of UAVs, and the radio signals of UAVs have stronger anti-interference ability and perform more stably in various complex environments. UAVs control flight and image transmission through radio frequency signals, so the captured radio frequency signals of UAVs can be used for identification of UAVs.

[0004] In terms of the current research status of UAV radio frequency fingerprint identification, existing methods mainly include two types: traditional machine learning-based methods and deep learning-based methods. Traditional methods usually rely on manually designed features, such as signal envelope features and high-order statistical features, combined with support vector machines, random forests, and other classifiers for identification. This type of method has certain effect in specific environments, but the feature extraction process is complex and the generalization ability is limited. In recent years, deep learning-based methods mainly use convolutional neural networks to process time-frequency graphs, achieving end-to-end identification. However, most studies only use single modal features, which have insufficient feature representation ability in low signal-to-noise ratio environments, making it difficult to fully extract the essential features of signals and effectively cope with identification challenges in complex electromagnetic environments.

[0005] It needs to be pointed out that the existing interference identification method in the prior art is essentially different from the present application in the technical path. The existing interference identification method faces the mixture of "communication signal + interference", aiming to solve the detection and classification problem of interference signal, using the "separation-identification" paradigm, realizing interference separation through autoencoder and then analyzing the inter-class macro features of interference through time-frequency and cyclic spectrum. The present application, however, faces the pure unmanned aerial vehicle OFDM communication signal, aiming to solve the identity recognition problem of communication transmitting source, using the "deep fusion-discrimination" paradigm, directly extracting and fusing the double-modal features of the target signal. Compared with the existing general multi-modal fusion scheme based on CNN and Vision Transformer, the innovation of the present application lies in the deep customization for the specific scenario of unmanned aerial vehicle radio frequency fingerprint identification: through the introduction of deformation convolution to enhance the geometric deformation modeling ability of time-frequency features, combined with Vision Transformer to excavate the global stationary characteristics of cyclic spectrum, and innovatively using channel and spatial dual attention mechanism to realize fine fusion of cross-modal features, and through the excavation of the intra-class subtle features of cyclic spectrum at the symbol rate to realize individual discrimination. The two have fundamental differences in problem domain, technical path and feature utilization level. The present application realizes the technical leap from macro classification to micro discrimination, effectively solving the problem of insufficient capture of subtle features of radio frequency signal in low signal-to-noise ratio environment for traditional methods.

[0006] In view of the limitations of the existing method, the present application proposes an innovative method based on multi-modal feature fusion, which excavates the deep features of unmanned aerial vehicle image transmission signal (OFDM) such as obvious periodic variation characteristics, and makes a significant breakthrough in feature representation and model robustness. Experimental verification shows that the present application has a significant advantage over single modal method in improving the final recognition accuracy to 98% in 0-10dB low signal-to-noise ratio complex electromagnetic environment, effectively solving the technical problem of low unmanned aerial vehicle recognition accuracy in low signal-to-noise ratio environment. SUMMARY

[0007] The purpose of the present application is to propose a radio frequency fingerprint identification system based on unmanned aerial vehicle image transmission signal in low-altitude complex electromagnetic environment. In traditional radio frequency fingerprint identification technology, different fingerprint features of unmanned aerial vehicle remote control signal are usually extracted. Unmanned aerial vehicle remote control signal can only be collected with sufficient individual information under the condition of high sampling rate, so the cost of collection equipment is high, and the fingerprint feature extraction is easily affected by channel changes, the recognition rate is low, and the robustness is poor. In order to solve this problem, the present application proposes a radio frequency fingerprint identification scheme based on unmanned aerial vehicle image transmission signal. Image transmission signal can improve the recognition rate in low signal-to-noise ratio due to its characteristics of continuous stability, rich features and easy capture.

[0008] The application is based on the study of existing unmanned aerial vehicle radio frequency fingerprint identification technology, and proposes a dual-mode DRF-Net radio frequency fingerprint identification scheme based on time-frequency graph and cyclic spectrum. DRF-Net is specially designed for precise unmanned aerial vehicle model identification in complex radio frequency environment. The model analyzes the time-frequency characteristics and cyclostationary characteristics of the signal through a dual-mode feature fusion architecture, effectively overcoming the limitations of traditional single feature identification methods. This technology can solve the problem of low recognition rate of low-altitude unmanned aerial vehicles in complex electromagnetic environment and low signal-to-noise ratio conditions.

[0009] The technical solution adopted by the application to solve its technical problems comprises the following steps:

[0010] A radio frequency fingerprint identification method based on unmanned aerial vehicle image transmission signal, comprising the following steps:

[0011] Step 1: Perform signal scanning in the 2.4GHz and 5.8GHz frequency bands to obtain radio frequency signals, and use a cyclostationary feature detection method to identify non-frequency hopping communication signals;

[0012] Step 2: Calculate and determine whether the non-frequency hopping signal is an OFDM signal based on the fourth-order cumulant of the non-frequency hopping signal, and if it is determined to be an OFDM signal, the non-frequency hopping signal is used as the image transmission signal of the unmanned aerial vehicle;

[0013] Step 3: Perform short-time Fourier transform and cyclic spectrum analysis on the image transmission signal respectively to obtain time-frequency graph and cyclic spectrum graph respectively;

[0014] Step 4, input the time-frequency graph and the cyclic spectrum graph into the CNN and deformation convolution-based time-frequency feature extraction stream and the Vision Transformer-based cyclic spectrum feature extraction stream respectively to extract local time-frequency features and global cyclostationary features of the signal;

[0015] Step 5, fuse the local time-frequency features and the global cyclostationary features through a multi-modal feature fusion module to obtain fused features, the multi-modal feature fusion module contains channel attention mechanism and spatial attention mechanism for adaptively weighting important feature channels and key time-frequency regions;

[0016] Step 6, input the fused features into a classification head composed of global average pooling, fully connected layer and Dropout layer, and output the probability distribution of each unmanned aerial vehicle model through the Softmax activation function, determine the final unmanned aerial vehicle model classification result according to the output probability distribution, calculate the classification confidence, and complete the radio frequency fingerprint identification of the unmanned aerial vehicle.

[0017] Preferably, in step 1: the step of performing signal scanning to obtain radio frequency signals specifically includes: for the signal received by signal scanning, first calculating the energy of the received signal, and comparing the energy with a preset threshold. If the energy exceeds the preset threshold, then the signal is used as a signal for signal scanning to obtain radio frequency signals.

[0018] The steps of the energy detection and cyclostationary feature detection method are as follows:

[0019] 1. Energy detection

[0020] (1) Let the scanned signal be x(n), n=0,1,…,N-1, where N is the number of samples in one observation.

[0021] (2) Calculate the signal energy:

[0022]

[0023] (3) Collect background noise signals ω(m) from M samples during the no-signal period, m=0,1,…,M-1,

[0024]

[0025] (3) Set the threshold: Th_energy=E_noise*γ, where γ is a coefficient set according to the false alarm probability.

[0026] If E signal If the value of Th_energy is found, it indicates the presence of a signal, and the system proceeds to cyclic stationary feature detection; otherwise, the scanning continues.

[0027] 2. Cyclic stationary feature detection

[0028] (1) Calculate the cyclic autocorrelation function for the signal x1(n) that passes through the energy detector:

[0029]

[0030] Where α is the cycle frequency and τ is the time delay.

[0031] (2) Search for peak values ​​in the cyclic frequency α domain. For non-frequency hopping signals, peak values ​​typically appear at twice the carrier frequency and integer multiples of the symbol rate.

[0032] (3) Set the threshold for the peak value in the cyclic frequency domain. If the peak value exceeds the threshold, it is determined to be a non-frequency hopping signal.

[0033] Preferably, step 2, the determination of whether a signal is an OFDM signal based on the fourth-order cumulant of the non-frequency hopping signal, specifically includes the following steps:

[0034] First, the non-frequency hopping signal is preprocessed to convert it into a zero-mean complex baseband signal x2(n);

[0035] Secondly, calculate the second moment M2 = E(|x2[n]| of the preprocessed signal. 2 ) and the fourth moment M4=E(|x2[n]| 4 );

[0036] Then, calculate the fourth-order cumulant: cum4 = M4 - 2 × (M2) 2 ;

[0037] Finally, a threshold decision is made: if |cum4| < the threshold, then it is determined to be an OFDM signal.

[0038] Preferably, in step 3, the steps for obtaining the time-frequency diagram and the cyclic spectrum are as follows:

[0039] 1. Time-frequency diagram

[0040] The OFDM signal is x3[n], and its discrete Fourier transform is:

[0041]

[0042] Where m is the time frame index, k is the frequency index, N1 is the number of FFT points, and ω[n] is the window function in the discrete case.

[0043] The power spectrum is usually calculated, and the time-frequency plot shows a logarithmic scale. The formula is:

[0044] L(m,k)=10log 10 (|STFT(m,k)| 2 +ε)

[0045] Where ε is a very small positive number.

[0046] 2. Cyclic spectrum

[0047] Cyclic spectral density can be calculated using frequency domain smoothing methods. The cyclic autocorrelation function of the OFDM signal x3[n] is defined as...

[0048]

[0049] Where m is the discrete time delay and α is the cycle frequency.

[0050] The cyclic spectral density is the Fourier transform of the cyclic autocorrelation function with respect to the time delay m:

[0051]

[0052] Where f represents the frequency spectrum.

[0053] Preferably, in step 4, the recurrent spectral feature extraction stream based on Vision Transformer and the time-frequency feature extraction stream based on CNN and deformable convolution specifically include:

[0054] The cyclic spectral feature extraction stream based on Vision Transformer receives the cyclic spectrogram input, divides the input image into fixed-size blocks through the image segmentation module, and after linear projection and position encoding, it is input together with the CLS classification token into the Transformer encoder containing 12 encoding layers. Each encoding layer contains a multi-head self-attention module, a layer normalization module and a feedforward neural network. Finally, the cyclic spectral feature vector F_cs is output from the CLS token to capture global cyclic stationary features.

[0055] The time-frequency feature extraction stream based on CNN and deformable convolution receives a time-frequency map as input and extracts features sequentially through three convolutional blocks. The first convolutional block uses a 7×3 convolutional kernel and connects batch normalization, ReLU activation function and 3×3 max pooling layer. The second convolutional block uses a 3×3 convolutional kernel to output a 64-channel feature map. The third convolutional block uses a 3×3 convolutional kernel to output a 128-channel feature map. Then, the receptive field is adaptively adjusted through a deformable convolutional module to capture irregular feature patterns. Finally, global average pooling is used to obtain the time-frequency feature vector F_tf to represent local time-frequency features.

[0056] Preferably, the multimodal feature fusion module in step 5 includes:

[0057] First, the time-frequency feature vector F_tf and the cyclic spectrum feature vector F_cs are concatenated through a feature concatenation operation. Then, the concatenated features are weighted according to their importance through spatial attention and channel attention modules. The spatial attention module is used to highlight key spatial regions in the feature map, while the channel attention module is used to recalibrate the weight distribution of each feature channel. Finally, the feature weighted fusion layer achieves adaptive fusion of the two modal features, thereby generating a more discriminative fused feature representation for the final classification decision.

[0058] The beneficial effects of this invention are as follows: This invention proposes a UAV radio frequency fingerprinting method based on dual-modal features of time-frequency diagrams and cyclic spectrum diagrams, building upon traditional UAV identification technology based on raw radio frequency signals. This method acquires and detects UAV image transmission signals in the 2.4GHz and 5.8GHz frequency bands, accurately identifies OFDM modulated signals using fourth-order cumulant features, and generates dual-modal feature representations of time-frequency diagrams and cyclic spectrum diagrams, respectively.

[0059] Preferably, this invention employs a dual-stream network architecture consisting of a time-frequency feature extraction stream based on CNN and deformable convolution, and a cyclic spectrum feature extraction stream based on VisionTransformer. This extracts local time-frequency features and global cyclic stationary features from the time-frequency map and cyclic spectrum map, respectively. Through a multimodal feature fusion module, the time-frequency feature vector F_tf and the cyclic spectrum feature vector F_cs are first concatenated. Then, the key feature regions and important feature channels are adaptively weighted through a spatial attention module and a channel attention module. Finally, feature weighting fusion and a classification head are used to effectively enhance the feature discrimination capability and classification accuracy of UAV RF fingerprints.

[0060] This invention significantly improves the recognition performance and adaptability of UAV identification systems in complex electromagnetic environments, especially under multipath effects and low signal-to-noise ratio conditions. Through in-depth mining and intelligent fusion of dual-modal features, it achieves high-precision identification of low-altitude UAVs. Experimental results show that the method described in this invention significantly improves the recognition accuracy of low-altitude UAVs on the target test dataset compared to traditional methods, effectively solving the technical challenge of low UAV recognition rate in low signal-to-noise ratio environments. Attached Figure Description

[0061] Figure 1 This is a flowchart of the OFDM signal detection process for image transmission.

[0062] Figure 2 The flowchart shows the bimodal recognition process based on time-frequency plots and cyclic spectra.

[0063] Figure 3 This is a time-frequency diagram based on the short-time Fourier transform.

[0064] Figure 4 This is a cyclic spectrum based on cyclic spectrum analysis;

[0065] Figure 5 This is a diagram of a bimodal recognition network based on time-frequency plots and cyclic spectra;

[0066] Figure 6 For channel attention mechanism module;

[0067] Figure 7 This is a spatial attention mechanism module;

[0068] Figure 8 For classification head model blocks;

[0069] Figure 9 This is a performance comparison chart of drone recognition solutions. Detailed Implementation

[0070] The present invention will now be described in detail with reference to the accompanying drawings.

[0071] The overall process of this invention is as follows: Figure 1 and Figure 2 As shown, the process mainly includes three core stages: signal detection and recognition, feature extraction and fusion, and classification and recognition. First, signal scanning is performed in the 2.4GHz and 5.8GHz frequency bands. Valid non-frequency hopping communication signals are identified through energy detection and cyclostationary feature detection. Then, the fourth-order cumulant of the signal is calculated to identify OFDM-modulated UAV image transmission signals. Short-time Fourier transform and cyclic spectrum analysis are performed on the identified image transmission signals to generate time-frequency maps and cyclic spectrum maps. The two types of images are input into a time-frequency feature extraction stream based on CNN and deformable convolution, and a cyclic spectrum feature extraction stream based on VisionTransformer, respectively. The two types of features are fused through a multimodal feature fusion module. Finally, the UAV model identification result is output through a classification head.

[0072] The specific steps will be explained below:

[0073] 1) Signal detection and recognition

[0074] The detailed process for this stage is as follows: Figure 1 As shown, it specifically includes:

[0075] (1) Energy detection

[0076] For the scanned signal x(n), n=0,1,…,N-1 (where N is the number of samples in a single observation), process it according to the following steps:

[0077] 1. Calculate signal energy:

[0078]

[0079] 2. Collect background noise signals ω(m) from M samples during periods without signal, m = 0, 1, ..., M-1

[0080]

[0081] 3. Set the detection threshold:

[0082] Th_energy=E noise *γ

[0083] Where γ is a coefficient set according to the false alarm probability.

[0084] 4. If E signal If the value of Th_energy is found, it indicates the presence of a signal, and the system proceeds to cyclic stationary feature detection; otherwise, the scanning continues.

[0085] (2) Cyclic stationarity feature detection

[0086] The signal x1(n) that passes through the energy detection is processed as follows:

[0087] 1. Calculate the cyclic autocorrelation function:

[0088]

[0089] Where α is the cycle frequency and τ is the time delay.

[0090] 2. Searching for peak values ​​in the cyclic frequency α domain, non-frequency hopping signals will show obvious peak values ​​at twice the carrier frequency and integer multiples of the symbol rate.

[0091] 3. Set a peak threshold; if the peak value exceeds the threshold, it is determined to be a non-frequency hopping signal.

[0092] (3) OFDM signal recognition

[0093] 1. Signal preprocessing: The identified non-frequency hopping signal is converted into a zero-mean complex baseband signal x2(n), where n = 0, 1, ..., N-1, and N is the total number of signal samples.

[0094] 2. Calculation of second moment

[0095] M2=E(|x2[n]| 2 3. Calculation of fourth-order moments

[0096] M4=E(|x2[n]| 4 4. Calculation of fourth-order cumulants

[0097] cum4 = M4 - 2 × (M2) 2

[0098] 5. Threshold Decision Criteria

[0099] According to the theoretical characteristics of fourth-order cumulants, OFDM signals satisfy the condition that if |cum4| < threshold, they are determined to be OFDM signals; otherwise, the scanning continues.

[0100] 2) Generation of time-frequency graphs and cyclic spectra

[0101] (1) Time-frequency graph generation

[0102] Time-frequency graph as follows Figure 3 As shown.

[0103] For the OFDM signal x3[n], the short-time Fourier transform is defined as:

[0104]

[0105] Where m is the time frame index, k is the frequency index, N1 is the number of FFT points, and ω[n] is the window function in the discrete case.

[0106] Specific generation steps:

[0107] Signal framing: The signal is divided into overlapping frames, with a frame length of 512 points and an overlap of 256 points.

[0108] Windowing: Apply Hanning window to each frame.

[0109] FFT calculation: Perform a 4096-point FFT on each frame to obtain the spectrum.

[0110] Matrix construction: Arrange the spectrum in chronological order to form a time-frequency matrix.

[0111] Image generation: Calculate the logarithmic power spectrum L(m,k) = 10log 10 (|STFT(m,k)| 2 +ε), where ε is a very small positive number to avoid pairing zeros. Then L(m,k) is normalized to the range [0,1] and interpolated to a 224×224 three-channel image.

[0112] (2) Generation of Cyclic Spectra

[0113] Cyclic spectrum as follows Figure 4 As shown.

[0114] Cyclic autocorrelation function:

[0115]

[0116] Where m is the discrete time delay and α is the cycle frequency.

[0117] Cyclic spectral density:

[0118]

[0119] Where f represents the frequency spectrum.

[0120] Specific generation steps:

[0121] Signal segmentation: Dividing a signal into multiple overlapping segments.

[0122] Frequency domain calculation: Window each segment and perform FFT to calculate... Among them, X i (f) is x i Fourier transform of [n].

[0123] Averaging: Averaging all segments yields the estimated cyclic spectral density.

[0124] Image generation: Take the absolute value of the result and convert it into a 224×224 three-channel image.

[0125] 3) Dual-stream feature extraction

[0126] The network architecture at this stage is as follows: Figure 5As shown, it includes two parallel feature extraction streams:

[0127] (1) Time-frequency feature extraction flow

[0128] Network architectures based on CNN and deformable convolution include:

[0129] Convolutional block 1: Uses 7×3 convolution, followed by batch normalization and ReLU activation function, and downsampled through a 3×3 max pooling layer.

[0130] Convolutional block 2: Uses 3×3 convolution to output a 64-channel feature map.

[0131] Convolutional block 3: Uses 3×3 convolution to output a 128-channel feature map.

[0132] Deformed Convolution Module: Introduces learnable offsets to enhance the ability to model irregular time-frequency structures.

[0133] Global average pooling: Converts the feature map into a time-frequency feature vector F_tf.

[0134] (2) Cyclic Spectrum Feature Extraction Flow

[0135] The network architecture based on Vision Transformer includes:

[0136] Image segmentation: Dividing the input cyclic spectrum into fixed-size patches.

[0137] Linear projection and position encoding: Mapping tiles to vectors and adding position information.

[0138] CLS Token Addition: Insert a classification token for global feature extraction.

[0139] Transformer encoder: 12-layer structure, each layer includes a multi-head self-attention (MSA) mechanism, layer normalization, and a feedforward network.

[0140] CLS token output: Extract global cyclic stationary features to form a cyclic spectrum feature vector F_cs.

[0141] 4) Multimodal feature fusion

[0142] The fusion module employs a dual-path attention mechanism to effectively fuse the cyclic enhancement vector F_cs from the visual Transformer path with the time-frequency feature vector F_tf from the convolutional expansion path.

[0143] (1) Feature splicing

[0144] The two feature vectors F_tf and F_cs are concatenated (Concat) to form the multimodal feature representation before fusion.

[0145] (2) Channel Attention Module

[0146] Channel attention module model such as Figure 6 As shown.

[0147] The input feature map is subjected to global average pooling (GAP) and global max pooling (GMP) respectively to obtain two pooled feature vectors.

[0148] These two feature vectors are sequentially input into the same shared two-layer MLP (Multilayer Perceptron). The first layer of the MLP compresses the dimension to C / 16 (C is the number of input channels) and uses the ReLU activation function, while the second layer restores the dimension to the original number of channels C.

[0149] The two feature vectors after MLP processing are added element by element.

[0150] The summation result is passed through the Sigmoid activation function to generate channel attention weights with values ​​in the range [0,1].

[0151] Finally, the weights are multiplied channel by channel with the original input feature map, thereby amplifying the feature responses of important channels and suppressing unimportant channels.

[0152] (3) Spatial Attention Module

[0153] Spatial attention module model such as Figure 7 As shown.

[0154] Average pooling and max pooling are performed along the channel dimension to obtain two spatial feature maps.

[0155] The two feature maps are concatenated along the channel dimension.

[0156] Spatial attention weights are generated using 7×7 convolution and the Sigmoid function.

[0157] The weights are multiplied position by position in the original feature map to enhance the features of key regions.

[0158] (4) Feature-weighted fusion

[0159] By combining channel and spatial attention weights, the spliced ​​features are adaptively weighted and fused to form more discriminative multimodal fusion features.

[0160] 5) Classification and Recognition

[0161] Classification head structure such as Figure 8 As shown, it includes:

[0162] Global average pooling: First, global average pooling is performed on the input fused features to convert them into a 1024-dimensional feature vector.

[0163] The first fully connected layer and normalization: 1024-dimensional features are input into a fully connected layer with 512 neurons, and then batch normalization is performed and a non-linear transformation is performed using the ReLU activation function.

[0164] Random deactivation: A dropout layer with a dropout rate of 0.5 is introduced to prevent the model from overfitting.

[0165] The second fully connected layer and normalization: The processed features are further input into a fully connected layer with 256 neurons, and are also processed by batch normalization and ReLU activation function.

[0166] Output Layer and Activation: Finally, the features are fed into the output layer, where the number of neurons corresponds to the number of drone models to be identified, and the Softmax activation function is used to generate the probability distribution for all models.

[0167] Classification decision: During reasoning, the model corresponding to the highest probability value in the probability distribution is taken as the final identification result, and this probability value is also used as the confidence level of this identification.

[0168] Example: Implementation of a UAV Radio Frequency Signal Identification System

[0169] 1. Data Acquisition and Preprocessing

[0170] Signals from 13 mainstream drones were acquired using a real receiver, including {Herelink HX4, DJI FPV COMBO, DJI AVTA2, DJI MAVIC3 PRO, DJI MINI3, DJI MINI4 PRO, YunZhuo H16, Phantom 4 Pro, Mavic Pro, Mini 2, MATRICE 300, AVATA, MATRICE 600Pro}. The drone signals were sampled at 100MHz, operating in the 2.4GHz and 5.8GHz frequency bands.

[0171] 2. Signal Detection and Recognition

[0172] The detection and identification process of drone radio frequency signals is as follows: Figure 1 As shown, a multi-level detection mechanism is employed to ensure accurate capture and identification of the target signal:

[0173] (1) Non-frequency hopping signal detection

[0174] The received 13 types of UAV radio frequency signals were analyzed using a two-stage detection method based on energy detection and cyclostationary characteristics.

[0175] Energy detection: Calculate the energy of the received signal and compare it with a dynamic threshold set based on the statistical characteristics of background noise to preliminarily determine whether the signal exists.

[0176] Cyclic stationary feature detection: Calculate the cyclic autocorrelation function of the signal that passes the energy detection, search for characteristic peaks in the cyclic frequency domain, and identify non-frequency hopping communication signals with cyclic stationary characteristics.

[0177] (2) OFDM signal recognition

[0178] For the identified non-frequency hopping signals, OFDM modulation identification is performed based on higher-order statistical characteristics:

[0179] Signal preprocessing: Converting the non-frequency hopping signal into a zero-mean complex baseband signal.

[0180] Fourth-order cumulant calculation: Calculate the second and fourth moments of the signal, and then calculate the fourth-order cumulant cum4.

[0181] Threshold decision: The fourth-order cumulant characteristics are judged by setting a threshold condition. When |cum4| < threshold, it is determined to be an OFDM modulated UAV image transmission signal.

[0182] 3. Generation of time-frequency feature images

[0183] 4096 points were selected from the detected starting point as a set of signal data points, and short-time Fourier transform and cyclic spectrum analysis were performed respectively:

[0184] Time-frequency graph generation: A Hamming window with a window length of 512 points, an overlap of 256 points, and 4096 FFT points were used to calculate the logarithmic power spectrum and convert it into a 224×224 three-channel image.

[0185] Cyclic spectrum generation: The cyclic spectral density is calculated using a frequency domain smoothing method, and the absolute value of the result is taken and converted into a 224×224 three-channel image.

[0186] Figure 3 and Figure 4 Compare the generated time-frequency plot and cyclic spectrum.

[0187] For each UAV model, 400 time-frequency plots and 400 cyclic spectra were selected as the training set, and 100 time-frequency plots and 100 cyclic spectra were selected as the test set.

[0188] 4. Two-stream feature extraction network

[0189] Figure 5 The overall structure of a drone recognition network based on dual-modal feature fusion is demonstrated. This network comprises two parallel feature extraction paths:

[0190] (1) Expand the feature extraction stream:

[0191] The input image is first processed through a convolutional block (Conv2D).

[0192] Deep feature extraction is then performed using multiple concatenated convolutional blocks (Conv2D++).

[0193] A deformable convolution module is introduced to enhance the model's adaptability to geometric deformation and scale changes.

[0194] Finally, the time-frequency feature vector F_tf is obtained through a global average pooling layer.

[0195] (2) Looping Enhanced Feature Extraction Flow:

[0196] Based on the Vision Transformer architecture, the input image is segmented into a fixed sequence of tiles.

[0197] Add a learnable location code along with an additional CLS token.

[0198] Global context information is modeled using a 12-layer Transformer encoder.

[0199] Finally, the output of the CLS token is used as the global feature representation to obtain the cyclic enhanced feature vector F_cs.

[0200] 5. Multimodal feature fusion and classification

[0201] The feature vectors extracted by the two-stream network are processed through a feature fusion module that includes a dual attention mechanism. The specific process is as follows:

[0202] (1) Feature concatenation: The time-frequency feature vector F_tf is concatenated with the cyclic enhancement feature vector F_cs to form the initial fused features.

[0203] (2) Enhanced attention:

[0204] Channel attention mechanism: Channel attention weights are generated through global average pooling and global max pooling, and the concatenated features are recalibrated in terms of channel dimension to highlight information-rich feature channels.

[0205] Spatial attention mechanism: Pooling operation is performed on the channel dimension to generate spatial attention weights, and features are recalibrated in the spatial dimension to focus on features in key regions.

[0206] (3) Feature-weighted fusion

[0207] The features calibrated by the dual attention mechanism are then weighted and fused.

[0208] (4) Classification and recognition

[0209] The fused features are then subjected to global average pooling to convert them into feature vectors.

[0210] The vector is then refined through a classification head that includes a fully connected layer, batch normalization, ReLU activation function, and Dropout layer.

[0211] Finally, the output layer uses the Softmax activation function to calculate the probability distribution of each drone model, and takes the model with the highest probability as the recognition result.

[0212] Experimental verification

[0213] To verify the performance of this invention, 13 types of UAV signals were tested within a signal-to-noise ratio range of 0–10 dB. The experimental results are as follows: Figure 9 The results show that the recognition accuracy of the method of the present invention steadily improves from 49% at 0dB to 98% at 10dB, demonstrating good signal-to-noise ratio adaptability. At the critical signal-to-noise ratio of 5dB, the recognition accuracy reaches 87%, representing improvements of 4% and 9% compared to the single time-frequency map method (83%) and the single cyclic spectrum method (78%), respectively. Particularly in the low signal-to-noise ratio region (0–5dB), the advantages of the method of the present invention compared to the single-modal method are more significant, with an average accuracy improvement of 6.5%. Experimental results demonstrate that the present invention effectively enhances feature discrimination capability through dual-modal feature fusion, solving the technical problem of low recognition accuracy of UAVs in complex electromagnetic environments.

[0214] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A radio frequency fingerprinting method based on UAV image transmission signals, characterized in that, Includes the following steps: Step 1: Scan the RF signals in the 2.4GHz and 5.8GHz frequency bands and use the cyclic stationary feature detection method to identify non-frequency hopping signals; Step 2: For the identified non-frequency hopping signal, calculate and determine whether it is an OFDM signal based on the fourth-order cumulant of the non-frequency hopping signal. If it is determined to be an OFDM signal, then the non-frequency hopping signal is used as the image transmission signal of the UAV. Step 3: Perform short-time Fourier transform and cyclic spectrum analysis on the image transmission signal to obtain the time-frequency diagram and cyclic spectrum, respectively; Step 4: Input the time-frequency graph and cyclic spectrum graph into the time-frequency feature extraction stream based on CNN and deformable convolution and the cyclic spectrum feature extraction stream based on Vision Transformer, respectively, to extract the local time-frequency features and global cyclic stationary features of the signal; Step 5: The local time-frequency features and global cyclic stationary features are fused through a multimodal feature fusion module to obtain the fused features. The multimodal feature fusion module includes a channel attention mechanism and a spatial attention mechanism, which are used to adaptively weight important feature channels and key time-frequency regions. Step 6: Input the fused features into the classification head, which consists of global average pooling, fully connected layers, and dropout layers. Output the probability distribution of each drone model through the Softmax activation function. Determine the final drone model classification result based on the output probability distribution and calculate the classification confidence to complete the radio frequency fingerprint recognition of the drone.

2. The radio frequency fingerprinting method based on UAV image transmission signals as described in claim 1, characterized in that, In step 1, the step of performing signal scanning to obtain radio frequency signals specifically includes: for the signal received by signal scanning, first calculating the energy of the received signal, and comparing the energy with a preset threshold. If the energy exceeds the preset threshold, then the signal is used as a signal for signal scanning to obtain radio frequency signals.

3. The radio frequency fingerprinting method based on UAV image transmission signals as described in claim 2, characterized in that, In step 1, the The non-frequency hopping communication signal is identified using a cyclic stationary feature detection method, which includes the following steps: For a signal x1(n) that passes energy detection, where n is a non-negative integer, the cyclic autocorrelation function of signal x1(n) is calculated using the following formula: Where α is the cycle frequency and τ is the time delay; The peak value is searched in the cyclic frequency α domain. For non-frequency hopping signals, if a peak value appears cyclically at twice the carrier frequency and an integer multiple of the symbol rate, and the peak value exceeds a preset threshold, it is determined to be a non-frequency hopping signal.

4. The radio frequency fingerprinting method based on UAV image transmission signals as described in claim 1, characterized in that, Step 2, the determination of whether a signal is an OFDM signal based on the fourth-order cumulant of the non-frequency hopping signal, specifically includes the following steps: First, the non-frequency hopping signal is preprocessed to convert it into a zero-mean complex baseband signal x2(n); Secondly, calculate the second moment M2 = E(|x2[n]| of the preprocessed signal. 2 ) and the fourth moment M4=E(|x2[n]| 4 ); where E represents the expected value; Then, calculate the fourth-order cumulant: cum4 = M4 - 2 × (M2) 2 ; Finally, a threshold decision is made: if |cum4| < the threshold, then it is determined to be an OFDM signal.

5. The radio frequency fingerprinting method based on UAV image transmission signals as described in claim 1, characterized in that, Step 3, obtaining the time-frequency diagram and the cyclic spectrum, specifically includes the following steps: For the OFDM signal x3[n], perform a discrete Fourier transform: Where m is the time frame index, k is the frequency index, N1 is the number of FFT points, and ω[n] is the window function in the discrete case; The formula for calculating the time-frequency diagram is: L(m,k)=10log 10 (|STFT(m,k)| 2 +ε) Where ε is a positive number; For the cyclic spectrum: the cyclic spectral density is calculated using a frequency domain smoothing method, and the cyclic autocorrelation function of the OFDM signal x3[n] is defined as... Where m is the discrete time delay and α is the cycle frequency; The cyclic spectral density is the Fourier transform of the cyclic autocorrelation function with respect to the time delay m: Where f represents the frequency spectrum.

6. The radio frequency fingerprinting method based on UAV image transmission signals as described in claim 1, characterized in that, In step 4, The time-frequency feature extraction stream based on the CNN model and deformable convolution receives the time-frequency map as input and extracts features sequentially through three convolutional blocks: the first convolutional block uses a 7×3 convolutional kernel and connects batch normalization, ReLU activation function and 3×3 max pooling layer; the second convolutional block uses a 3×3 convolutional kernel to output a 64-channel feature map; and the third convolutional block uses a 3×3 convolutional kernel to output a 128-channel feature map. After the three convolutional blocks, a deformable convolutional module is set to adaptively adjust the receptive field to capture irregular feature patterns. Finally, global average pooling is used to obtain the time-frequency feature vector F_tf to represent local time-frequency features. The Vision Transformer-based cyclic spectral feature extraction stream receives a cyclic spectrogram input, segments the input image into fixed-size blocks through an image segmentation module, and inputs them along with a CLS classification token into a Transformer encoder containing 12 coding layers. Each coding layer contains a multi-head self-attention module, a layer normalization module, and a feedforward neural network. Finally, the cyclic spectral feature vector F_cs is output from the CLS token to capture global cyclic stationary features.

7. The radio frequency fingerprinting method based on UAV image transmission signals as described in claim 6, characterized in that, In step 5, the processing procedure of the multimodal feature fusion module includes: First, the time-frequency feature vector F_tf and the cyclic spectrum feature vector F_cs are concatenated through a feature concatenation operation. Then, the concatenated features are weighted according to their importance through spatial attention and channel attention modules. The spatial attention module is used to highlight key spatial regions in the feature map, while the channel attention module is used to recalibrate the weight distribution of each feature channel. Finally, the feature weighted fusion layer achieves adaptive fusion of the two modal features, thereby generating a more discriminative fused feature representation for the final classification decision.

8. The radio frequency fingerprinting method based on UAV image transmission signals as described in claim 7, characterized in that, In step 5, The channel attention mechanism performs global average pooling and global max pooling on the input feature map to obtain two pooled feature vectors. These two pooled feature vectors are sequentially input into the same shared two-layer MLP. The first layer of the MLP compresses the dimension to 1 / 16 of the number of input channels and uses the ReLU activation function, while the second layer restores the dimension to the number of input channels. After the two feature vectors processed by MLP are added element by element, channel attention weights are generated by applying the Sigmoid activation function. The channel attention weights are then multiplied with the feature map channel by channel to amplify the feature responses of important channels and suppress unimportant channels.

9. The radio frequency fingerprinting method based on UAV image transmission signals as described in claim 7, characterized in that, In step 5, The spatial attention module is used to model the importance of features in the spatial dimension. The spatial attention module obtains two spatial feature maps by performing average pooling and max pooling operations on the channel dimension respectively. After concatenating these two spatial feature maps, spatial attention weights are generated by convolution and the Sigmoid function. The spatial attention weights are multiplied with the feature maps position by position to enhance the features of key regions.

Citation Information

Cited By

  • Feature extraction and identification method for radio frequency signal of unmanned aerial vehicle

    CN122065131A