Radio frequency fingerprint identification method and system based on double-path cross attention fusion network

Through the combination of Wigner-Ville distribution and dual-channel cross-attention fusion network, the problem of low accuracy of RF fingerprint recognition in low signal-to-noise ratio environment is solved, and high recognition accuracy is achieved under complex noise conditions.

CN120302298APending Publication Date: 2025-07-11SOUTHEAST UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510662311.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The recognition accuracy of RF fingerprint recognition in the existing technology is limited in the low signal-to-noise ratio environment, and it is difficult to effectively improve.

Method used

The Wigner-Ville distribution theory and deep learning combined with the dual-channel cross attention fusion network is used to construct a DAFFM-RFF network for identification by converting the radio frequency signal into a WVD time-frequency diagram and using the dual-channel cross attention module for feature fusion.

Benefits of technology

Maintaining a high recognition accuracy in a low signal-to-noise ratio environment significantly improves the robustness of recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302298A_ABST
    Figure CN120302298A_ABST
Patent Text Reader

Abstract

The invention discloses a radio frequency fingerprint identification method and system based on a double-path cross attention fusion network, and is suitable for an application scene under a low signal-to-noise ratio condition. In an Internet of Things system, prevention of security threats such as illegal monitoring, data tampering and equipment pretending is a basis for guaranteeing user communication security. The radio frequency fingerprint identification technology based on physical layer feature authentication is generated at the right time, has the characteristics of passive authentication, no need of a secret key and the like, and is widely used for enhancing the identity authentication capability of a wireless network. However, in the prior art, the recognition performance is limited in a low signal-to-noise ratio environment, so that how to improve the recognition accuracy in the scene becomes a key problem to be broken through urgently. According to the radio frequency fingerprint identification method based on feature fusion provided by the invention, the one-dimensional radio frequency signal and the two-dimensional time-frequency graph are subjected to multi-modal fusion through the double-path cross attention fusion network, and useful complementary features are further extracted from noise, so that the influence of the noise is reduced, and the identification accuracy of unknown equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of mobile communication technology and wireless network security, and proposes a radio frequency fingerprint recognition method and system based on a dual-path cross-attention fusion network, which is suitable for the identity authentication scenario of Internet of Things devices. Background Art

[0002] In the IoT system, preventing security threats such as illegal monitoring, data tampering, and device impersonation is the basis for ensuring user communication security. To this end, non-password radio frequency fingerprint recognition technology has emerged to enhance the identity authentication capabilities of wireless networks. However, the current technology has limited recognition performance in low signal-to-noise ratio environments, so how to improve the recognition accuracy in such scenarios has become a key problem that needs to be solved urgently. Summary of the invention

[0003] The technical problem to be solved by the present invention is to provide a radio frequency fingerprint recognition strategy that can ensure a higher recognition accuracy in a low signal-to-noise ratio environment.

[0004] To solve the above technical problems, the present invention combines the Wigner-Ville distribution theory, deep learning theory, dual-path cross attention fusion network (DAFFM), and recognition module. It includes the following steps:

[0005] (1) Divide the RF signal dataset s into a training set, a validation set, and a test set;

[0006] (2) For each signal x in the training set and the validation set i Perform Wigner-Ville distribution transformation to obtain the corresponding WVD time-frequency diagram p i ;

[0007] (3) Build a radio frequency fingerprint recognition network based on a dual-path cross attention fusion module, denoted as DAFFM-RFF;

[0008] (4) Each RF signal x in the training set and the validation set i And its corresponding time-frequency diagram p i As input, DAFFM-RFF is trained and validated through deep learning theory to obtain the optimal weight W of the DAFFM-RFF network;

[0009] (5) Testing phase: for each signal y in the test set i Perform Wigner-Ville distribution transformation to obtain the corresponding WVD time-frequency diagram r i ;

[0010] (6) Each RF signal y in the test set i And its corresponding time-frequency diagram ri As input, the DAFFM-RFF network with weight W is input together, and the test samples are identified and predicted using the weight W.

[0011] Preferably, the specific steps of step (2) include:

[0012] (21) Input the I / Q complex signal x i ;

[0013] (22) Obtain the WVD time-frequency diagram p according to the Wigner-Ville distribution transformation formula i , that is

[0014]

[0015] where t represents time, f represents frequency, τ represents time delay, and x i * (t) is the conjugate of x i (t), and e represents the natural base.

[0016] Preferably, the specific steps of step (3) include:

[0017] (31) Build a dual-path cross-attention module, denoted as DMCAM;

[0018] (32) Use the original radio frequency signal X and its corresponding WVD time-frequency diagram P as the two inputs of DMCAM respectively to obtain the output H of DMCAM DMCAM ;

[0019] (33) Build a multi-head self-attention module;

[0020] (34) Use H DMCAM as the input of the query, key, and value vectors in the multi-head self-attention module to obtain the output H of the multi-head self-attention module DAFFM ;

[0021] (35) Build an identification module;

[0022] (36) Input H DAFFM into the identification module to complete radio frequency fingerprint identification.

[0023] Preferably, the specific steps of step (31) include:

[0024] (311) Input the original radio frequency signal X and the WVD time-frequency diagram P;

[0025] (312) Project the original radio frequency signal X and the WVD time-frequency diagram P through multiple linear transformation layers respectively to obtain the corresponding query vector, key vector, and value vector, denoted as Q x , K x , Vx and Q P 、K P 、V P ;

[0026] (313)Construct a multi-head cross-attention module;

[0027] (314)Take Q x and K P 、V P as the inputs of the query, key, and value vectors in the first multi-head cross-attention module respectively, and obtain the output M1 of the first multi-head cross-attention module;

[0028] (315)Construct another multi-head cross-attention module;

[0029] (316)Take Q P and K x 、V x as the inputs of the query, key, and value vectors in the second multi-head cross-attention module respectively, and obtain the output M2 of the second multi-head cross-attention module;

[0030] (317)Add the output results M1 and M2 of the two cross-attention modules to obtain the fused feature representation H after cross-enhancement between the two modalities DMCAM .

[0031] Preferably, the specific steps of step (313) include:

[0032] (3131)Input the query, key, and value vectors, denoted as Q, K, V;

[0033] (3132)Input Q, K, V into the multi-head cross-attention mechanism (Multi-head CrossAttention, MCA) respectively. The MCA module contains 4 parallel attention heads. Each attention head independently calculates the attention weights and outputs the feature subspace representation. Subsequently, all subspace representations are concatenated and linearly mapped to obtain the attention output H MCA ;

[0034] (3133)Add the output H MCA of the MCA module to the input Q to obtain Z1;

[0035] (3134)Perform layer normalization on the residual connection result Z1. This operation calculates the mean μ and standard deviation σ for each sample s i ∈Z1 on its feature dimension and performs a normalization transformation on each feature value, specifically as follows:

[0036]

[0037] where μ is the mean in the feature dimension; σ is the standard deviation; ∈ is a very small constant to avoid division by zero; γ and β are learnable scaling factors and bias terms.

[0038] (3135) Build a feed-forward neural network layer (FFN), which consists of two fully-connected layers with a non-linear activation function in between, specifically expressed as:

[0039] FFN(H LN ) = ReLU(H LN W1 + b1)W2 + b2

[0040] where ReLU(·) represents the non-linear activation function, W1 and W2 represent the weight matrices in the two fully-connected layers respectively, and b1 and b2 represent the biases in the two fully-connected layers respectively.

[0041] (3136) Input H LN into the feed-forward neural network to obtain the output H FFN of the feed-forward neural network;

[0042] (3135) Add H FFN and H LN and send the result into the second normalization layer, and the output is used as the final representation result M1 of the cross-attention module.

[0043] Preferably, the specific steps of step (3132) include:

[0044] (31321) Linearly map the input Q, K, and V into sub-space vector groups Q i , K i , V i in a specific dimension respectively, where i ∈ {1, 2, 3, 4} represents the i-th attention head, and each attention head has an independent projection matrix and

[0045] (31322) In each attention head, use the scaled dot-product attention mechanism to calculate the attention weight matrix A i , and the specific calculation formula is:

[0046]

[0047] where d k is the dimension of the key vector, which is used to scale the dot product to stabilize the gradient;

[0048] (31323) Perform weighted summation of the attention weight matrix A i and the corresponding value vector V i to obtain the output H i of this attention head:

[0049] H i = A i V i

[0050] (31324) Concatenate (Concat) the outputs H1, H2, H3, and H4 of the four attention heads in the feature dimension to form an integrated attention representation vector H concat :

[0051] H concat = Concat(H1, H2, H3, H4)

[0052] (31325) For the concatenated vector H concat Perform a unified dimensionality mapping through a set of linear transformations W O to obtain the final output representation H of this multi-head cross-attention module MCA :

[0053] H MCA = W O H concat

[0054] Preferably, the specific steps of step (33) are the same as those of step (313).

[0055] Preferably, the specific steps of step (35) include:

[0056] (351) Pass through a linear transformation layer to linearly transform the input fused feature vector to adjust the dimension;

[0057] (352) Perform feature normalization on the output of the linear transformation layer through a normalization layer;

[0058] (353) Pass through a ReLU activation function to introduce non-linear transformation ability;

[0059] (354) Pass through a linear transformation layer again to perform a second linear mapping on the ReLU-activated features, for outputting a vector with the same dimension as the number of classification categories.

[0060] (355) Execute the Softmax function on the output result of the linear transformation layer to calculate the probability distribution of each RF signal category for the final classification decision output.

[0061] Preferably, the specific steps of step (4) include:

[0062] (41) In the model training stage, use the Adam optimizer to optimize the parameters and evaluate the model performance with the cross-entropy loss function. This loss function is defined as follows:

[0063]

[0064] Among them, N represents the total number of samples, K represents the total number of categories, and y ik represents the label of the i-th sample. If the i-th sample belongs to category k, then y ik = 1; otherwise, it is 0; represents the predicted probability that the i-th sample belongs to category k;

[0065] (42) Initialize the minimum validation loss Loss min to a maximum value;

[0066] (43) Calculate the current training loss Loss train based on the training set data;

[0067] (44) Backpropagate Loss train to adjust the DAFFM-RFF model parameters;

[0068] (45) Calculate the validation loss Loss val using the validation set data;

[0069] (46) If Loss val < Loss min , then update Loss min = Loss val and save the current model weights;

[0070] (47) Iteratively execute steps (43)-(46) until the validation loss continuously maintains Loss val ≥ Loss min .

[0071] The present invention also provides a radio frequency fingerprint recognition system based on a dual-channel cross-attention fusion network, including:

[0072] (1) A Wigner-Ville transform module for transforming a one-dimensional radio frequency signal into a WVD time-frequency diagram;

[0073] (2) A dual-channel cross-attention fusion network for fusing a one-dimensional radio frequency signal and a WVD time-frequency diagram to obtain more abundant complementary information;

[0074] (3) An identification module for identifying the fused feature map.

[0075] The present invention has the following beneficial effects: By introducing the Wigner-Ville distribution method, the original one-dimensional radio frequency signal is effectively converted into a two-dimensional representation (WVD time-frequency diagram) with joint time-frequency characteristics, and then combined with a dual-channel cross-attention fusion network and an end-to-end recognition module, a new radio frequency fingerprint recognition method suitable for low signal-to-noise ratio environments is constructed. This method can still maintain a high recognition accuracy under the condition of complex noise interference. The simulation experiment and numerical evaluation results show that the recognition framework proposed by the present invention still exhibits excellent robustness in low signal-to-noise ratio scenarios, and has achieved a significant improvement in recognition performance compared with the existing technical means. Brief Description of the Drawings

[0076] Appendix Figure 1 is the network structure diagram of the DMCAM module of the present invention;

[0077] Appendix Figure 2 is the network structure diagram of the recognition module of the present invention;

[0078] Appendix Figure 3 is the network structure schematic diagram of the dual-channel cross-attention fusion network of the present invention. Detailed Embodiment

[0079] The present invention will be further explained below with reference to the embodiments and the drawings.

[0080] Embodiment 1

[0081] In this embodiment, the publicly available ORACLE dataset will be used to verify the present invention. In the simulation, additive white Gaussian noise is selected as the influence brought by the AWGN channel. To reduce the influence of the channel, a dataset with a transmitter-receiver distance of 2 ft is selected for the experiment, and the number of device categories to be recognized is 16.

[0082] The following method is adopted for recognition, including the following steps:

[0083] (1) Consider the construction of the dataset. First, the original ORACLE dataset is randomly divided into a training set, a validation set, and a test set according to a ratio of 8:1:1. For each one-dimensional radio frequency signal in the training set, the validation set, and the test set, the WVD transform is used to generate its corresponding two-dimensional WVD time-frequency diagram.

[0084] (2) Consider the recognition problem of the radio frequency fingerprint recognition strategy under low signal-to-noise ratio. The radio frequency fingerprint features in a low signal-to-noise ratio environment are easily submerged by noise, resulting in incorrect recognition results. Using such as Figure 1The shown dual-path cross-attention network performs multimodal fusion on the original one-dimensional RF signal and the two-dimensional WVD time-frequency diagram, which is used to extract richer complementary information from the noise, thereby reducing the interference of the noise on the RF fingerprint features. Then, a multi-head self-attention module is used to further extract more discriminative feature representations.

[0085] (3) The RF fingerprint recognition problem based on the dual-path cross-attention fusion network. The recognition module shown as Figure 2 is used to recognize the feature map fused in step (1), and the training and verification of the DAFFM-RFF network shown as Figure 3 are completed through deep learning theory, and the cross-entropy loss and accuracy are calculated. Iterate repeatedly until the cross-entropy loss curve converges to obtain the optimal weight W of the DAFFM-RFF network.

[0086] (4) The problem of verifying the RF fingerprint recognition effect based on the dual-path cross-attention fusion network. The test sets under different signal-to-noise ratios are passed through the optimal weight W in step (2) to obtain the output results, so as to obtain the RF fingerprint recognition accuracy of the DAFFM-RFF network under different signal-to-noise ratios, which is used to test the DAFFM-RFF network in a low signal-to-noise ratio environment.

[0087] The RF fingerprint recognition based on the dual-path cross-attention fusion network in step (3) includes the following steps:

[0088] (31) Build an RF fingerprint recognition network based on the dual-path cross-attention fusion network;

[0089] (32) In the model training stage, select Adam as the optimizer, set the batch size to 128, set the learning rate to 0.0001, use the cross-entropy loss as the loss function, and initialize the best loss Loss min to be infinite;

[0090] (33) Calculate the current training loss Loss train based on the training set data;

[0091] (34) Backpropagate Loss train to adjust the DAFFM-RFF model parameters;

[0092] (35) Calculate the validation loss Loss val using the validation set data;

[0093] (36) If Loss val < Loss min , then update Loss min = Loss val and save the current model weights;

[0094] (37) Repeat steps (23)-(26) until Loss has been maintained for 10 consecutive times val ≥Loss min .

[0095] In step (4), for the verification of the RF fingerprint recognition effect based on the dual-path cross-attention fusion network, calculate the recognition accuracy at different signal-to-noise ratios, including the following steps:

[0096] (41) Sequentially input the test sets with different signal-to-noise ratios into the DAFFM-RFF network with weight W;

[0097] (42) Calculate the recognition accuracy of the DAFFM-RFF network on the test sets with different signal-to-noise ratios according to the output and the corresponding labels.

[0098] Embodiment 2

[0099] Based on the above method, this embodiment provides a RF fingerprint recognition system based on a dual-path cross-attention fusion network, including:

[0100] (1) Wigner-Ville transform module: Receive the one-dimensional RF signal input, convert it into a two-dimensional time-frequency diagram through the Wigner-Ville distribution transform formula, and extract the time-frequency joint features of the signal to enhance the feature expression ability in a noisy environment.

[0101] (2) Dual-path cross-attention fusion network (DAFFM): Input the original one-dimensional RF signal and its corresponding WVD time-frequency diagram into the dual-path cross-attention module (DMCAM) respectively, and realize the feature interaction and complementary fusion between the two modalities through the multi-head cross-attention mechanism to generate a fused high-dimensional feature representation. Then, further realize feature extraction through the multi-head self-attention module (MSAM) to extract more discriminative fused features.

[0102] (3) Recognition module: Perform linear transformation, normalization, and non-linear activation processing on the fused features, and finally output the class probability of the RF fingerprint through the Softmax classifier to complete the device identity recognition.

[0103] The system optimizes the network weights through end-to-end training, iteratively updates the parameters using the cross-entropy loss function and the Adam optimizer, and evaluates the recognition accuracy based on the test sets with different signal-to-noise ratios during the verification phase to ensure high robustness even in a low signal-to-noise ratio environment.

[0104] Although the present invention has been illustrated and described with respect to the preferred embodiments, those skilled in the art should understand that various changes and modifications can be made to the present invention as long as they do not exceed the scope defined by the claims of the present invention.

Claims

1. A radio frequency fingerprint recognition method based on a dual-channel cross-attention fusion network, characterized in that It includes the following steps: (1) Divide the radio frequency signal dataset s into a training set, a validation set, and a test set; (2) For each signal x in the training set and the validation set i perform the Wigner-Ville distribution transformation to obtain its corresponding WVD time-frequency diagram p i ; (3) Build a radio frequency fingerprint recognition network DAFFM-RFF based on a dual-path cross-attention fusion module; (4) For each radio frequency signal \(x\) in the training set and the validation set i and its corresponding time-frequency diagram \(p\) i input them into the DAFFM-RFF network for training and validation to obtain the optimal weights \(W\) of the DAFFM-RFF network; (5) For each signal y in the test set i perform the Wigner-Ville distribution transformation to obtain its corresponding WVD time-frequency diagram r i ; (6) For each RF signal y in the test set i and its corresponding time-frequency diagram r i input them into the DAFFM-RFF network with weight W to complete RF fingerprint recognition.

2. The method according to claim 1, wherein The specific steps of step (2) include: (21) Input I / Q complex signal x i ; (22) The WVD time-frequency diagram p is obtained according to the Wigner-Ville distribution transformation formula i , that is where \(t\) represents time, \(f\) represents frequency, \(\tau\) represents time delay, and \(x^*\) i (t) is the conjugate of \(x\) i (t), and \(e\) represents the natural base.

3. The method according to claim 1, wherein The specific steps of step (3) include: (31) Build a dual-path cross-attention module DMCAM; use the original RF signal X and its corresponding WVD time-frequency diagram P as the two inputs of DMCAM respectively to obtain the output H of DMCAM DMCAM ; (32) Build a multi-head self-attention module; Use H DMCAM as the input of the query, key, and value vectors in the multi-head self-attention module to obtain the output H DAFFM ; (33)Build an identification module; perform linear transformation, normalization, activation function processing, and Softmax classification on H DAFFM ​ 4. The method according to claim 3, wherein The implementation of the DMCAM module in step (31) includes: (311) Project the original RF signal X and the WVD time-frequency diagram P through multiple linear transformation layers respectively to obtain the corresponding query vectors Q x and Q P , key vectors K x and K P , and value vectors V x and V P ; (312)Construct a multi-head cross-attention module; use Q x and K P and V P as the inputs of the query, key, and value vectors in the first multi-head cross-attention module respectively, and obtain the output M1 of the first multi-head cross-attention module; (313) Reconstruct a multi-head cross-attention module again; take Q P and K x , V x as the inputs of the query, key, and value vectors in the second multi-head cross-attention module respectively, and obtain the output M2 of the second multi-head cross-attention module; (314) Add the output results M1 and M2 of the two cross-attention modules to obtain the fused feature representation H after cross-enhancement between the two modalities. DMCAM .

5. The method according to claim 4, characterized in that, The implementation of the first multi-head cross-attention module and the second multi-head cross-attention module includes: (321) The input Q, K, and V are respectively mapped into subspace vectors Q, K, and V of four attention heads through independent projection matrices, where i ∈ {1, 2, 3, 4}; i , K i , V i , where i ∈ {1, 2, 3, 4}; (322) In each attention head, the attention weight matrix is calculated through scaled dot product where d k is the dimension of the key vector; (323) Sum A i and the corresponding V i with weighted summation to obtain the output H of each attention head i ; (324) Concatenate the outputs H1, H2, H3, H4 of all attention heads, and perform a linear mapping W O to obtain the multi-head cross-attention output H MCA .

6. The method according to claim 5, wherein After step (324), it also includes: (325) Connect H MCA with the input Q through a residual connection to obtain Z1; (326) Perform layer normalization on Z1, specifically as follows: where μ is the mean on the feature dimension; σ is the standard deviation; ∈ is a very small constant to avoid division by zero errors; γ, β are learnable scaling factors and bias terms; (327) Input the normalization result into the feedforward neural network layer, and output H FFN ; (328) After connecting H FFN with H LN and performing residual connection and then normalization again, it is used as the final output of the cross-attention module.

7. The method according to claim 6, wherein The implementation of the feed-forward neural network layer is: FFN(H LN ) = ReLU(H LN W1 + b1)W2 + b2 where ReLU(·) represents a non-linear activation function, W1 and W2 respectively represent the weight matrices in two fully connected layers, and b1 and b2 respectively represent the biases in two fully connected layers.

8. The method according to claim 3, characterized in that The implementation of the recognition module in step (33) includes: (331) Adjust the dimension of the fused feature vector through a linear transformation layer; (332) Perform layer normalization and ReLU activation on the output; (333) Map it to the dimension of the number of classification categories again through a linear transformation layer; (334) Apply the Softmax function to calculate the class probability distribution and complete the classification decision.

9. The method according to claim 1, characterized in that The specific steps of step (4) include: (41) In the model training stage, use the Adam optimizer to optimize the parameters and evaluate the model performance with the cross-entropy loss function, and this loss function is defined as follows: Among them, N represents the total number of samples, K represents the total number of categories, and y ik represents the label of the i-th sample. If the i-th sample belongs to category k, then y ik = 1; otherwise it is 0. represents the predicted probability that the i-th sample belongs to category k; (42) Initialize the minimum validation loss min to a maximum value; (43) Calculate the current training loss Loss based on the training set data train ; (44) Backpropagate Loss train to adjust the DAFFM-RFF model parameters; (45) Calculate the validation loss using the validation set data val ; (46) If Loss val <Loss min , then update Loss min = Loss val and save the current model weights; (47) Iteratively execute steps (43)-(46) until the validation loss has remained Loss val ≥Loss min for multiple consecutive times.

10. A radio frequency fingerprint recognition system based on the method according to any one of claims 1-9, characterized in that It includes: (1) A Wigner-Ville transform module, which is used to transform a one-dimensional radio frequency signal into a WVD time-frequency diagram; (2) A DAFFM multi-modal fusion network module, which is used to perform multi-modal fusion on a one-dimensional radio frequency signal and a WVD time-frequency diagram to obtain richer complementary information; (3) A recognition module, which recognizes the fused feature map.

Citation Information

Cited By

  • Internet of vehicles broadcast frame radio frequency fingerprint identification method based on LMMSE channel estimation

    CN120825713A

  • Multi-modal radio frequency authentication method based on multi-scale signal representation

    CN121418820A