A radio frequency fingerprinting method based on sample enhancement and deep learning

By using densely connected convolutional networks and transfer learning in radio frequency fingerprint recognition, the problems of gradient vanishing and long training time of traditional CNN networks are solved, improving the accuracy and robustness of recognition, and making it suitable for radio frequency fingerprint recognition of wireless communication devices.

CN116257750BActive Publication Date: 2025-12-05NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310119656.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-11-09
Filing Date
2023-02-15
Publication Date
2025-12-05
Estimated Expiration
2043-02-15

AI Technical Summary

Technical Problem

Traditional CNN networks suffer from problems such as vanishing gradients, long training time, and high computational cost in radio frequency fingerprint recognition. Furthermore, the difference in the distribution of training data and test data leads to poor model robustness.

Method used

A densely connected convolutional network (DenseNet) algorithm is used for radio frequency fingerprint recognition, and sample augmentation is performed through transfer learning (TL). The DSEN-TL network is constructed, and raw-IQ data is used for feature extraction and classification. The loss function is optimized to improve the model's generalization ability.

Benefits of technology

It alleviates the gradient vanishing problem, reduces training parameters and time, improves the robustness of RF fingerprint recognition, is applicable to data recognition on different dates and with different receivers, and enhances the model's mobility and recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116257750B_ABST
    Figure CN116257750B_ABST
Patent Text Reader

Abstract

A radio frequency fingerprinting method based on sample enhancement and deep learning, S1, receiving and storing raw radio frequency data collected by a wireless communication device to a PC end; S2, preprocessing the Raw-IQ data samples collected by the wireless communication device receiving end, dividing the data into a training set and a test set, and further dividing the training set into a source domain and a target domain; S3, constructing a DSEN-TL neural network and inputting the training set to train it, determining whether the H divergence reaches balance during the training process, if it is too large or too small, the number of layers and parameters of the radio frequency individual identification network and the domain classification network must be adjusted and retrained, and the loss function is optimized; S4, inputting the test set into the DSEN-TL network and outputting the device identification type; S5, embedding the trained DSEN-TL network into an actual board-level system for testing, accurately identifying the radio frequency device individual while transmitting and receiving data, and realizing communication and perception integration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of neural network and radio frequency fingerprint recognition technology, specifically relating to a radio frequency fingerprint recognition method based on sample augmentation and deep learning. Background Technology

[0002] With the rapid development of wireless communication technology and the increasingly dense signal environment, the number of wireless communication devices is growing exponentially, leading to challenges in device privacy and security in both military and civilian fields. For example, in human-machine collaborative broadband self-organizing network baseband transceivers, due to cost and power consumption limitations, the data transmitted from the transmitter to the receiver is often unencrypted or weakly encrypted, which poses security risks. Since minor hardware defects are unavoidable in radio circuits, and manufacturing tolerances and drift tolerances exist in electronic components and printed circuit boards during manufacturing and use, radio frequency fingerprinting technology for wireless communication devices can identify wireless electronic devices by extracting features from radio frequency signals.

[0003] Radio frequency (RF) fingerprinting methods mainly include manual feature extraction and machine learning. Feature extraction analyzes the instantaneous amplitude, frequency, and phase information of statistical signals in their time domain, frequency domain, and power spectrum to extract certain statistical features as the basis for identification. With the development of artificial intelligence technology, the integration of deep learning and the field of communication is becoming increasingly close, and there has been relatively successful progress in signal modulation recognition. RF fingerprinting is essentially a type of pattern recognition. The characteristic of deep learning lies in constructing suitable deep network structures, completing the selection and extraction of original data features through nonlinear activation transformations in each layer. Utilizing the self-learning mechanism of deep neural networks, it is possible to improve the recognition rate of communication devices in complex environments and solve the influence of various real-time environmental factors such as noise, interference, fading, inter-symbol interference, and hardware defects. This is an inevitable trend in the research of RF fingerprinting technology. The deep learning methods applied to RF fingerprinting mainly include convolutional neural networks (CNNs) and recurrent neural networks (RNNs).

[0004] The current field of radio frequency fingerprint recognition has the following problems: (1) As CNN networks become deeper, traditional CNNs require a large amount of training data and have a long training time. During the training process, the gradient hardly increases as the number of layers increases, resulting in the gradient vanishing problem and thus a decrease in recognition accuracy. (2) The traditional data preprocessing method is to first perform time synchronization, frequency offset compensation, phase compensation and other operations on the Raw-IQ data before inputting it into the neural network model. However, the data features after these preprocessing methods may cover the radio frequency fingerprint features in the original IQ data. (3) Although CNN, ResNet, LSTM and other networks have achieved good recognition accuracy, and some studies have performed data augmentation on data with different signal-to-noise ratios (SNR), the problems in the field of radio frequency fingerprint recognition have not been deeply considered: the training data and test data are not collected on the same day, and the data is collected using different receivers. These will lead to differences in the distribution between the training set and the test set, thereby degrading the performance of the deep learning model and making the network model less robust. Summary of the Invention

[0005] The purpose of this invention is to overcome the problems of gradient vanishing, excessive training time, and high computational cost in traditional CNN networks. This invention proposes a method for applying the DenseNet algorithm to radio frequency fingerprint recognition in wireless communication devices. It collects raw IQ samples (Raw IQ, representing the instantaneous IQ value of a signal at equal sampling intervals over a certain period; a single IQ sample point represents the instantaneous IQ value of the signal at a specific moment) from the device receiver. Through transfer learning (TL), sample augmentation is performed, increasing the generalization ability even when there are significant differences in the distribution of the training and test sets, such as when data is collected on different dates or received from the same RF source using different receivers. This invention names this neural network the DSEN-TL network, a simplified abbreviation of the deep learning network used, DenseNet (Densely Connected Convolutional Network).

[0006] To achieve the above-mentioned objectives, the present invention provides a radio frequency fingerprint recognition method based on sample augmentation and deep learning, comprising the following steps:

[0007] Step S1: Collect raw radio frequency data from the receiving end of the wireless communication device and store it on the PC.

[0008] Step S2: Preprocess the Raw-IQ data samples collected by the receiving end of the wireless communication device, divide the data into training set and test set, and further divide the training set into source domain and target domain.

[0009] Step S3: Construct the DSEN-TL neural network and input the training set to train it. During the training process, determine whether the H divergence has reached a balance. If it is too large or too small, the number of layers and parameters of the radio frequency individual identification network and the domain classification network must be adjusted and retrained to optimize the loss function.

[0010] Step S4: Input the test set into the DSEN-TL network and output the device identification type;

[0011] Step S5: Embed the trained DSEN-TL network into the actual board-level system for testing. While transmitting and receiving data, it accurately identifies individual radio frequency devices and realizes integrated communication and sensing.

[0012] Furthermore, in step S1, the wireless communication device used is the official ADRV9361-Z7035 development board from AD. The test platform consists of four transmitters and four receivers, labeled T1-T4 and R1-R4 respectively. The specific data acquisition method is as follows: the transmitters transmit data through antennas, and the receivers continuously capture raw IQ samples over the air interface in real time, saving them as files and transferring them to the PC. The dataset uses IQ samples transmitted using the IEEE 802.11a (WiFi) standard. To increase robustness, datasets transmitted by different devices at different dates and distances, as well as datasets received by different receivers from the same RF source, were collected.

[0013] Furthermore, the specific steps for data sample preprocessing and data partitioning in step S2 are as follows:

[0014] S21. First, divide the data collected on different dates into training and test sets in a ratio of 8:2. Then, further divide the training set into source and target domains. The training sets collected on different dates are divided as follows: First, set the data collected on the first day as the initial source domain (D1), which consists of labeled training signals (training set X1). Set the data collected on the second day as the initial target domain (D2), which consists of unlabeled signals to be identified (training set X2). After completing one round of training and testing on the first and second days, one round of network parameters is obtained. Then, set the data collected on the third day as the new target domain (D3, i.e., training set X3) for training, and so on, until a high accuracy is achieved. Similarly, the data collected from different receivers of the same RF source is divided as follows: First, divide all collected data into training and test sets in an 8:2 ratio. Then, first, set the data collected on R1 as the initial source domain (D2). R1 That is, training set X R1 The data collected by R2 is used as the initial target domain (D). R2 That is, training set X R2After training and testing, a new target domain is set. The source and target domains share the same RF source fingerprint features and categories, but the inconsistent feature distribution of data samples collected on different dates and from different receivers leads to a deterioration in the performance of traditional neural networks. Through transfer learning, the recognition network learned from the source domain signals exhibits better recognition performance in the target domain.

[0015] S22 and Raw-IQ data samples are unprocessed raw data, with no information loss and no suppression of confounding factors detrimental to RFID fingerprint recognition. The captured Raw-IQ samples are preprocessed into a 2×1024 two-dimensional array, with the first row representing in-phase sampling and the second row representing orthogonal sampling. The training set is fed into the DenseNet neural network as 1×2×1024 black-and-white images.

[0016] Furthermore, the DSEN-TL neural network in step S3 comprises three parts: a DenseNet feature extraction network, an RFID fingerprint individual identification network, and a domain classification network.

[0017] Furthermore, the specific structure of the DenseNet feature extraction network in step S3 is as follows:

[0018] S31, the DenseNet feature extraction network consists of an input layer, a first convolutional layer, a first pooling layer, three DenseBlock modules, and two transition layers and a second pooling layer connected in sequence. During the first round of training, the network's input layer includes a labeled source domain signal D1 (the training signal collected on the first day) and an unlabeled target domain signal D2 (the signal to be identified collected on the second day).

[0019] S311. The input to the first convolutional layer is 2×N two-dimensional data, used to extract latent features from the input data. Conv2D is used to perform two-dimensional convolution on the data. The first pooling layer selects the MaxPool max pooling strategy to calculate local maxima. Specifically, it first passes through a BN layer, then through a ReLU activation function layer, and finally performs max pooling with a pooling region size of (2×3) and a stride of (2×2) to reduce the number of parameters calculated by the model and reduce overfitting.

[0020] The S312 and DenseBlock modules each have 3 units. Each DenseNet network contains 5 convolutional layers, each consisting of 'a' convolutional layers of dimension (b, c). Compared to ResNet, DenseNet does not pass features through summation. Instead, it concatenates the inputs of each previous convolutional layer after each convolution, using this concatenation as the new input for the next convolutional layer. DenseNet does not extract representational power from extremely deep or wide architectures; instead, it leverages the network's potential through feature reuse, producing a condensed model that is easy to train and highly parameter-efficient. By connecting feature maps learned from different layers, it increases the variation of inputs to subsequent layers and improves efficiency. A transition layer is used between every two DenseNet layers. The transition layer structure is BN-ReLU-AveragePool, using mean pooling with a pooling region size of (2×2) and a stride of (2×2). The transition layer not only connects two DenseBlock modules but also reduces the model width by decreasing the number of feature maps input to the transition layer from the previous DenseBlock module, making the model more concise.

[0021] S313. The second pooling layer selects the MaxPool max pooling strategy, performing max pooling with a pooling region size of (2×3) and a step size of (2×2). The data after max pooling are input into the RF fingerprint individual identification network and the domain classification network, respectively, for training and transfer learning in the source domain and target domain.

[0022] Furthermore, the radio frequency fingerprint individual identification network and domain classification network in step S3 are specifically as follows:

[0023] S32. The feature vector f extracted from the source domain data in the second pooling layer is input into the RFID fingerprint individual identification network to obtain the RFID device label y, i.e., the RFID fingerprint identification result. At the same time, the feature vectors of the source domain signal and the target domain signal are jointly input into the domain classification network to obtain the domain label d.

[0024] S321, the radio frequency fingerprint individual identification network contains two fully connected layers. A softmax function is used as the output layer to calculate the probability of each category and output y1 to y2. m , representing the probability of identifying m different devices. Simultaneously, a dropout technique is added before the fully connected layer, with the dropout coefficient set to 0.5, meaning that only half of the neurons are active at any given time. This prevents overfitting during training due to an overly deep neural network, excessively long training time, or insufficient data.

[0025] S322. The domain recognition network consists of a gradient inversion layer (GRL), fully connected layers, and a SoftMax classifier. Feature vectors input to the domain classification network layers pass through a gradient inversion layer (GRL). This layer performs an identity transformation during forward propagation and automatically inverts the gradient during backpropagation. Gradient inversion multiplies the loss of the domain classifier by a negative coefficient δ, making the training objectives of the networks before and after it opposite, achieving a generative adversarial effect similar to that of generative adversarial networks (GANs), i.e., the feature extraction network G... f G domain classification network d The training objective is adversarial. During the training phase, the value of δ changes from 0 to 1 as the number of iterations increases, as shown in the following equation:

[0026]

[0027] Furthermore, the H-divergence mentioned in step S3 is a metric set to measure the difference between the distributions of D1 and D2. It can be used to determine whether a data sample belongs to the source domain or the target domain, thus deriving the condition for the classification network to transfer to the target domain. That is, it is necessary to simultaneously minimize the sum of the H-divergence and the classification errors in both the source and target domains. When the H-divergence is large enough, the data differences between the source and target domains are significant and easily distinguishable, resulting in minimal classification errors. However, to transfer the network to the target domain, the differences between the source and target domains must be minimized, ensuring that the data distributions in the two domains are similar. Therefore, the network training phase is an adversarial optimization process similar to that of a Generative Adversarial Network (GAN).

[0028] H-divergence provides a method for quantifying the differences between different domains. A classifier trained on data from one domain predicts results in two domains, and the difference between these results serves as the upper bound of the differences between the two domains.

[0029] Furthermore, the loss function to be optimized in step S3 is obtained by combining the parameters of the feature extraction network, the RFID fingerprint individual identification network, and the domain classification network. By optimizing the loss function of the entire model, the feature extraction network can extract features that are both discriminative and domain-invariant, thereby solving the problem of inconsistent signal distribution across different dates and receivers. During training, it is necessary to balance the RFID fingerprint individual identification network and the domain classification network. Specifically, when the H divergence is large, meaning the two domains are easily distinguishable, the domain classifier network may be overtrained, resulting in a small percentage of gradient backpropagation that is ineffective in guiding the feature extraction network to extract domain-invariant features. In this case, it is necessary to reduce the performance of the domain classifier by adjusting the number of layers and nodes in both networks.

[0030] The beneficial effects of this invention are as follows: This invention collects raw IQ samples (Raw-IQ) and constructs a DSEN-TL network. It alleviates the gradient vanishing problem in traditional CNN networks through feature reuse, reducing training parameters and training time. Furthermore, it enhances the sample through transfer learning, increasing the robustness of the RF fingerprint recognition network. It also has the following characteristics:

[0031] (1) Using DenseNet convolutional neural network for radio frequency fingerprint recognition, based on traditional CNN, the features of previous scale layers are fused, which makes more effective use of features, alleviates the gradient vanishing problem, and reduces training parameters and training time.

[0032] (2) The dataset uses raw samples (Raw-IQ), resulting in no information loss and no suppression of confounding factors detrimental to RF fingerprint recognition. While traditional RF fingerprint recognition employs sample augmentation techniques such as adding noise, establishing channel fading models, frequency offset simulation, and pseudo-random integration, it does not consider the recognition of data from different dates and receivers. This invention uses transfer learning to train samples from different dates and receivers, increasing the robustness of the RF fingerprint recognition network for wireless communication devices and providing mobility.

[0033] (3) It has been embedded in a small wireless broadband self-organizing network baseband system, which can be further applied to the field of large-scale communication and sensing integrated intelligent communication. Attached Figure Description

[0034] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0035] Figure 1 This is a technical flowchart of the present invention.

[0036] Figure 2 These are the system structure diagram and data capture structure diagram of the present invention.

[0037] Figure 3 This is a schematic diagram of the baseband receiver structure of the present invention.

[0038] Figure 4A This is a schematic diagram illustrating the division of the dataset collected on different dates according to the present invention. Figure 4B A schematic diagram illustrating the division of datasets collected by different receivers.

[0039] Figure 5 This is a schematic diagram of the DSN-TL network structure constructed in this invention.

[0040] Figure 6 This is a schematic diagram of the DenseBlock module connection in the DSN-TL network structure of the present invention. Detailed Implementation

[0041] The specific embodiments of the present invention will be described below to enable those skilled in the art to understand the present invention.

[0042] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0043] like Figure 1 As shown, a radio frequency fingerprint recognition method based on sample augmentation and deep learning includes the following steps:

[0044] Step S1: Collect raw radio frequency data from the receiving end of the wireless communication device and store it on the PC.

[0045] Step S2: Preprocess the Raw-IQ data samples collected by the receiving end of the wireless communication device, divide the data into training set and test set, and further divide the training set into source domain and target domain.

[0046] Step S3: Construct the DSEN-TL neural network and input the training set to train it. During the training process, determine whether the H divergence has reached a balance. If it is too large or too small, the number of layers and parameters of the radio frequency individual identification network and the domain classification network must be adjusted and retrained to optimize the loss function.

[0047] Step S4: Input the test set into the DSEN-TL network and output the device identification type;

[0048] Step S5: Embed the trained DSEN-TL network into the actual board-level system for testing. While transmitting and receiving data, it accurately identifies individual radio frequency devices and realizes integrated communication and sensing.

[0049] like Figure 2 As shown, the dataset in this paper originates from a 2×2 MIMO wireless communication system built based on the 802.11a (WiFi) physical layer transmission protocol. This system is implemented in FPGA hardware based on MATLAB theoretical simulation and tested at the board level using the Analog Devices ADRV9361-Z7035 official development board. The ADRV9361-Z7035 combines the Analog Devices AD9361 integrated RF agile transceiver with the Xilinx Z7035 Zynq-7000 All Programmable SoC, providing a wideband 2×2 receive and transmit path in the 70MHz to 6GHz range. Because the AD9361 has automatic gain control (AGC) functionality, after adjusting the AGC parameters, it can maintain a bit error rate of less than 1% at the receiver even in mobile devices, facilitating data acquisition in mobile environments.

[0050] like Figure 3The diagram shows the baseband receiver structure. In this invention, the sample acquisition point is after the data ADC and before timing synchronization. The WiFi standard uses Orthogonal Frequency Division Multiplexing (OFDM) and therefore uses multiple subcarriers to transmit each digital symbol. BPSK, QPSK, 16QAM, or 64QAM modulation can be used, with different levels of convolutional coding (1 / 2 or 3 / 4). L-STF is a traditional short training field, mainly used for initial timing synchronization, initial frequency offset estimation, and AGC settings. L-LTF is a traditional long training field, mainly used for precise timing synchronization, precise frequency offset estimation, and channel estimation. The transmitter processes the raw data through modules such as CRC, scrambling, BCC coding, puncturing interleaving, modulation, IFFT, cyclic shifting, and CP insertion. The receiver obtains the received data through modules such as coarse synchronization, frequency offset estimation, FFT, channel estimation, maximum ratio combining, and decoding.

[0051] The data samples of this invention are transmitted at a center frequency of 2.4 GHz and a sampling rate of 20 MHz. The modulation and coding scheme is MCS0 (i.e., BPSK modulation, code rate 1 / 2), and the transmission rate is approximately 5.34 Mbps. At the receiving end, the gain control mode is set to AGC, which can maintain the amplitude and power of the received signal stable even when the antenna distance changes and the mobile device moves slowly, thereby increasing the diversity of the collected data.

[0052] Furthermore, the specific steps for sample export in step S1 are as follows:

[0053] S11. Data collection was conducted on different devices at different distances in an open indoor corridor on different dates. On the first day, receiver R1 was fixed in an open corridor. Then, transmitter T1 was placed 1m away from R1 for 30 seconds of data acquisition. Next, T1 was slowly moved to a distance of 5m from R1 for another 30 seconds of data acquisition. Finally, T1 was slowly moved to a distance of 10m from R1 for another 30 seconds of data acquisition. Then, transmitters T2, T3, and T4 were used in sequence, and the above process was repeated. On the second day, the receiver was fixed in the same location at the same time, and different transmitters were used to repeat the acquisition process of the first day. Next, on the same day, transmitter T1 was fixed in a fixed position. Receiver R1 was placed at 1m, 5m, and 10m respectively, and then R2, R3, and R4 were placed in the same way.

[0054] S12. Single device capture preparation process: Connect the official development board ADRV9361-Z7035 to the host via serial cable and network cable and power it on. Start the data transmission switch in the Linux system on the board and confirm that the receiving end is receiving data normally.

[0055] S13. Run the Python capture program on the Linux system on the board, save the air interface data received by the development board in real time as a txt file and transmit it to the PC. The file contains real-time I and Q channel data.

[0056] like Figure 4A and Figure 4B The following are the specific steps for data sample partitioning in step 2:

[0057] S21. First, divide the data collected on different dates into training and test sets in a ratio of 8:2. Then, further divide the training set into source and target domains. The training sets collected on different dates are divided as follows: First, set the data collected on the first day as the initial source domain (D1), which consists of labeled training signals (training set X1). Set the data collected on the second day as the initial target domain (D2), which consists of unlabeled signals to be identified (training set X2). After completing one round of training and testing on the first and second days, one round of network parameters is obtained. Then, set the data collected on the third day as the new target domain (D3, i.e., training set X3) for training, and so on, until a high accuracy is achieved. Similarly, the data collected from different receivers of the same RF source is divided as follows: First, divide all collected data into training and test sets in an 8:2 ratio. Then, first, set the data collected on R1 as the initial source domain (D2). R1 That is, training set X R1 The data collected by R2 is used as the initial target domain (D). R2 That is, the training set X R2 After training and testing, a new target domain is set. The source and target domains share the same RF source fingerprint features and categories, but the inconsistent feature distribution of data samples collected on different dates and from different receivers leads to a deterioration in the performance of traditional neural networks. Through transfer learning, the recognition network learned from the source domain signals exhibits better recognition performance in the target domain.

[0058] S22 and Raw-IQ data samples are raw data that has not undergone timing synchronization, frequency correction, FFT, channel estimation, and equalization. No information is lost, and no confounding factors detrimental to RF fingerprint recognition are suppressed. The captured Raw-IQ samples are preprocessed into a 2×1024 two-dimensional array. The first row of the Raw-IQ samples represents in-phase samples, and the second row represents quadrature samples. The training set is fed into the DenseNet neural network as 1×2×1024 black and white images.

[0059] like Figure 5 As shown, the DSEN-TL neural network in step S3 includes three parts: DenseNet feature extraction network, radio frequency fingerprint individual identification network, and domain classification network.

[0060] S31, the DenseNet feature extraction network consists of an input layer, a first convolutional layer, a first pooling layer, three DenseBlock modules, and two transition layers and a second pooling layer connected in sequence. During the first round of training, the network's input layer includes a labeled source domain signal D1 (the training signal collected on the first day) and an unlabeled target domain signal D2 (the signal to be identified collected on the second day).

[0061] S311. The input to the first convolutional layer is 2×N two-dimensional data, used to extract latent features from the input data. Conv2D is used to perform two-dimensional convolution on the data. In a convolutional neural network, each neuron in each layer is only connected to a portion of the neurons in the previous layer, and neurons in the same layer share parameters, exhibiting characteristics of local perception, weight sharing, and shift invariance. The first pooling layer uses the MaxPool max pooling strategy to calculate local maxima. Specifically, it first passes through a BN layer, then a ReLU activation function layer, and finally performs max pooling with a pooling region size of (2×3) and a stride of (2×2), reducing the number of parameters calculated for the model and minimizing overfitting.

[0062] The number of S312 and DenseBlock modules is 3, such as Figure 6 As shown, each DenseNet network contains five convolutional layers. The first convolutional layer consists of 32 kernels with dimensions (1,3), the second and fifth convolutional layers consist of 32 kernels with dimensions (2,3), the third and fourth convolutional layers consist of 32 kernels with dimensions (1,3), and the fifth convolutional layer consists of 32 kernels with dimensions (2,3). Compared to ResNet, DenseNet does not pass features through summation. Instead, it concatenates the inputs of each previous convolutional layer after each convolution and feeds them as the new input to the next convolutional layer. DenseNet does not draw representational power from extremely deep or wide architectures. Instead, it utilizes the potential of the network through feature reuse, producing a condensed model that is easy to train and has high parameter efficiency. By connecting feature maps learned from different layers, it increases the variation of inputs to subsequent layers and improves efficiency. Specifically, the first convolutional layer takes the original input x0 as input and outputs x1; the second convolutional layer takes the input of x0 and x1 as a new matrix and outputs x2; the third layer takes the input of x0, x1, and x2 as a new matrix, and so on. Each layer's input comes from the outputs of all previous layers, and is concatenated with all preceding layers as input.

[0063] x l =H l ([x0,x1,…,x l-1 ])

[0064] Hl (·) represents a nonlinear transformation function, which is a combination operation that includes a series of Batch Normalization (BN), ReLU, and Conv operations. l This represents the output of layer l. Each neural network layer extracts features from the input data, and these features become more prominent as the layer depth increases. The feature propagation method involves directly concatenating the features from all preceding layers before passing them to the next layer, rather than having an arrow pointing from each preceding layer to all subsequent layers. ReLU indicates an activation layer with the ReLU function; using ReLU in all five convolutional layers can facilitate non-linear fitting.

[0065] A transition layer is used between every two DenseNet modules. The transition layer structure is BN-ReLU-AveragePool, which uses mean pooling with a pooling region size of (2×2) and a stride of (2×2). The transition layer not only connects the two DenseBlock modules, but also reduces the model width by decreasing the number of feature maps input to the transition layer from the previous DenseBlock module, making the model more concise.

[0066] S313. The second pooling layer selects the MaxPool max pooling strategy, performing max pooling with a pooling region size of (2×3) and a step size of (2×2). The data after max pooling are input into the RF fingerprint individual identification network and the domain classification network, respectively, for training and transfer learning in the source domain and target domain.

[0067] Furthermore, the radio frequency fingerprint individual identification network and domain classification network in step S3 are specifically as follows:

[0068] S32. The feature vector f extracted from the source domain data in the second pooling layer is input into the RFID fingerprint individual identification network to obtain the RFID device label y, i.e., the RFID fingerprint identification result. At the same time, the feature vectors of the source domain signal and the target domain signal are jointly input into the domain classification network to obtain the domain label d.

[0069] S321, the radio frequency fingerprint individual identification network contains two fully connected layers. A softmax function is used as the output layer to calculate the probability of each category and output y1 to y2. m , representing the probability of identifying m different devices. Simultaneously, a dropout technique is added before the fully connected layer, with the dropout coefficient set to 0.5, meaning that only half of the neurons are active at any given time. This prevents overfitting during training due to an overly deep neural network, excessively long training time, or insufficient data.

[0070] S322. The domain classification network consists of a gradient inversion layer (GRL), fully connected layers, and a SoftMax classifier. Feature vectors input to the domain classification network layers pass through a gradient inversion layer (GRL). This layer performs an identity transformation during forward propagation and automatically inverts the gradient during backpropagation. Gradient inversion multiplies the loss of the domain classifier (the difference between the predicted and true values) by a negative coefficient δ, making the training objectives of the networks before and after it opposite. This achieves a generative adversarial effect similar to that of generative adversarial networks (GANs), i.e., the feature extraction network G... f G domain classification network d The training objective is adversarial. During the training phase, the value of δ changes from 0 to 1 as the number of iterations increases, as shown in the following equation:

[0071]

[0072] Here, γ is a hyperparameter, typically set to a constant of 10. p changes from 0 to 1 during training, representing the current iteration number divided by the total number of iterations.

[0073] Furthermore, the definition of H-divergence in step S3 is as follows:

[0074] To measure the difference in distributions between D1 and D2, the H-divergence is used as the metric.

[0075]

[0076] Where X is the feature space mapped by the neural network, H is a hypothesis space, and h is a function in this space used for binary classification, i.e., h:X→{0,1}. h can effectively distinguish between source and target domain data if it satisfies the following conditions: the source domain data is classified as class 1 with a probability very close to 1, and the target domain data is classified as class 1 with a probability very close to 0.

[0077]

[0078]

[0079] Based on the H divergence, the formula for the HΔH distance is defined as follows:

[0080]

[0081] Where h1 and h2 are discriminant functions used to determine whether data belongs to the source domain or the target domain. From the above equation, we can see that HΔH can be used to measure the probability that the discriminant results of the two discriminant functions h1 and h2 are not equal.

[0082] HΔH={η:η(x * =1)}

[0083]

[0084] Therefore, by finding the extreme value of the HΔH distance, the maximum error of the discrimination results h1 and h2 between the source and target domains can be found, i.e., the H divergence. It should satisfy the following equation:

[0085]

[0086] Where λ is a constant. The above equation shows that the error generated by the neural network trained on source domain data when classifying target domain data is mainly caused by two factors: first, the inherent classification error of the network in the source domain, which is relatively small if the source domain dataset is sufficient; second, the HΔH distance between the data distributions of the source and target domains. To reduce the classification error of the classification network on the target domain data, it is necessary to reduce the classification error of the network in the source domain and to reduce the distance between the data distributions of the source and target domains. The above equation can be simplified to:

[0087] d HΔH (D1,D2)=2(1-min(∈1(h)+∈2(h)))

[0088] According to the above formula, in order to blur the boundary between the source domain and the target domain, it is necessary to maximize the sum of the classification errors of the network in the source domain and the target domain. Therefore, the essence of the problem is to optimize the minimax function, as shown in the following formula:

[0089]

[0090] The above equation shows that for a classification network to transfer to the target domain, it needs to simultaneously minimize the H-divergence and the sum of the classification errors in both the source and target domains. When the H-divergence is large enough, the data differences between the source and target domains are significant and easily distinguishable, resulting in minimal classification errors. However, to transfer the network to the target domain, the differences between the source and target domains must be minimized, ensuring that the data distributions in both domains are similar. Therefore, the network training phase is an adversarial optimization process similar to that of a Generative Adversarial Network (GAN).

[0091] Furthermore, the loss function in step 3 is calculated as follows:

[0092] The training data is first transformed into a one-dimensional feature vector through the feature extraction network: f = G f (a;θ f Then the feature vector branches to two networks, the radio frequency fingerprint individual identification network G. y (a;θ y ) and domain classification network G d (a;θ d The loss function for the entire model is calculated as follows:

[0093] E(θ f ,θy ,θ d )

[0094] =∑ i=1,2…N L y (G y (G f (a i ;θ f );θ y ),y i )+δ

[0095] ∑ i=1,2…N L d (G d (R λ G f (a i ;θ f );θ d ),y i )

[0096] By optimizing the aforementioned loss function, the feature extraction network can extract features that are both discriminative and domain-invariant, thereby solving the problem of inconsistent signal distributions collected from different dates and receivers. During training, it is necessary to balance the RF fingerprint individual identification network and the domain classification network. Specifically, when the H divergence is large, meaning the two domains are easily distinguishable, the domain classifier network may be overtrained, resulting in a small percentage of gradient backpropagation that is ineffective in guiding the feature extraction network to extract domain-invariant features. In this case, it is necessary to adjust the number of layers and nodes in both networks to reduce the performance of the domain classifier.

Claims

1. A radio frequency fingerprint recognition method based on sample augmentation and deep learning, characterized in that, Includes the following steps: Step S1: Collect raw radio frequency data from the receiving end of the wireless communication device and store it on the PC. Step S2: Preprocess the Raw-IQ data samples collected by the receiving end of the wireless communication device, divide the data into training set and test set, and further divide the training set into source domain and target domain. Step S3: Construct the DSEN-TL neural network and input the training set to train it. During the training process, determine whether the H divergence has reached a balance. If it is too large or too small, the number of layers and parameters of the radio frequency individual identification network and the domain classification network must be adjusted and retrained to optimize the loss function. Step S4: Input the test set into the DSEN-TL network and output the device identification type; Step S5: Embed the trained DSEN-TL network into the actual board-level system for testing. While transmitting and receiving data, it accurately identifies individual radio frequency devices and realizes integrated communication and sensing. The DSEN-TL neural network in step S3 comprises three parts: a DenseNet feature extraction network, a radio frequency fingerprint individual identification network, and a domain classification network; the specific structure of the DenseNet feature extraction network in step S3 is as follows: S31. The DenseNet feature extraction network consists of an input layer, a first convolutional layer, a first pooling layer, three DenseBlock modules, and two transition layers and a second pooling layer connected in sequence. During the first round of training, the input layer of the network includes a labeled source domain signal D1, which is the training signal collected on the first day, and an unlabeled target domain signal D2, which is the signal to be identified collected on the second day. S311. The input to the first convolutional layer is 2×N two-dimensional data, which is used to extract the latent features in the input data. Conv2D is used to perform two-dimensional convolution on the data. The first pooling layer selects the MaxPool max pooling strategy to calculate the local maximum. Specifically, it first goes through the BN layer, then through the ReLU activation function layer, and finally performs max pooling with a pooling region size of 2×3 and a stride of 2×2 to reduce the number of parameters calculated for the model and reduce overfitting. S312, the number of DenseBlock modules is 3, each DenseNet network contains 5 convolutional layers, each convolutional layer consists of a convolutional layers with dimensions (b,c); compared with the ResNet network, the DenseNet network does not pass features by summation, but concatenates the inputs of each previous convolutional layer after each convolution and sends them as new inputs to the next convolutional layer. DenseNet does not draw representational power from extremely deep or wide architectures, but rather leverages the potential of networks through features; resulting in condensed models that are easy to train and highly parameter-efficient. S313. The second pooling layer selects the MaxPool max pooling strategy to perform max pooling with a pooling region size of 2×3 and a step size of 2×2. The data after max pooling are input into the radio frequency fingerprint individual identification network and the domain classification network respectively for training and transfer learning in the source domain and target domain respectively.

2. The radio frequency fingerprint recognition method based on sample augmentation and deep learning according to claim 1, characterized in that, In step S1, the wireless communication device used is ADRV9361-Z7035. The test platform includes 4 transmitters and 4 receivers, labeled T1-T4 and R1-R4 respectively. The specific data acquisition method is as follows: the transmitter transmits data through the antenna, and the receiver continuously captures the original IQ samples of the air interface in real time and saves them as files to be transferred to the PC. The dataset uses IQ samples transmitted according to the IEEE 802.11a standard. Data sets transmitted by different devices at different dates and distances, as well as datasets received by different receivers from the same radio frequency source, were collected.

3. The radio frequency fingerprint recognition method based on sample augmentation and deep learning according to claim 1, characterized in that, The specific steps for data sample preprocessing and data partitioning in step S2 are as follows: S21. First, divide the data collected on different dates into training and test sets in a ratio of 8:

2. Then, further divide the training set into source and target domains. The training sets collected on different dates are divided as follows: First, set the data collected on the first day as the initial source domain D1, which is the labeled training signal (i.e., training set X1). Set the data collected on the second day as the initial target domain D2, which is the unlabeled signal to be identified (i.e., training set X2). After completing one round of training and testing on the first and second days, one round of network parameters is obtained. Then, set the data collected on the third day as the new target domain D3 (i.e., training set X3) for training, and so on, until a high accuracy is achieved. Similarly, the data collected from different receivers of the same RF source is divided as follows: First, divide all collected data into training and test sets in an 8:2 ratio. Then, first, set the data collected by R1 as the initial source domain D1. R1 That is, training set X R1 The data collected by R2 is used as the initial target domain D. R2 That is, training set X R2 After training and testing, a new target domain is set. The source domain and target domain have the same fingerprint features and categories from the same radio frequency source. However, due to the inconsistent feature distribution of data samples collected on different dates and by different receivers, the performance of traditional neural networks deteriorates. Through transfer learning, the recognition network learned from the source domain signals has a better recognition effect on the target domain. S22. The Raw-IQ data samples are unprocessed raw data, which have not lost any information and have not suppressed any confounding factors that are not conducive to RFID fingerprint recognition. The captured Raw-IQ samples are preprocessed into a 2×1024 two-dimensional array. The first row of the Raw-IQ samples is in-phase sampling and the second row is orthogonal sampling. The training set will be input into the DenseNet neural network in the form of 1×2×1024 black and white images.

4. The radio frequency fingerprint recognition method based on sample augmentation and deep learning according to claim 1, characterized in that, The radio frequency fingerprint individual identification network and domain classification network in step S3 are as follows: S32. The feature vector f extracted from the source domain data in the second pooling layer is input into the radio frequency fingerprint individual identification network to finally obtain the radio frequency device label y, i.e., the radio frequency fingerprint identification result; at the same time, the feature vectors of the source domain signal and the target domain signal are jointly input into the domain classification network to obtain the domain label d. S321, the radio frequency fingerprint individual identification network contains two fully connected layers. A softmax function is used as the output layer to calculate the probability of each category and output y1 to y2. m , representing the probability of identifying m different devices; at the same time, a dropout technique is added before the fully connected layer, with the dropout coefficient set to 0.5, meaning that only half of the neurons are active at any given time, to prevent overfitting caused by the neural network being too deep, the training time being too long, or the lack of sufficient data during the training process. S322. The domain recognition network consists of a gradient inversion layer (GRL), fully connected layers, and a SoftMax classifier. Feature vectors input to the domain classification network layers pass through a gradient inversion layer (GRL). This layer performs an identity transformation during forward propagation and automatically inverts the gradient during backpropagation. Gradient inversion multiplies the loss of the domain classifier by a negative coefficient δ, making the training objectives of the networks before and after it opposite, achieving a generative adversarial effect similar to that of Generative Adversarial Networks (GANs), i.e., the feature extraction network G... f G domain classification network d The training objective is adversarial; during the training phase, the value of δ changes from 0 to 1 as the number of iterations increases, as shown in the following formula: The H divergence mentioned in step S3 is a metric set to measure the difference between the distributions of D1 and D2. It is used to determine whether the data sample belongs to the source domain or the target domain, thereby deriving the condition that the classification network can transfer to the target domain, that is, it is necessary to minimize the sum of the H divergence and the classification error in both the source domain and the target domain. When the H divergence is large enough, the data differences between the source domain and the target domain are large and easy to distinguish, and the classification error is very small. However, in order to transfer the network to the target domain, the differences between the source domain and the target domain must be smaller to ensure that the data distributions of the two domains are similar. Therefore, the training phase of the network is an adversarial optimization process similar to that of a generative adversarial network (GAN).

5. The radio frequency fingerprint recognition method based on sample augmentation and deep learning according to claim 1, characterized in that, The loss function to be optimized in step S3 is obtained by combining the parameters of the feature extraction network, the radio frequency fingerprint individual identification network, and the domain classification network. By optimizing the loss function of the entire model, the feature extraction network can extract features that are both discriminative and domain invariant, thereby solving the problem of inconsistent signal distribution at different dates and different receivers. During the training process, it is necessary to balance the radio frequency fingerprint individual identification network and the domain classification network.

6. The radio frequency fingerprint recognition method based on sample augmentation and deep learning according to claim 1, characterized in that, When the H divergence is large, meaning the two domains are easily distinguishable, the domain classifier network may be trained too well, resulting in a small percentage of gradient backpropagation that is ineffective and fails to guide the feature extraction network to extract domain-invariant features. In this case, it is necessary to adjust the number of layers and nodes in the two networks to reduce the performance of the domain classifier.

Citation Information

Patent Citations

  • Electroencephalogram emotion classification method based on transfer learning

    CN114492560A

  • Small sample radio frequency fingerprint intelligent identification system and method

    CN114980122A