Radio frequency fingerprinting method and related device based on deep reinforcement learning and raw i / q
By combining deep reinforcement learning and Raw I/Q, the overfitting problem of RF fingerprint recognition under small-scale datasets is solved, and high-accuracy device recognition is achieved. Training is carried out using I/Q sample data, avoiding dependence on large-scale datasets and complex feature extraction steps.
Patent Information
- Application Number
- CN202310567059.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-18
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-05-18
AI Technical Summary
With small datasets, radio frequency fingerprint recognition technology is prone to overfitting, resulting in low recognition rates.
A radio frequency fingerprint recognition method based on deep reinforcement learning and Raw I/Q is adopted. By customizing the sample environment, a one-dimensional neural network model CNN is built, and DQN reinforcement learning is combined to design a reward function to train the sample data. I/Q sample data is then used for recognition.
With a small sample size, the recognition accuracy reaches over 98%, rapidly improving the device's recognition accuracy and avoiding the dependence on large-scale datasets and complex feature extraction steps in traditional methods.
Smart Images

Figure CN116680563B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of radio frequency fingerprint identification, and particularly relates to a radio frequency fingerprint identification method based on deep reinforcement learning and Raw I / Q and a related device. BACKGROUND
[0002] Radio frequency fingerprint identification technology is a category of automatic identification technology, which is initially applied to enemy aircraft identification, and then gradually applied to wireless device identification with the continuous development and research of researchers. Radio frequency fingerprint is composed of various unique hardware properties generated by electronic components of a device during production. Such characteristics will not change due to the modulation mode of wireless transmission and the transmission information content, so it is widely used in the field of device identification.
[0003] With the continuous development of machine learning and cross-application in various fields, the feature extraction step of radio frequency fingerprint is changed from traditional manual design features to using neural network models to learn hidden features. Studies have shown that manually designed features can be forged, for example, operations on baseband signals can change the carrier frequency offset and phase offset, and deep learning methods are more difficult to forge because there is no explicit feature design scheme.
[0004] However, due to the increasingly complex electromagnetic environment, and the difficulty of collecting data sets in some unknown environments, it is necessary to study how to fully utilize small-scale and limited data sets. Deep learning methods have shown excellent results in radiation source identification of large-scale data sets, but often overfitting occurs for small-scale data sets, resulting in low recognition rate. SUMMARY
[0005] The purpose of the present application is to provide a radio frequency fingerprint identification method based on deep reinforcement learning and Raw I / Q and a related device to solve the problem of overfitting of small-scale data sets, which often leads to low recognition rate.
[0006] To achieve the above purpose, the application adopts the following technical solutions:
[0007] In a first aspect, the application provides a radio frequency fingerprint identification method based on deep reinforcement learning and Raw I / Q, comprising:
[0008] Collecting I / Q sample data for user UE devices;
[0009] Customizing a sample environment for identifying user UE devices, and building a one-dimensional neural network model CNN;
[0010] Combining the customized sample environment and the built neural network model CNN, performing DQN reinforcement learning on the sample data;
[0011] The reward function is designed in combination with DQN reinforcement learning, and the sample data is trained. The simulation verification of the change of recognition accuracy at different training steps is carried out.
[0012] Optionally, when collecting I / Q sample data, the properties of the LTE radio frequency transmitter are modified, including co-directional quadrature IQ imbalance, phase noise and power amplifier gain, and five hardware property different devices are distinguished.
[0013] Optionally, 100 IQ data samples are collected for each device, each sample length is 7680*1, and the data set is divided into training set, verification set and test set according to 7:2:1.
[0014] Optionally, the environment includes a sample selection function, an action reward function, an action execution function and a state reset function.
[0015] Optionally, in combination with the custom sample environment and the built neural network model CNN, the sample data is subjected to DQN reinforcement learning:
[0016] Define the environment: the state of the environment is defined as a sample, each different sample represents a different state, and after executing the judgment action, it enters the next sample, and the reward is returned by the custom reward and punishment mechanism;
[0017] E-greedy strategy: the epsilon parameter is set between (0.01, 1) linearly decaying, and the decay coefficient is set to 0.0001; when the random number is less than epsilon, randomly select the action, otherwise select the action with the maximum q value;
[0018] The CNN model fits the Q table, and outputs the q value of each action under different states, and the q value is updated according to the following formula:
[0019] q_new(critic)=old_q(target)+alpha*(R+gamma*max(q’(target))-old_q);
[0020] The global parameters of the experiment are set as alpha=0.5, gama=0.5, the CNN includes four convolutional layers, two maximum pooling layers, the activation function uses the tanh function, and the last output layer uses the Dense layer;
[0021] Experience Replay, first build an experience pool memorry, continuously explore and save the results (s, a, r, s') of each exploration when the minimum sampling length is not reached; when the minimum sampling length of the experience pool is reached, randomly sample the data in the experience pool for training and learning;
[0022] The DQN is trained, first, the environment is initialized to obtain an initial state, then an action is selected by a greedy function and a corresponding q value is obtained, then the environment is interacted to obtain a next state and a reward of the action, a loss is calculated from the q value of the next state, the training is returned, and the state is updated to the next state, thus completing a learning; after reaching a preset learning number, the critic network parameters are copied to the target network until the cycle ends.
[0023] Optionally, a reward function is designed, and the reward function is inconsistent for different numbers of device identification; when the number of devices is 5, the reward function is designed as: 10 points are rewarded for judging the correct action, 1 point is deducted for a difference of 1 between the action and the actual target, 2 points are deducted for a difference of 2, 3 points are deducted for a difference of 3, and 4 points are deducted for a difference of 4; the penalty for a wrong judgment is determined by the difference between the judgment action and the actual target, and the larger the difference, the smaller the reward.
[0024] Optionally, the training process is as follows:
[0025] Simulation verification of the change of the identification accuracy under the training step numbers of 200, 400, 600, 1000, 2000 and 4000 is completed; and simulation training of the signal-to-noise ratios of 20 dB, 30 dB and 40 dB is completed under the training step number of 4000.
[0026] In a second aspect, the present application provides a radio frequency fingerprint identification system based on deep reinforcement learning and Raw I / Q, comprising:
[0027] A data acquisition module is configured to acquire I / Q sample data of a user UE device;
[0028] An environment building module is configured to build a sample environment for identifying the user UE device, and build a one-dimensional neural network model CNN;
[0029] A reinforcement learning module is configured to perform DQN reinforcement learning on sample data in combination with the custom sample environment and the built neural network model CNN;
[0030] A training output module is configured to perform training on sample data in combination with a reward function designed based on DQN reinforcement learning, and perform simulation verification of the change of the identification accuracy under different training step numbers.
[0031] In a third aspect, the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the radio frequency fingerprint identification method based on deep reinforcement learning and Raw I / Q when executing the computer program.
[0032] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program, when executed by a processor, implements the steps of the radio frequency fingerprinting method based on deep reinforcement learning and Raw I / Q.
[0033] Compared with the prior art, the present application has the following technical effects:
[0034] The present application proposes a radio frequency fingerprinting technology based on reinforcement learning for the scenario where the channel environment and signal modulation are unknown and sample data are difficult to collect in large quantities. The advantages of reinforcement learning, which does not require a large number of data labels and does not require a very fine feature extraction step, are used to train the I / Q sample data collected by the UE device, and the combination of DQN and radio frequency fingerprinting technology is realized. Experimental results show that, compared with supervised learning, deep reinforcement learning can improve the device recognition accuracy faster under a small number of samples, and the accuracy is more than 98%. The advantages are as follows:
[0035] First, the scheme uses a radio frequency fingerprint based on Raw IQ, which uses the difference in the IQ modulator attribute in the transmission end radio frequency transmission to distinguish different devices. Compared with the traditional radio frequency fingerprinting technology, the present scheme does not perform complex feature extraction on the collected samples, and does not require prior knowledge of the modulation method, channel environment, etc. The amplitude and phase features of the original IQ data are directly used for learning.
[0036] Second, the scheme combines a deep reinforcement learning framework. Under the condition of expensive sample collection at the receiving end, limited data samples cannot make the model fully learn the nonlinear mapping relationship between the data and the label, so that the recognition accuracy cannot reach the expected value; while deep reinforcement learning can balance between exploration and utilization, use the collected sample data to obtain rewards, and can remember and quickly learn potential features in the initial learning, so that better action selection can be obtained in the future, so that the sample data can be fully utilized.
[0037] Third, for the combination of device recognition and DQN algorithm, the present scheme defines a reinforcement learning environment suitable for device recognition, and designs a reward function for the reward sparsity problem caused by discrete action space and state space. The accuracy difference under different reward modes is compared and analyzed, and the experimental results show that the designed reward function can significantly improve the recognition accuracy of the algorithm. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 is a flowchart of the DQN of the present application;
[0039] Figure 2 is a data collection module diagram of the present application;
[0040] Figure 3 is a CNN model of the present application;
[0041] Figure 4 is a data preprocessing flowchart of the present application;
[0042] Figure 5 is a comparison chart of the results before and after the reward and punishment function optimization of the present application. DETAILED DESCRIPTION
[0043] The present application is further described below in conjunction with the accompanying drawings:
[0044] Please refer to Figures 1 to 5 The present application provides a UE wireless device identification method based on deep reinforcement learning, which can identify and classify different UE devices in the case of difficulty in labeling data label values and sample data collection, and has a high accuracy of 98%. The present application also compares and analyzes the performance changes of the algorithm under different signal-to-noise ratios. In the optimization of the DQN algorithm, the reward function is designed and adjusted to avoid falling into a local optimal strategy in multiple iterations.
[0045] The present application proposes a radio frequency fingerprint identification technology based on deep reinforcement learning for the case of unknown channel environment and signal modulation, and difficulty in collecting a large amount of sample data. The advantages of reinforcement learning, which does not require a large number of data labels and does not require a very fine feature extraction step, are used to train the I / Q sample data collected by the UE device, realizing the combination of DQN and radio frequency fingerprint identification technology. In the optimization of the DQN algorithm, the reward function is designed and adjusted: the reward function is designed to reward 10 points for correct action, deduct 1 point for 1 unit difference between action and actual target, deduct 2 points for 2 unit difference, deduct 3 points for 3 unit difference, and deduct 4 points for 4 unit difference. The punishment for the wrong judgment is determined by the difference between the judgment action and the actual target, the larger the difference, the smaller the reward. And to prevent the agent from falling into the situation of taking a step forward in exploration, all wrong rewards are deducted, this design not only prevents the agent from falling into a local optimal dead loop, but also urges the agent to learn correct judgment more quickly.
[0046] The present application is implemented by the following technical solutions:
[0047] A UE device identification simulation method based on deep reinforcement learning:
[0048] Step 1: According to the experimental purpose, the LTE RF Transmitter example needs to be simulated on Simulink Figure 2), and 5 data samples of different UE devices are collected. The application distinguishes 5 devices with different hardware properties by modifying the properties of the LTE radio frequency transmitter, including in-phase / quadrature (IQ) imbalance, phase noise and power amplifier gain, and the device property parameters are as shown in Table 1. 100 IQ data samples are collected for each device, each sample has a length of 7680*1, and the data set is divided into a training set, a validation set and a test set according to a ratio of 7:2:1. The processing process is as shown in Figure 4
[0049] Table 1: LTE radio frequency transmitter parameter settings
[0050] Device ID I / Q gain I / Q phase VGA noise HPA gain HPA noise HPA IP3 1 1.18 5 2.5 13 7.3 48 2 1.6 4.131 3.4 19 7.2 25 3 4 4 3.7 20 7 39 4 1.2 4.9 2.7 23 7.3 45 5 1.1 5.2 2.5 3 7.1 50
[0051] Step 2: According to the requirement analysis, the DQN code simulation is written. The environment for UE device identification is customized, and the environment framework is imitated from the openAI environment framework structure. The written environment includes a sample selection function, an action reward function, an action execution function and a state reset function.
[0052] Step 3: A one-dimensional CNN (model structure as shown in Figure 3 ) is built to output the q values of different states.
[0053] Step 4: A complete DQN framework (pseudo code flow as shown in the table below) is built, which combines the customized sample environment and the built CNN to explore and learn the sample data.
[0054]
[0055]
[0056] Step 5: The reward function is designed, and the reward function is inconsistent for different number of device identification. The number of experimental devices in the application is 5, and the reward function is designed as follows: 10 points are rewarded for correct action, 1 point is deducted for a difference of 1 between the action and the actual target, 2 points are deducted for a difference of 2, 3 points are deducted for a difference of 3, and 4 points are deducted for a difference of 4. The punishment for incorrect judgment is determined by the difference between the judgment action and the actual target, and the larger the difference, the smaller the reward. In order to prevent the agent from being stuck in the same place during exploration, all incorrect rewards are deducted. This design not only prevents the agent from falling into a local optimal dead loop, but also urges the agent to learn the correct judgment more quickly.
[0057] Step 6: After completing all the simulation part design, the samples are trained. The simulation verification of the identification accuracy change under the training step number of 200, 400, 600, 1000, 2000 and 4000 is completed. And under the training step number of 4000, the simulation training of the signal-to-noise ratio of 20dB, 30dB and 40dB is completed.
[0058] Signal to noise ratio / db 20 30 40 Recognition accuracy 0.5 0.84 0.98
[0059] Referring to Figure 2 As shown in the figure, the application collects the original RF data sample as the input state of reinforcement learning. The data collection environment is set as an awgn channel, the noise effect is Gaussian white noise, the signal-to-noise ratio of the channel is set to 40db, the parameters of the RFTransmitter are modified, and the collection stop time is set to 0.1s.
[0060] Referring to Figure 1 The DQN implementation flowchart is shown in the figure:
[0061] (1) Define the environment: the state of the environment is defined as our sample, and each different sample represents a different state. After executing the judgment action, it enters the next sample, and the reward is returned by the self-defined reward and punishment mechanism.
[0062] (2) E-greedy strategy: the epsilon parameter is set between (0.01, 1) linearly decaying, and the decay coefficient is set to 0.0001. When the random number is less than epsilon, a random action is selected, otherwise the action with the maximum q value is selected.
[0063] (3) CNN model fitting Q table, outputting the Q value of each action under different states. The q value is updated according to the following formula:
[0064] q_new(critic) = old_q(target) + alpha*(R + gamma*max(q'(target))-old_q);
[0065] The global parameters of the experiment are set to alpha = 0.5 and gama = 0.5, the CNN contains 4 convolutional layers, two maximum pooling layers, the activation function uses the tanh function, and the last output layer uses the Dense layer. The specific convolutional channel number and convolution kernel size are set as shown in the figure. Figure 2
[0066] (4) Experience Replay, first establish an experience pool memory, continuously explore and save the results of each exploration (s, a, r, s') when the minimum sampling length is not reached; when the minimum sampling length of the experience pool is reached, randomly sample the data in the experience pool for training and learning.
[0067] (5) Start training DQN, first initialize the environment, get the initial state, then select the action by the greedy function and get the corresponding q value, then interact with the environment, get the next state and the reward of the action. The loss is calculated by the q value of the next state, and the state is updated to the next state, thus completing a learning. After a certain number of learning, the critic network parameters are copied to the target network until the end of the cycle.
[0068] Referring to Figure 5 Compared with the CNN-based method, the UE device identification method based on deep reinforcement learning can learn to make correct judgments from small batches of samples faster when the number of training steps is less than 1000. Before the improvement of the reward and punishment mechanism based on DQN, the reward and punishment mechanism is to add one point for correct judgment and deduct one point for incorrect judgment, and the final identification accuracy is only 0.8, and the accuracy basically does not change after 2000 training steps. The reward and punishment mechanism of the present scheme is set to add 10 points for correct judgment and deduct points for incorrect judgment, and the punishment size is determined by the difference between the judgment made and the correct target (for example, 1 point is deducted for a difference of 1, 3 points are deducted for a difference of 3, and the punishment intensity is gradually increased), and finally the identification accuracy reaches 0.98.
[0069] In another embodiment of the present application, a radio frequency fingerprint identification system based on deep reinforcement learning and Raw I / Q is provided, which can be used to implement the radio frequency fingerprint identification method based on deep reinforcement learning and Raw I / Q described above. Specifically, the system comprises:
[0070] The data acquisition module is used to acquire I / Q sample data of the user UE device.
[0071] The environment building module is used to customize the sample environment for identifying the user UE device, and build a one-dimensional neural network model CNN.
[0072] The reinforcement learning module is used to combine the customized sample environment and the built neural network model CNN to perform DQN reinforcement learning on the sample data.
[0073] The training output module is used to combine the DQN reinforcement learning to design a reward function, train the sample data, and simulate and verify the change of the identification accuracy under different training steps.
[0074] The division of the modules in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, another division mode can be used. In addition, the function modules in each embodiment of the present application can be integrated in one processor, or can be physically separated, or two or more modules can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software function module.
[0075] In another embodiment of the present application, a computer device is provided, which comprises a processor and a memory, the memory is configured to store a computer program, the computer program comprises program instructions, and the processor is configured to execute the program instructions stored in the computer storage medium. The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, and are suitable for implementing one or more instructions, and are particularly suitable for loading and executing one or more instructions in the computer storage medium to implement a corresponding method flow or a corresponding function; the processor in the embodiments of the present application can be used for the operation of the radio frequency fingerprinting method based on deep reinforcement learning and Raw I / Q.
[0076] In another embodiment of the present application, the present application further provides a storage medium, specifically a computer readable storage medium (Memory), which is a memory device in a computer device, and is configured to store programs and data. It can be understood that the computer readable storage medium herein can include an internal storage medium in the computer device, and of course can also include an expansion storage medium supported by the computer device. The computer readable storage medium provides a storage space, and the storage space stores an operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and the instructions can be one or more computer programs (including program codes). It should be noted that the computer readable storage medium herein can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory. One or more instructions stored in the computer readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the radio frequency fingerprinting method based on deep reinforcement learning and Raw I / Q in the above embodiments.
[0077] Those skilled in the art will appreciate that embodiments of the application can be devised for a method, a system, or a computer program product. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.
[0078] The present application is described in reference to the flowchart and / or block diagrams of the method, apparatus (system) and computer program product according to embodiments of the application. It will be understood that each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing device or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 means for carrying out each of the one or more functions specified in the flowchart and / or block diagram block or blocks.
[0079] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 means for carrying out each of the one or more functions specified in the flowchart and / or block diagram block or blocks.
[0080] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 means for carrying out each of the one or more functions specified in the flowchart and / or block diagram block or blocks.
[0081] Finally, it should be noted that the above-mentioned embodiments are merely intended for describing and illustrating, not limiting, the technical solution of the present application. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that the specific embodiments of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, and any modification or equivalent replacement without departing from the spirit and scope of the present application should be covered in the protection scope of the claims of the present application.
Claims
1. A radio frequency fingerprint recognition method based on deep reinforcement learning and Raw I / Q, characterized in that, include: Collect I / Q sample data from user UE devices; Customize and build a sample environment for user UE device identification, and build a one-dimensional neural network model CNN; By combining a custom sample environment with a built neural network model (CNN), DQN reinforcement learning is performed on the sample data. A reward function was designed using DQN reinforcement learning, and the sample data was used for training. Simulation verification was performed to show the changes in recognition accuracy under different training steps. By combining a custom sample environment and a built neural network model (CNN), DQN reinforcement learning is performed on the sample data: Define the environment: The state of the environment is defined as a sample. Each different sample represents a different state. After performing a judgment action, the process moves to the next sample, and a reward is returned by a custom reward and punishment mechanism. ϵ-greedy strategy: The epsilon parameter is set to decay linearly between (0.01, 1), and the decay coefficient is set to 0.0001; when the random number is less than epsilon, an action is randomly selected; otherwise, the action with the largest q value is selected. The CNN model fits a Q-table and outputs the q-values for each action in different states. The q-values are updated according to the following formula: q_new(critic) =old_q(target) + alpha * (R + gamma * max(q'(target))-old_q); The experiment was conducted with global parameters alpha=0.5 and gama=0.
5. The CNN consisted of 4 convolutional layers, 2 max pooling layers, and the activation function was tanh. The final output layer was a Dense layer. Experience Replay first establishes an experience pool memory. Before reaching the minimum sampling length, it continuously explores and saves the results of each exploration (s, a, r, s'). After the experience pool reaches the minimum sampling length, it randomly samples data from the experience pool for training and learning. To train DQN, first initialize the environment to obtain the initial state, then select an action by a greedy function and obtain the corresponding q value, then interact with the environment to obtain the next state and the reward for this action; The loss is calculated from the q-value of the next state, the training is returned, and the state is updated to the next state, thus completing one learning cycle; After reaching the preset number of learning iterations, the parameters of the critic network are copied to the target network until the loop ends.
2. The radio frequency fingerprint recognition method based on deep reinforcement learning and Raw I / Q according to claim 1, characterized in that, When collecting I / Q sample data, the attributes of the LTE radio frequency transmitter were modified, including in-phase quadrature IQ imbalance, phase noise, and power amplifier gain, to distinguish five devices with different hardware attributes.
3. The radio frequency fingerprint recognition method based on deep reinforcement learning and Raw I / Q according to claim 2, characterized in that, Each device collects 100 IQ data samples, each sample is 7680*1 in length, and the dataset is divided into training set, validation set and test set in a 7:2:1 ratio.
4. The radio frequency fingerprint recognition method based on deep reinforcement learning and Raw I / Q according to claim 1, characterized in that, The environment includes a sample selection function, an action reward function, an action execution function, and a state reset function.
5. The radio frequency fingerprint recognition method based on deep reinforcement learning and Raw I / Q according to claim 1, characterized in that, The reward function is designed to vary depending on the number of devices recognizing the target: When there are 5 devices, the reward function is designed as follows: a correct action is rewarded with 10 points, and a difference of 1 point between the action and the actual target is deducted for every 1 point, 2 points for every 2 points, 3 points for every 3 points, and 4 points for every 4 points. The penalty for incorrect judgment is determined by the difference between the judged action and the actual target; the greater the difference, the smaller the reward.
6. The radio frequency fingerprint recognition method based on deep reinforcement learning and Raw I / Q according to claim 1, characterized in that, Training process: Complete the simulation verification of the recognition accuracy variation under training steps of 200, 400, 600, 1000, 2000, and 4000. Simulation training was completed with signal-to-noise ratios of 20dB, 30dB, and 40dB at a training step count of 4000.
7. A radio frequency fingerprint recognition system based on deep reinforcement learning and Raw I / Q, characterized in that, include: The data acquisition module is used to collect I / Q sample data from the user UE device; The environment setup module is used to customize and build the sample environment for user UE device recognition and to build a one-dimensional neural network model CNN. The reinforcement learning module is used to perform DQN reinforcement learning on sample data by combining a custom sample environment and a built neural network model (CNN). The training output module is used to design a reward function in conjunction with DQN reinforcement learning, train on sample data, and simulate and verify the changes in recognition accuracy under different training steps. By combining a custom sample environment and a built neural network model (CNN), DQN reinforcement learning is performed on the sample data: Define the environment: The state of the environment is defined as a sample. Each different sample represents a different state. After performing a judgment action, the process moves to the next sample, and a reward is returned by a custom reward and punishment mechanism. ϵ-greedy strategy: The epsilon parameter is set to decay linearly between (0.01, 1), and the decay coefficient is set to 0.0001; when the random number is less than epsilon, an action is randomly selected; otherwise, the action with the largest q value is selected. The CNN model fits a Q-table and outputs the q-values for each action in different states. The q-values are updated according to the following formula: q_new(critic) =old_q(target) + alpha * (R + gamma * max(q'(target))-old_q); The experiment was conducted with global parameters alpha=0.5 and gama=0.
5. The CNN consisted of 4 convolutional layers, 2 max pooling layers, and the activation function was tanh. The final output layer was a Dense layer. Experience Replay first establishes an experience pool memory. Before reaching the minimum sampling length, it continuously explores and saves the results of each exploration (s, a, r, s'). After the experience pool reaches the minimum sampling length, it randomly samples data from the experience pool for training and learning. To train DQN, first initialize the environment to obtain the initial state, then select an action by a greedy function and obtain the corresponding q value, then interact with the environment to obtain the next state and the reward for this action; The loss is calculated from the q-value of the next state, the training is returned, and the state is updated to the next state, thus completing one learning cycle; After reaching the preset number of learning iterations, the parameters of the critic network are copied to the target network until the loop ends.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the radio frequency fingerprinting method based on deep reinforcement learning and Raw I / Q as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the radio frequency fingerprinting method based on deep reinforcement learning and Raw I / Q as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Electromagnetic radiation source identification method based on deep reinforcement learning
CN113221454A
Decision-making method based on deep reinforcement learning
WO2022083029A1