Radio frequency fingerprint identification method and system based on convolution-attention mechanism and multi-packet reasoning
By combining the radio frequency fingerprint recognition method of convolutional neural network and Transformer encoder, the problem of insufficient feature extraction in a noisy environment is solved, the balance between global and local features is achieved, and the recognition accuracy and stability are improved.
Patent Information
- Application Number
- CN202510820264.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The existing RF fingerprint recognition model has degraded performance in noisy environments, insufficient feature extraction, and it is difficult to achieve a balance between global and local features.
Using a radio frequency fingerprint recognition method based on convolution-attention mechanism and multi-packet inference, combined with a convolutional neural network and a Transformer encoder, local features are extracted through convolutional layer, Transformer encoder captures global dependencies, and introduces a multi-packet adaptive fusion method to improve recognition performance.
It improves the recognition accuracy and robustness of the model in a noisy environment, reduces the computational complexity, and improves the accuracy and stability of device recognition.
Smart Images

Figure CN120337998A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a radio frequency fingerprint recognition method and system based on a convolutional-attention mechanism and multi-packet inference, and belongs to the technical fields of communication networks and artificial intelligence. Background Art
[0002] Radio Frequency Fingerprinting (RFF) is an important technology for device authentication and security guarantee in wireless communication systems. By extracting the unique features generated by wireless devices during signal transmission, RFF can effectively identify device identities. Traditional radio frequency fingerprint recognition methods usually rely on specific feature engineering and classifiers based on statistical methods. These methods often show low robustness and classification accuracy when facing complex channel conditions, signal noise interference, and similarity in device characteristics. In recent years, with the development of deep learning technology, radio frequency fingerprint recognition models based on deep learning have gradually become mainstream. These models can automatically extract the time-frequency domain features of signals and greatly improve the accuracy of device classification.
[0003] Currently, the application of deep learning models in the field of radio frequency fingerprint recognition mainly focuses on classic structures such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory Networks (LSTMs). However, a single model may face limitations when dealing with complex signal features. For example, CNNs perform excellently in capturing local features but lack the ability to model global time series information; recurrent networks such as LSTMs can capture the long dependencies of time series, but have low computational efficiency and are difficult to process high-dimensional signal data.
[0004] To address the above problems, Transformer has gradually become a research hotspot for processing sequential signals due to its powerful global feature modeling ability and parallel computing efficiency. However, directly applying Transformer to radio frequency fingerprint recognition still faces some challenges, such as the locality of signal features not being fully utilized, and the global modeling ability may be restricted by the expression ability of input features. To solve these problems, there is an urgent need for a radio frequency fingerprint recognition method that can simultaneously capture the local features and global dependencies of signals to improve the classification accuracy of the model and the robustness to noisy data. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technologies, the present invention provides a radio frequency fingerprint recognition method and system based on a convolutional-attention mechanism and multi-pack inference, which are used to solve the problems that the performance of traditional radio frequency fingerprint recognition models degrades in a noisy environment, feature extraction is insufficient, and it is difficult to achieve a balance between global and local features for complex data relying on a single model. By deeply analyzing the local and global feature distribution laws of radio frequency signals, the present invention proposes a deep learning model architecture-RFF-CAT that combines a convolutional neural network (CNN) and a Transformer encoder. This model uses convolutional layers to strengthen the local feature extraction ability, extracts local features from the original radio frequency signals, and effectively removes noise and retains key signal features through layer-by-layer convolutional operations. At the same time, the Transformer encoder module captures the global dependencies of radio frequency signals through the multi-head self-attention mechanism and enhances the sequence modeling ability by combining explicit position encoding.
[0006] In terms of module design, aiming at the problems of high computational complexity and redundant module interaction existing in the traditional model fusion mechanism, the present invention adopts a module separation strategy to independently optimize the convolutional module and the Transformer module, which not only reduces the computational complexity but also improves the stability of the model. To further improve the accuracy of the model, the present invention also introduces a multi-pack adaptive fusion method (MPF). This method divides the data into multiple packs and performs weighted fusion on different packs by weighted summation, thereby improving the recognition performance in a low signal-to-noise ratio (SNR) environment.
[0007] The technical solution of the present invention is as follows: The first aspect of the present invention provides a radio frequency fingerprint recognition method based on a convolutional-attention mechanism and multi-pack inference, including: Step 1, capture the device transmission signal and preprocess the transmission signal to obtain a spectrogram; Step 2, construct a radio frequency fingerprint recognition model, input the obtained spectrogram into the radio frequency fingerprint recognition model for training, and obtain a trained radio frequency fingerprint recognition model; The radio frequency fingerprint recognition model includes a convolutional layer, a position encoding layer, and a Transformer encoder; The convolutional layer converts the input feature map into a higher-order feature representation; The position encoding layer is used to add position information (absolute or relative position) to enable the model to have the ability to process sequential information; The Transformer encoder includes a multi-head self-attention mechanism and a feed-forward neural network layer (FFN). The multi-head self-attention mechanism captures the global dependencies in the entire sequence, and the feed-forward neural network layer further enhances the feature expression ability; Step 3: Use the trained radio frequency fingerprint recognition model for prediction, and combine the multi-packet inference method to obtain the final prediction result, realizing the authentication and recognition of the device.
[0008] Preferably according to the present invention, capture the device transmission signal and preprocess the transmission signal to obtain a spectrogram; including: First, collect the transmission signal from the LoRa device and record it as r(n), and the LoRa device uses the LoRa protocol; Then, calculate the carrier frequency offset (CFO) as shown in the following formula: ; where arg(⋅) represents the phase of the transmission signal, n represents the time variable, represents the carrier frequency offset; Compensate for the carrier frequency offset (CFO) to reduce the error caused by the carrier frequency offset, as shown in the following formula: ; where, is the received signal, is the signal after compensation, j represents the imaginary unit, CFO represents the carrier frequency offset, i.e., , and t represents the time variable; Use data normalization to normalize the compensated signal to unit amplitude, eliminating the influence of the difference in transmission power between devices on classification, as shown in the following formula: ; where r norm (n) represents the normalized received signal, whose amplitude is limited within the range of [-1, 1], used to unify the energy scale of signals from different devices, enhance the sensitivity of the model to the structural features in the device fingerprint, and reduce the interference of the transmission power difference on the classification result; Finally, perform short-time Fourier transform (STFT) to convert it into a channel-independent spectrogram to reduce the channel influence, as shown in the following formula: ; where, represents the spectrogram, represents the matrix window with length N, and is the sliding window function, specifically represented as a rectangular window function with length N = 64. Select a segment of signal at time point n to each moment t for transformation, j represents the imaginary unit, and f represents the frequency.
[0009] Preferably according to the present invention, construct a radio frequency fingerprint recognition model, input the obtained spectrogram into the radio frequency fingerprint recognition model for training, and obtain the trained radio frequency fingerprint recognition model; including: The convolutional layer includes a normalization layer, a pointwise convolutional layer, a GLU activation function layer, a depthwise separable convolutional layer, a batch normalization layer, a Swish activation function layer, a concatenated convolutional layer, and a Dropout layer; First, the obtained spectrogram is input into the convolutional layer, and the input spectrogram is normalized so that its mean is 0 and variance is 1 to improve training stability, obtaining a more uniformly distributed spectral feature, and retaining the normalized output of the original structure; the channel dimension is adjusted using 1x1 convolution through the pointwise convolutional layer to fuse or expand the feature information, obtaining a feature map with adjusted channel numbers; the key information of the feature map is enhanced through the GLU activation function layer, and the redundant information of the feature map is suppressed; the number of parameters of the feature map is reduced through the depthwise separable convolutional layer to enhance the spatial features; a feature map with stable distribution is obtained through the batch normalization layer; an activated feature map is obtained through the Swish activation function layer, retaining the small gradient information in the negative value region; a feature representation with optimized final channel numbers is obtained through the concatenated convolutional layer; and finally, a regularized feature map is obtained through the Dropout layer ; This module extracts low-level local features in the signal, such as the edges and textures of the spectrogram, which contribute to subsequent pattern recognition of the signal; Then, in order to retain the time and order information, a position encoding layer is introduced after the convolutional layer operation to obtain a position-encoded feature map, ensuring appropriate adjustment of the output of the convolutional layer to capture the time relationships present in the data. Specifically, the position embedding is added to the output of the convolutional layer, and the formula is as follows: ; where PE(X) represents the position-encoded feature map, X is the output of the convolutional layer, is the position encoding. This explicit encoding helps the model understand the order and dependencies of the features in the sequence, which is a key factor for time series analysis and signal processing tasks; Finally, the position-encoded feature map is input into the Transformer encoder. The Transformer encoder uses the multi-head self-attention mechanism to capture the global dependencies in the entire sequence; simulating the long-distance relationships between features enables the model to understand the context and global patterns crucial for accurate classification. Specifically, the operation is as follows: head i =Attention(W i Q X1,W i K X1,W i V X1); where X1 is the input, i.e., the position-encoded feature map, and W i Qis the query matrix of the i-th attention head, which is used to map the input feature map into a query vector; W i K is the key matrix of the i-th attention head, which is used to map the input feature map into a key vector; W i V is the value matrix of the i-th attention head, which is used to map the input feature map into a value vector; the output of the final multi-head self-attention mechanism is generated by h attention heads, and the output is concatenated and linearly transformed as shown in the following formula: MHA(X1)=W O (head1||... || head h ); where MHA(X1) represents the output of the multi-head self-attention mechanism, and W O is the output weight matrix; the combined output of the multi-head self-attention mechanism is mapped to ensure compliance with the final output dimension of the Transformer model; the attention mechanism enables the model to selectively focus on important features at different positions in the sequence, thereby more effectively simulating global context and dependencies; the feed-forward neural network layer is used to further refine the feature representation; After being processed by the Transformer encoder, through a one-dimensional global average pooling layer, the time dimension of the feature map is compressed into a feature vector of a fixed length; finally, the feature vector is fed into a fully connected layer and a softmax layer for classification, and after training, the trained RFF-CAT model, that is, the radio frequency fingerprint recognition model, is obtained; the design of the RFF-CAT model ensures independent processing of local and global feature extraction; the convolutional layer is specifically used to capture local features, while the Transformer encoder is good at simulating global dependencies; by separating these tasks, the model reduces the computational complexity, improves flexibility, and improves performance in a noisy environment; compared with the traditional Conformer model (which usually fuses CNN and Transformer operations into one module), this separation also makes the model architecture simpler and easier to interpret; wherein, the training of the RFF-CAT model adopts the following process, including: 1) Use the RMSProp optimization algorithm to minimize the cross-entropy loss function; 2) Batch process the training data and perform multiple rounds of iterative training until convergence; 3) During the training process, introduce a learning rate scheduling strategy to improve the training efficiency.
[0010] Preferably according to the present invention, use the trained radio frequency fingerprint recognition model for prediction, and combine the multi-pack inference method to obtain the final prediction result; including: In the prediction and inference stage, for the RF signal data from unknown devices, it is first preprocessed and converted into a channel-independent spectrogram, and then the spectrogram is input into the trained RFF-CAT model; the output of the Softmax layer in the RFF-CAT model is a probability vector, where each element represents the confidence associated with one of the devices, and the probability vector indicates the likelihood that the data packet belongs to each potential device, thus achieving accurate and reliable device identification; The output obtained from the RFF-CAT model is then passed through a multi-packet inference method, namely the adaptive fusion method (MPF), to obtain the final output probability. The multi-packet inference method (MPF) includes: The incoming RF signal data is divided into multiple data packets, and the initial prediction value, that is, the initial prediction value obtained through the RFF-CAT model, is calculated for each data packet. An initial weight set is defined, and then a weighted sum is performed to obtain the fused prediction probability, as shown below: ; where p represents the fused prediction probability, represents the prediction probability of the nth data packet, that is, the initial prediction value; W n represents the weight of the prediction probability of the nth data packet; through the above formula, the various contributions of different data packets can be quantified, and different weights are set for the data packets; Calculate the error between the initial prediction value and the fused prediction probability; update the weights to adapt to the prediction values of each data packet, usually following the update rules of the error balancing algorithm; at each iteration, the update rule of the weights is shown in the following formula: ; where, represents the array of weights of the prediction probabilities of the data packets with index n; represents the array of weights after updating the weights of index n, preparing for the next iteration; represents the learning rate, which controls the magnitude of the weight adjustment in each iteration; represents the error between the initial prediction probability, that is, the initial prediction value and the fused probability after weighted summation, reflecting the difference between the initial prediction probability and the prediction probability after weighted summation; is an array including the initial prediction probabilities of the first N data packets starting from index n; Use the updated weights to perform a weighted sum on the prediction values of each data packet again to obtain a new fused prediction probability, that is, identify the RF fingerprint of the corresponding device. The model not only outputs the identity of the device, but also, based on the uniqueness of the RF signal, can distinguish different devices, realizing the authentication and identification of the devices.
[0011] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a radio frequency fingerprint recognition method based on a convolution-attention mechanism and multi-pack inference are implemented.
[0012] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of a radio frequency fingerprint recognition method based on a convolution-attention mechanism and multi-pack inference are implemented.
[0013] The second aspect of the present invention provides a radio frequency fingerprint recognition system based on a convolution-attention mechanism and multi-pack inference, including: A preprocessing module configured to: capture a device transmission signal and preprocess the transmission signal to obtain a spectrogram; A model training module configured to: construct a radio frequency fingerprint recognition model, input the obtained spectrogram into the radio frequency fingerprint recognition model for training, and obtain a trained radio frequency fingerprint recognition model; The radio frequency fingerprint recognition model includes a convolutional layer, a position encoding layer, and a Transformer encoder; The convolutional layer converts the input feature map into a high-order feature representation; The position encoding layer is used to add position information; The Transformer encoder includes a multi-head self-attention mechanism and a feed-forward neural network layer. The multi-head self-attention mechanism captures global dependencies in the entire sequence, and the feed-forward neural network layer further enhances the feature expression ability; A multi-pack inference module configured to: use the trained radio frequency fingerprint recognition model for prediction and combine a multi-pack inference method to obtain a final prediction result.
[0014] The beneficial effects of the present invention are as follows: 1. By modeling radio frequency fingerprint recognition as a deep learning problem based on a convolutional layer and a Transformer encoder and proposing a multi-pack adaptive fusion method, the present invention effectively improves the feature expression and classification accuracy of radio frequency signals by combining the feature extraction capabilities of the convolutional layer and the Transformer during the model training process.
[0015] 2. RFF-CAT shows higher accuracy, stability, and robustness when processing complex spatio-temporal data and high-dimensional feature maps. It is a deep learning model with broad application prospects. The present invention performs excellently in a low signal-to-noise ratio environment, can significantly improve the overall performance of a distributed radio frequency fingerprint recognition system, and enhance the recognition accuracy and robustness of devices. Description of the Drawings
[0016] Figure 1Schematic diagram of the system model and working process of radio frequency fingerprint recognition of the present invention; Figure 2 Schematic diagram of the RFF-CAT model of the present invention; Figure 3 Time domain diagram of the dataset used in the experiment of the present invention; Figure 3 In (a), it is the time domain diagram of the dataset of the LoRa device collected when the spreading factor of the LoRa device is set to 7; Figure 3 In (b), it is the time domain diagram of the dataset of the LoRa device collected when the spreading factor of the LoRa device is set to 8; Figure 3 In (c), it is the time domain diagram of the dataset of the LoRa device collected when the spreading factor of the LoRa device is set to 9; Figure 4 Spectrum diagram generated after preprocessing the dataset used in the present invention; Figure 5 Overall comparison diagram of the prediction performance models of RFF-CAT and the baseline; Figure 6 Schematic diagram of the prediction performance comparison of different signal-to-noise ratios under the MPF condition and the inference based on the average value of the present invention; Figure 6 In (a), it is the accuracy schematic diagram of using MPF and not using MPF when the signal-to-noise ratio is 0; Figure 6 In (b), it is the accuracy schematic diagram of using MPF and not using MPF when the signal-to-noise ratio is 5; Figure 6 In (c), it is the accuracy schematic diagram of using MPF and not using MPF when the signal-to-noise ratio is 10; Figure 6 In (d), it is the accuracy schematic diagram of using MPF and not using MPF when the signal-to-noise ratio is 20; Figure 7 Classification schematic diagram between the MPF method (weighted sum) used in the present invention and the traditional average-based method. Specific implementation manner
[0017] The present invention will be further described below through examples in conjunction with the drawings, but not limited thereto.
[0018] Example 1 Radio frequency fingerprint recognition method based on convolutional-attention mechanism and multi-packet inference, as Figure 1 shown, includes: Step 1, capture the device transmission signal and preprocess the transmission signal to obtain a spectrum diagram; Step 2, construct a radio frequency fingerprint recognition model, input the obtained spectrum diagram into the radio frequency fingerprint recognition model for training, and obtain a trained radio frequency fingerprint recognition model; The radio frequency fingerprint recognition model includes a convolutional layer, a position encoding layer, and a Transformer encoder; The convolutional layer transforms the input feature map into a high-order feature representation; The position encoding layer is used to add position information (absolute or relative position) to enable the model to have the ability to process sequential information; The Transformer encoder includes a multi-head self-attention mechanism and a feed-forward neural network layer (FFN). The multi-head self-attention mechanism captures global dependencies in the entire sequence, and the feed-forward neural network layer further enhances the feature expression ability; Step 3: Use the trained radio frequency fingerprint recognition model for prediction, and combine the multi-packet inference method to obtain the final prediction result, realizing the authentication and recognition of the device.
[0019] Embodiment 2 The radio frequency fingerprint recognition method based on the convolutional-attention mechanism and multi-packet inference according to Embodiment 1 is characterized in that: Capture the device transmission signal, and the time-domain diagram of the transmission signal is as Figure 3 shown, Figure 3 where (a) in is the time-domain diagram of the dataset of the LoRa device collected when the spreading factor of the LoRa device is set to 7; Figure 3 where (b) in is the time-domain diagram of the dataset of the LoRa device collected when the spreading factor of the LoRa device is set to 8; Figure 3 where (c) in is the time-domain diagram of the dataset of the LoRa device collected when the spreading factor of the LoRa device is set to 9, and preprocess the transmission signal to obtain a spectrogram, including: First, collect the transmission signal from the LoRa device and record it as r(n). The LoRa device adopts the LoRa protocol; A LoRa device refers to a terminal device that uses LoRa (Long Range) modulation technology for wireless communication; such devices usually have the ability of low power consumption, long distance, and low rate data transmission, and are widely used in Internet of Things scenarios such as smart metering, environmental monitoring, agricultural Internet of Things, and smart city; LoRa devices include sensor nodes, gateways, modules, etc., and can operate for a long time without frequent charging, and are suitable for deployment in application environments sensitive to energy consumption; The LoRa protocol usually refers to the LoRaWAN (LoRa Wide Area Network) protocol, which is an open low-power wide-area network communication protocol based on LoRa physical layer technology; LoRaWAN adopts a star topology structure, supports mechanisms such as device hierarchical management, security authentication, and end-to-end encryption, and has advantages such as low power consumption, long distance, and large connection number, and is especially suitable for application scenarios with remote, intermittent, and small data volume transmission; Then, calculate the carrier frequency offset (CFO) as shown in the following formula: ; where arg(⋅) represents the phase of the transmitted signal, n represents the time variable, represents the carrier frequency offset; Compensate for the carrier frequency offset (CFO) to reduce the error caused by the carrier frequency offset, as shown in the following formula: ; where, is the received signal, is the signal after compensation, j represents the imaginary unit, CFO represents the carrier frequency offset, i.e., , and t represents the time variable; Adopt data normalization to normalize the compensated signal to unit amplitude, eliminate the influence of the difference in transmission power between devices on classification, as shown in the following formula: ; where r norm (n) represents the normalized received signal, whose amplitude is limited within the range of [-1, 1], used to unify the energy scale of signals from different devices, enhance the sensitivity of the model to the structural features in the device fingerprint, and reduce the interference of the transmission power difference on the classification result; Finally, perform short-time Fourier transform (STFT) to convert it into a channel-independent spectrogram to mitigate the channel impact, as shown in the following formula: ; where, represents the spectrogram, represents a matrix window of length N, is the sliding window function, specifically represented as a rectangular window function with length N = 64, selects a segment of signal at time point n to each moment t for transformation, j represents the imaginary unit, and f represents the frequency.
[0020] Construct a radio frequency fingerprint recognition model, input the obtained spectrogram into the radio frequency fingerprint recognition model for training to obtain a trained radio frequency fingerprint recognition model; as Figure 2 shown, including: The convolutional layer includes a normalization layer, a point convolutional layer, a GLU activation function layer, a depthwise separable convolutional layer, a batch normalization layer, a Swish activation function layer, a concatenated convolutional layer, and a Dropout layer; First, the obtained spectrogram is input into the convolutional layer. The input spectrogram is normalized so that its mean is 0 and variance is 1 to improve training stability, obtaining more uniformly distributed spectral features and retaining the normalized output of the original structure. The channel dimension is adjusted using 1x1 convolution through the pointwise convolutional layer to fuse or expand feature information, obtaining a feature map with adjusted channel numbers. The key information of the feature map is enhanced through the GLU activation function layer, and the redundant information of the feature map is suppressed. The number of parameters of the feature map is reduced through the depthwise separable convolutional layer to enhance spatial features. A feature map with stable distribution is obtained through the batch normalization layer. An activated feature map is obtained through the Swish activation function layer, retaining the small gradient information in the negative value region. The final feature representation with optimized channel numbers is obtained through the concatenated convolutional layer. Finally, a regularized feature map is obtained through the Dropout layer This module extracts low-level local features in the signal, such as the edges, textures, etc. of the spectrogram, which are helpful for subsequent pattern recognition of the signal; Then, in order to retain the time and order information, a position encoding layer is introduced after the convolutional layer operation to obtain a feature map with position encoding, ensuring appropriate adjustment of the output of the convolutional layer and capturing the time relationships existing in the data. Specifically, the position embedding is added to the output of the convolutional layer, and the formula is as follows: ; where PE(X) represents the feature map with position encoding, X is the output of the convolutional layer, is the position encoding. This explicit encoding helps the model understand the order and dependencies of the features in the sequence, which is a key factor for time series analysis and signal processing tasks; Finally, the feature map with position encoding is input into the Transformer encoder, i.e., the stacked multiple modules. The Transformer encoder uses the multi-head self-attention mechanism to capture the global dependencies in the entire sequence; simulates the long-distance relationships between features, enabling the model to understand the context and global patterns crucial for accurate classification. Specifically, the operation is as follows: head i =Attention(W i Q X1,W i K X1,W i V X1); where X1 is the input, i.e., the feature map with position encoding, W i Q is the query matrix of the i-th attention head, which is used to map the input features into query vectors; W i Kis the key matrix of the i-th attention head, which is used to map the input features into key vectors; W i V is the value matrix of the i-th attention head, which is used to map the input features into value vectors; the h attention heads generate the final output of the multi-head self-attention mechanism, and the output is concatenated and linearly transformed as shown in the following formula: MHA(X1)=W O (head1||... || head h ); Among them, MHA(X1) represents the output of the multi-head self-attention mechanism, and W O is the output weight matrix; it maps the combined output of the multi-head self-attention mechanism to ensure that it conforms to the final output dimension of the Transformer model; the attention mechanism enables the model to selectively focus on important features at different positions in the sequence, thereby more effectively simulating global context and dependencies; the feed-forward neural network layer is used to further improve the feature representation; After being processed by the Transformer encoder, it passes through a one-dimensional global average pooling layer to compress the time dimension of the feature map into a feature vector of a fixed length; finally, the feature vector is fed into a fully connected layer and a softmax layer for classification, and after training, the trained RFF-CAT model, that is, the radio frequency fingerprint recognition model, is obtained; the design of the RFF-CAT model ensures the independent processing of local and global feature extraction; the convolutional layer is specifically used to capture local features, while the Transformer encoder is good at simulating global dependencies; by separating these tasks, the model reduces the computational complexity, improves the flexibility, and improves the performance in a noisy environment; compared with the traditional Conformer model (which usually fuses CNN and Transformer operations into one module), this separation also makes the model architecture simpler and easier to interpret; Among them, the training of the RFF-CAT model adopts the following process, including: 1) Use the RMSProp optimization algorithm to minimize the cross-entropy loss function; 2) Batch process the training data and perform multiple rounds of iterative training until convergence; 3) During the training process, introduce a learning rate scheduling strategy to improve the training efficiency.
[0021] Use the trained radio frequency fingerprint recognition model for prediction and combine the multi-pack inference method to obtain the final prediction result; including: In the prediction and inference stage, for the radio frequency signal data from an unknown device, first perform preprocessing and convert it into a channel-independent spectrogram, such as Figure 4As shown, the spectrogram is then input into the trained RFF-CAT model; the output of the Softmax layer in the RFF-CAT model is a probability vector, where each element represents the confidence associated with one of the devices, and the probability vector represents the likelihood that the data packet belongs to each potential device, thus achieving accurate and reliable device identification; The output obtained from the RFF-CAT model is then passed through a multi-packet inference method, namely the adaptive fusion method (MPF), to obtain the final output probability. The multi-packet inference method (MPF) includes: The incoming RF signal data is divided into multiple data packets, and the initial prediction value for each data packet, i.e., the initial prediction value obtained through the RFF-CAT model, is calculated. An initial weight set is defined, and then a weighted sum is performed to obtain the fusion prediction probability, as shown below: ; where p represents the fusion prediction probability, represents the prediction probability of the nth data packet, i.e., the initial prediction value; represents the weight of the prediction probability of the nth data packet; through the above formula, the various contributions of different data packets can be quantified, and different weights are set for the data packets; Calculate the error between the initial prediction value and the fusion prediction probability; update the weights to adapt to the prediction values of each data packet, usually following the update rules of the error balancing algorithm; at each iteration, the update rule of the weights is shown in the following formula: ; where w n represents the array of weights of the prediction probabilities of the data packets with index n; w n+1 represents the array of weights after updating the weights of index n, preparing for the next iteration; represents the learning rate, which controls the magnitude of the weight adjustment in each iteration; e n represents the error between the initial prediction probability, i.e., the initial prediction value, and the fusion probability after weighted summation, reflecting the difference between the initial prediction probability and the prediction probability after weighted summation; p n is an array including the initial prediction probabilities of the first N data packets starting from index n.
[0022] The prediction values of each data packet are weighted and summed again using the updated weights to obtain a new fusion prediction probability, i.e., the RF fingerprint of the corresponding device is identified. The model not only outputs the identity of the device but also, based on the uniqueness of the RF signal, can distinguish different devices, achieving device authentication and identification.
[0023] After the entire model training is completed, it also includes: classifying the signal samples of unknown devices using the finally trained model; evaluating the accuracy, recall rate, and F1 value of the model based on the classification results; storing the evaluation results and periodically updating the model to adapt to the feature changes of new devices.
[0024] The results of the proposed radio frequency fingerprint recognition and multi-packet fusion method based on convolutional neural network and Transformer encoder are as Figure 5 , Figure 6 , Figure 7 shown; Figure 5 The classification accuracies of the RFF-CAT model are compared with those of the Transformer model, convolutional neural network model, and Transformer encoder model at different signal-to-noise ratio levels, and these classification accuracies are obtained without using MPF; as shown in the figure, at low SNR levels, the classification accuracy obtained using the RFF-CAT model is about 10% higher than that obtained using the Transformer model. At high signal-to-noise ratio levels, the classification accuracy of the RFF-CAT model is about 4 - 5% higher than that of the Transformer model. In addition, the classification accuracy obtained using the RFF-CAT model is generally about 10% higher than that obtained using the Transformer model. RFF-CAT uses convolutional operations to capture local patterns and features, which is particularly useful at low signal-to-noise ratio levels. Through simulations using the proposed RFF-CAT training model, the experimental results obtained show that the model exhibits excellent performance. Figure 6 In (a), it is a schematic diagram of the accuracy with and without using MPF when the signal-to-noise ratio is 0; Figure 6 In (b), it is a schematic diagram of the accuracy with and without using MPF when the signal-to-noise ratio is 5; Figure 6 In (c), it is a schematic diagram of the accuracy with and without using MPF when the signal-to-noise ratio is 10; Figure 6 In (d), it is a schematic diagram of the accuracy with and without using MPF when the signal-to-noise ratio is 20; It can be seen from Figure 6 and Figure 7 that the invented MPF method can significantly improve the RFF performance, especially in low signal-to-noise ratio scenarios.
[0025] Example 3 A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the radio frequency fingerprint recognition method based on convolutional-attention mechanism and multi-packet inference described in Example 1 or 2.
[0026] Example 4 A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the radio frequency fingerprint recognition method based on the convolutional-attention mechanism and multi-pack inference described in Embodiment 1 or 2 are implemented.
[0027] Embodiment 5 A preprocessing module is configured to: capture a device transmission signal and preprocess the transmission signal to obtain a spectrogram; A model training module is configured to: construct a radio frequency fingerprint recognition model, input the obtained spectrogram into the radio frequency fingerprint recognition model for training, and obtain a trained radio frequency fingerprint recognition model; The radio frequency fingerprint recognition model includes a convolutional layer, a position encoding layer, and a Transformer encoder; The convolutional layer converts the input feature map into a high-order feature representation; The position encoding layer is used to add position information; The Transformer encoder includes a multi-head self-attention mechanism and a feed-forward neural network layer. The multi-head self-attention mechanism captures global dependencies in the entire sequence, and the feed-forward neural network layer further enhances the feature expression ability; A multi-pack inference module is configured to: use the trained radio frequency fingerprint recognition model for prediction and combine the multi-pack inference method to obtain a final prediction result.
Claims
1. A radio frequency fingerprint recognition method based on a convolution-attention mechanism and multi-packet reasoning, characterized in that Including: Step 1: The capture device transmits a signal and preprocesses the transmitted signal to obtain a spectrogram. Step 2: Construct a radio frequency fingerprint recognition model, input the obtained spectrogram into the radio frequency fingerprint recognition model for training, and obtain a trained radio frequency fingerprint recognition model. The radio frequency fingerprint recognition model includes a convolutional layer, a position encoding layer, and a Transformer encoder. The convolutional layer converts the input feature map into a higher-order feature representation. The position encoding layer is used to add position information. The Transformer encoder includes a multi-head self-attention mechanism and a feed-forward neural network layer. The multi-head self-attention mechanism captures the global dependencies in the entire sequence, and the feed-forward neural network layer further enhances the feature expression ability. Step 3: Use the trained radio frequency fingerprint recognition model for prediction and combine it with the multi-pack inference method to obtain the final prediction result.
2. The radio frequency fingerprint recognition method based on convolutional-attention mechanism and multi-packet inference according to claim 1, wherein The capture device transmits a signal and preprocesses the transmitted signal to obtain a spectrogram. Including: First, collect the transmitted signal from the LoRa device and record it as r(n). The LoRa device uses the LoRa protocol. Then, calculate the carrier frequency offset as shown in the following formula: ; where arg(⋅) represents the phase of the transmission signal, and n represents the time variable, represents the carrier frequency offset; Compensate for the carrier frequency offset as shown in the following formula: ; Among them, is the received signal, is the signal after compensation, j represents the imaginary unit, and CFO represents the carrier frequency offset, that is , and t represents the time variable; Adopt data normalization to normalize the compensated signal to unit amplitude as shown in the following formula: ; where r norm (n) represents the normalized received signal; Finally, perform short-time Fourier transform to convert it into a spectrogram independent of the channel as shown in the following formula: ; Among them, represents a spectrogram, represents a matrix window of length N, and a section of signal is selected for transformation at each moment from time point n to t. j represents the imaginary unit, and f represents the frequency.
3. The radio frequency fingerprint recognition method based on convolutional-attention mechanism and multi-packet inference according to claim 2, wherein Construct a radio frequency fingerprint recognition model, input the obtained spectrogram into the radio frequency fingerprint recognition model for training, and obtain a trained radio frequency fingerprint recognition model. Including: The convolutional layer includes a normalization layer, a point convolutional layer, a GLU activation function layer, a depthwise separable convolutional layer, a batch normalization layer, a Swish activation function layer, a concatenated convolutional layer, and a Dropout layer. First, the obtained spectrogram is input into the convolutional layer, and the input spectrogram is normalized; the channel dimension is adjusted using 1x1 convolution through the point convolutional layer to obtain the feature map with adjusted number of channels; the key information of the feature map is enhanced and the redundant information of the feature map is suppressed through the GLU activation function layer; the number of parameters of the feature map is reduced and the spatial features are enhanced through the depthwise separable convolutional layer; the feature map with stable distribution is obtained through the batch normalization layer; the activated feature map is obtained through the Swish activation function layer, and the tiny gradient information in the negative value region is retained; the final feature representation with optimized number of channels is obtained through the concatenated convolutional layer; finally, the regularized feature map is obtained through the Dropout layer ; Then, introduce a position encoding layer after the convolutional layer operation to obtain a position-encoded feature map and capture the time relationship existing in the data. The formula is as follows: ; Among them, PE(X) represents the feature map of the positional encoding, where X is the output of the convolutional layer, is the positional encoding; Finally, input the position-encoded feature map into the Transformer encoder. The Transformer encoder uses the multi-head self-attention mechanism to capture the global dependencies in the entire sequence. The operation is as follows: head i =Attention(W i Q X1,W i K X1,W i V X1); Among them, X1 is the feature map of the input, i.e., the positional encoding, and W i Q is the query matrix of the i-th attention head; W i K is the key matrix of the i-th attention head; W i V is the value matrix of the i-th attention head; the h attention heads generate the output of the final multi-head self-attention mechanism, and the output is concatenated and linearly transformed as shown in the following formula: MHA(X1)=W O (head1||...||head h ); Among them, MHA(X1) represents the output of the multi-head self-attention mechanism, and W O is the output weight matrix; map the combined output of the multi-head self-attention mechanism and further improve the feature representation using a feed-forward neural network layer; After being processed by the Transformer encoder, it passes through a one-dimensional global average pooling layer to compress the time dimension of the feature map into a feature vector of a fixed length. Finally, the feature vector is sent to a fully connected layer and a softmax layer for classification. After training, the trained RFF-CAT model, that is, the radio frequency fingerprint recognition model, is obtained.
4. The radio frequency fingerprint recognition method based on convolution-attention mechanism and multi-packet inference according to claim 3, wherein Use the trained radio frequency fingerprint recognition model for prediction and combine it with the multi-pack inference method to obtain the final prediction result. Including: In the prediction and inference stage, for the radio frequency signal data from an unknown device, first preprocess it and convert it into a spectrogram independent of the channel, and then input the spectrogram into the trained RFF-CAT model. The output obtained from the RFF-CAT model is then passed through the multi-pack inference method, that is, the adaptive fusion method, to obtain the final output probability. The multi-pack inference method includes: The incoming radio frequency signal data is divided into multiple data packets. For each data packet, an initial prediction value, i.e., the initial prediction value obtained through the RFF-CAT model, is calculated. An initial weight set is defined, and then weighted summation is performed to obtain the fused prediction probability, as follows: ; where p represents the fusion prediction probability, represents the prediction probability of the nth data packet, i.e., the initial prediction value; represents the weight of the prediction probability of the nth data packet; Calculate the error between the initial prediction value and the fused prediction probability, and update the weights to adapt to the prediction values of each data packet; At each iteration, the weight update rule is shown in the following formula: ; Among them, represents the array of predicted probability weights for the data packet with index n; represents the weight array after updating the weight of index n, represents the learning rate, which controls the amplitude of weight adjustment in each iteration; represents the error between the initial prediction probability, i.e., the initial prediction value and the fused probability after weighted summation, is an array including the initial prediction probabilities of the first N data packets starting from index n; Use the updated weights to perform weighted summation on the prediction values of each data packet again to obtain a new fused prediction probability, that is, identify the radio frequency fingerprint of the corresponding device, and realize the authentication and identification of the device.
5. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the radio frequency fingerprint recognition method based on convolutional-attention mechanism and multi-packet inference according to any one of claims 1-4.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the radio frequency fingerprint recognition method based on convolutional-attention mechanism and multi-packet inference according to any one of claims 1-4.
7. A radio frequency fingerprint recognition system based on a convolutional-attention mechanism and multi-packet inference, characterized in that Including: A preprocessing module, configured to: capture the device transmission signal and preprocess the transmission signal to obtain a spectrogram; A model training module, configured to: construct a radio frequency fingerprint recognition model, input the obtained spectrogram into the radio frequency fingerprint recognition model for training, and obtain a trained radio frequency fingerprint recognition model; The radio frequency fingerprint recognition model includes a convolutional layer, a position encoding layer, and a Transformer encoder; The convolutional layer converts the input feature map into a higher-order feature representation; The position encoding layer is used to add position information; The Transformer encoder includes a multi-head self-attention mechanism and a feed-forward neural network layer. The multi-head self-attention mechanism captures the global dependencies in the entire sequence, and the feed-forward neural network layer further enhances the feature expression ability; A multi-packet inference module, configured to: use the trained radio frequency fingerprint recognition model for prediction and combine the multi-packet inference method to obtain the final prediction result.
Citation Information
Patent Citations
Multi-feature fusion wireless device radio frequency fingerprint extraction method based on attention mechanism
CN114118131A
Deep learning-based radio frequency fingerprint identification method for frequency equipment
CN114896887A
Robust radio frequency fingerprint identification method based on cross attention
CN118612744A
Method for re-recognizing object image based on multi-feature information capture and correlation analysis
US20220415027A1
Cited By
Industrial vehicle intelligent instrument control system and method based on multi-mode biological recognition
CN120526488A
Multi-sound-source direction-of-arrival estimation model training method, multi-sound-source direction-of-arrival estimation method, equipment, medium and product
CN120910566A
A training method for a multi-source direction-of-arrival (DOA) estimation model, a multi-source DOA estimation method, equipment, medium, and product.
CN120910566B
Unmanned aerial vehicle detection and classification method based on deep learning
CN121919560A
Deep learning-based drone detection and classification methods
CN121919560B