Physical layer authentication scheme based on sparse auto-encoder and knowledge distillation
By using sparse autoencoder to reduce data dimensionality in the physical layer certification scheme, and distilling the teacher model to the student model in combination with knowledge distillation technology, the efficiency and lightweight problems of physical layer certification in the existing technology are solved, and efficient and lightweight certification performance is achieved.
Patent Information
- Application Number
- CN202411878364.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-19
AI Technical Summary
The existing physical layer authentication schemes are difficult to achieve efficient and lightweight authentication in resource-constrained environments, resulting in high authentication delays and difficult to deploy in scenarios such as the Internet of Things.
The physical layer certification scheme based on sparse autoencoder and knowledge distillation is adopted to reduce data dimensionality through sparse autoencoder, and multiple large-scale teacher models are distilled to lightweight student models using knowledge distillation to reduce model complexity.
While ensuring authentication accuracy, it significantly reduces the complexity and computing burden of the model, improves training speed and resource utilization efficiency, and is suitable for resource-constrained IoT environments.
Smart Images

Figure SMS_1 
Figure SMS_3 
Figure SMS_5
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication security, and in particular to physical layer authentication, and specifically to a physical layer authentication scheme based on sparse autoencoder and knowledge distillation. Background Art
[0002] With the continuous development of wireless communication technology, the security challenges faced by wireless communication systems are also increasing. The physical layer is the bottom layer responsible for signal transmission in the communication system. The attack methods against the physical layer are constantly evolving, and the traditional key-based authentication method has gradually shown its limitations. Physical layer authentication uses the inherent properties of the wireless channel to authenticate identity. Compared with key-based authentication technology, it has the advantages of high security and low computational complexity, and has become a key technology in the current authentication field. Deep learning can significantly improve the accuracy of physical layer authentication, and physical layer authentication technology based on deep learning has attracted the attention of more and more researchers. However, with the rapid development of the Internet of Things, more and more devices need to perform complex authentication tasks in resource-constrained environments. The existing high-performance authentication schemes have a large number of parameters, resulting in high authentication delays and difficulty in direct deployment in resource-constrained environments such as the Internet of Things. Therefore, how to make the authentication model more lightweight and efficient while ensuring security authentication performance has become a hot issue in the field of physical layer security authentication.
[0003] For the physical layer authentication problem in resource-constrained scenarios, the solutions can be mainly divided into two categories: optimization solutions based on algorithm complexity and data compression, and distributed authentication solutions. The optimization solutions based on algorithm complexity and data compression reduce the computational and storage requirements of the model through algorithm-level optimization, mainly using lightweight algorithms, simplifying model structures, and using data compression technology to reduce data dimensions. Chen et al. proposed a physical layer authentication scheme based on convolutional denoising autoencoder (CDAE). The scheme uses CDAE and weighted-NN algorithm for feature extraction and classification, which can effectively eliminate noise and extract key features in CSI data. However, when the extracted features are not discriminable enough, the authentication scheme combined with the-NN algorithm will cause a significant decrease in accuracy. Senigagliesi L et al. reduced the number of authentication features through data compression technology. Experiments show that the three data compression technologies of principal component analysis, random sampling, and t-distributed random neighborhood embedding can significantly reduce the computational burden of authentication. However, when the attacker is close to the legitimate communication participants, the information loss caused by information compression causes a decrease in authentication accuracy. Distributed authentication schemes improve the efficiency and reliability of the authentication process by sharing authentication tasks among multiple nodes and utilizing network collaboration capabilities. Zhang et al. proposed a cooperative physical layer authentication scheme COPLA, which uses multiple supervisory nodes to share tasks and provides a lightweight deployment and high authentication performance solution in complex wireless environments. However, compared with other physical layer authentication schemes, COPLA is not suitable for physical layer devices in isolated environments, which limits its wide application in real scenarios. Zhao et al. proposed a physical layer authentication strategy for resource-constrained underwater acoustic networks (UANs). This strategy uses the support vector machine (SVM) algorithm to perform distributed authentication on a small training set, which can improve the node authentication accuracy while reducing resource energy consumption. However, this scheme is only applicable to underwater wireless acoustic networks and has poor generalization ability in other physical layer authentication scenarios. Summary of the invention
[0004] The purpose of the present invention is to propose a physical layer authentication scheme based on sparse autoencoder and knowledge distillation to solve the problem of limited wireless device resources.
[0005] The specific method is as follows:
[0006] The physical layer authentication scheme based on sparse autoencoder and knowledge distillation includes the following steps:
[0007] Step 1: Data collection and preprocessing: Use the Linux 802.11n CSI Tool to collect CSI data in real communication scenarios, then calculate the amplitude and phase of each subcarrier based on the subcarrier complex matrix, then use the Savitzky-Golay filter to denoise the CSI data, and finally normalize the data for subsequent analysis and processing.
[0008] Step 2: Perform data enhancement. Use a generative adversarial network based on Sensitivity-Specificity Loss to generate high-quality experimental data and improve the model's ability to recognize negative samples.
[0009] Step 3: Perform data dimensionality reduction. Use a sparse autoencoder based on the vulture optimization algorithm to reconstruct the input data. The vulture optimization algorithm can automatically find the hyperparameters of the sparse autoencoder and improve the performance of the sparse autoencoder.
[0010] Step 4: Use knowledge distillation strategy to train the model. First, use a large amount of labeled data to train multiple teacher models and perform complex security authentication task training. Then, in order to verify the effectiveness of the teacher model in practical applications, the performance of each teacher model is evaluated on an independent test set and the model is screened to ensure that it can provide high-accuracy security authentication in a variety of scenarios. Finally, the adaptive weighted multi-teacher knowledge distillation model AMTDF is used to distill the large teacher model into a lightweight student model, so that the student model can better simulate the output distribution of the teacher model and promote the effective transfer of complex knowledge.
[0011] The positive effects of the present invention are:
[0012] This scheme uses knowledge distillation to distill multiple large teacher models into a lightweight student model, effectively reducing the complexity of the model while ensuring the certification accuracy. Specifically, the present invention first proposes a generative adversarial network data enhancement model SensSpec-GAN based on Sensitivity-Specificity Loss, which enhances the diversity of data and the model's ability to recognize negative samples by introducing sensitivity and specificity constraints in the generative adversarial network. Subsequently, in response to the computational burden problem caused by the high dimension of CSI data, the present invention proposes a sparse autoencoder model SABES based on the vulture optimization algorithm to reduce the dimension of CSI data. SABES uses the vulture optimization algorithm to automatically adjust the model hyperparameters and enhance the model's feature expression ability. Finally, the present invention proposes an adaptive weighted multi-teacher knowledge distillation model AMTDF, which adjusts the contribution of different teacher models to the student model through an adaptive weight mechanism, thereby improving the performance of the model on the target task. Experimental results show that the accuracy of the scheme reaches 96.32%, and the training speed is 2.46 times faster than the baseline algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 This is the algorithm structure of SensSpec-GAN.
[0014] Figure 2 It is the performance indicator of different algorithms under 30% data enhancement ratio.
[0015] Figure 3 It is the impact of different dimensionality reduction techniques on the performance indicators.
[0016] Figure 4 This is the overall architecture diagram of AMTDF.
[0017] Figure 5 is the relationship between the number of teacher models and the accuracy of the student model. DETAILED DESCRIPTION
[0018] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings.
[0019] Step 1: Data collection and preprocessing: Use the Linux 802.11n CSI Tool to collect CSI data in real communication scenarios, then calculate the amplitude and phase of each subcarrier based on the subcarrier complex matrix, then use the Savitzky-Golay filter to denoise the CSI data, and finally normalize the data for subsequent analysis and processing.
[0020] CSI data is collected using three ThinkPad X201 laptops with Intel 5300 network cards. The three computers represent the sender Alice, the receiver Bob, and the attacker Eve in the communication system. The Ubuntu operating system with system version 14.06LTS and kernel version 4.4 is used as the information collection platform, and the Linux 802.11n CSITool toolkit is used as the CSI data collection tool. The received data type is a complex matrix with dimension (T, R, N), where T and R represent the number of antennas at the sender and receiver, respectively, and N represents the number of subcarriers between each pair of antennas in the CSI data.
[0021] To further improve the quality of CSI data, the present invention uses a Savitzky-Golay (SG) filter to denoise the CSI data. The SG filter smoothes the data by fitting a local polynomial to the data points in a sliding window, and replaces the original data points with the center value of the polynomial window. The SG filter can effectively remove random noise from the CSI data while retaining important features such as the peak value, valley value, and waveform trend of the signal. The SG filtering process can be expressed by the following formula:
[0022]
[0023] In the formula, is the filtered value, y i+j is the value of the i+jth point in the original data sequence, c j is the pre-calculated convolution coefficient, and m is the window radius. By adjusting the window size and polynomial order, the smoothness of the filter and the accuracy of the fit can be controlled.
[0024] Min-Max Normalization is a commonly used data preprocessing method, which scales the data on each feature dimension so that all feature values are limited to a fixed range. For each input data T in the lth dimension, l , the process of scaling it to [0,1] through maximum and minimum normalization can be expressed as:
[0025]
[0026] Where T l Represents the original data value, min(T l ) and max(T l ) represents the minimum and maximum values in the l-th dimension data, Represents the normalized result.
[0027] Step 2: Perform data enhancement. Use a generative adversarial network based on Sensitivity-Specificity Loss to generate high-quality experimental data and improve the model's ability to recognize negative samples.
[0028] SensSpec-GAN uses a fully connected network with two hidden layers as a generator. In order to fully utilize the sequence characteristics in CSI data, a deep neural network based on Sensitivity-Specificity Loss and cross entropy loss is used as a discriminator. The algorithm structure is as follows: Figure 1 shown.
[0029] Sensitivity-Specificity Loss is defined as follows:
[0030] L ss =w·sensitivity+(1-w)·specificity
[0031] In the formula
[0032]
[0033]
[0034] In SensSpec-GAN, the total loss function L of the discriminator is D It can be defined by combining Sensitivity-Specificity Loss and cross entropy loss, which can be expressed as follows:
[0035] L D =α·L SSD +β·L CED
[0036] Where α and β are weight coefficients used to balance the contribution of Sensitivity-Specificity Loss and cross entropy loss in the total loss of the discriminator. SSD and L CED Represent the specific values of Sensitivity-Specificity Loss and cross entropy loss respectively.
[0037] SensSpec-GAN uses the generator for the final data enhancement, but in the discriminator D, SensSpec-GAN introduces the Sensitivity-Specificity loss function, which is combined with the cross-entropy loss function to generate samples, so that the discriminator can achieve a balance between sensitivity and specificity. In actual use, the model's emphasis on sensitivity and specificity can be adjusted through parameters.
[0038] The performance indicators of different algorithms under 30% data enhancement ratio are as follows: Figure 2 shown.
[0039] Step 3: Perform data dimensionality reduction. Use a sparse autoencoder based on the vulture optimization algorithm to reconstruct the input data. The vulture optimization algorithm can automatically find the hyperparameters of the sparse autoencoder and improve the performance of the sparse autoencoder.
[0040] Using a sparse autoencoder to compress the CSI data before it is input into the neural network can not only effectively improve the efficiency of the physical layer authentication system, but also significantly reduce the delay in the authentication process. In addition, the sparse autoencoder can learn the inherent laws of the data, thereby efficiently extracting the key features of the CSI data, allowing the model to adapt to different wireless propagation environments and devices. However, the sparse autoencoder has a series of hyperparameters such as the sparsity penalty term and the L2 regularization term. Improper hyperparameter settings may lead to poor model performance, requiring a large number of experiments to optimize and adjust the hyperparameters, which limits its wide application in practice.
[0041] The present invention proposes a sparse autoencoder SABES (Sparse Autoencoder Based on the Bald Eagle Search Algorithm) based on a vulture optimization algorithm. The vulture optimization algorithm is an efficient global search algorithm, which imitates the search strategy of vultures when hunting and obtains the global optimal solution of model hyperparameters through iterative search. Before training the sparse autoencoder, the method optimizes the performance of the sparse autoencoder by automatically adjusting the sparsity penalty factor and other key hyperparameters. The vulture optimization algorithm is applied to the hyperparameter optimization of the sparse autoencoder, so that it can overcome the difficulties in hyperparameter optimization, thereby achieving more efficient and accurate physical layer authentication in a resource-constrained environment.
[0042] The specific steps for searching model hyperparameters in SABES are as follows:
[0043] (1) Initialization
[0044] Generate an initial population of vultures, each vulture representing a specific set of decision variable values (l i ,s i ), where l i represents the L2 regularization parameter, s i represents the sparse regularization parameter. For each parameter P i ∈(l i ,s i ), the values of these variables can be randomly initialized as follows to form the initial population:
[0045] Pi =P min +rand()*(P max -P min )
[0046] (2) Selection stage
[0047] In the selection phase, the vulture will fly near the current best individual, and then randomly select the search area and choose the area with the most prey to obtain the best location. The following formula is used to select the search area:
[0048]
[0049] In the formula, α is a constant in the range of [1.5, 2] to control the range of position change, and r is a random number in the range of [0, 1]. is the average position of the group.
[0050] (3) Search phase
[0051] During the search phase, the vulture will update its position around the current search space using an Archimedean spiral and speed up the search by moving in different directions. The specific implementation is as follows:
[0052] P i,new =P i +y(i)×(P i -P i+1 )+x(i)×(P i -P mean )
[0053]
[0054] xr(i)=r(i)×sin(θ(i)), yr(i)=r(i)×cos(θ(i))
[0055] θ(i)=α′×π×S, r(i)=θ(i)+R×S
[0056] Where α' and R are constant parameters that control the flight trajectory, where α'∈[5,10], R∈[0.5,2]. x(i) and y(i) represent the current position of the vulture in the coordinates, S is a random number in the range [0,1], and θ(i) and r(i) represent the polar angle and polar diameter of the spiral equation.
[0057] (4) Subduction stage
[0058] After the vulture locks on the prey in the search space during the search phase, the vulture will fly towards the prey in a spiral from its current position and update it using the following:
[0059] Pi,new =S×P best +x1(i)×(P i -c1×P mean )+y1(i)×(P i -c2×P best )
[0060]
[0061] xr(i)=r(i)×sinh(θ(i)), yr(i)=r(i)×cosh(θ(i))
[0062] θ(i)=α×π×S, r(i)=θ(i)
[0063] Where x(i) and y(i) represent the current position of the vulture in the coordinate system, c1 represents the flying speed to the optimal position, and its range is c1∈(-1,1), and c2 represents the flying speed to the average position, and its range is c1∈(1,2). These two parameters can increase the randomness of the vulture's movement intensity.
[0064] The impact of different dimensionality reduction techniques on performance indicators is as follows: Figure 3 shown.
[0065] Step 4: Use knowledge distillation strategy to train the model. First, use a large amount of labeled data to train multiple teacher models and perform complex security authentication task training. Then, in order to verify the effectiveness of the teacher model in practical applications, the performance of each teacher model is evaluated on an independent test set and the model is screened to ensure that it can provide high-accuracy security authentication in a variety of scenarios. Finally, the adaptive weighted multi-teacher knowledge distillation model AMTDF is used to distill the large teacher model into a lightweight student model, so that the student model can better simulate the output distribution of the teacher model and promote the effective transfer of complex knowledge.
[0066] This scheme is divided into two parts: teacher model screening and weight adaptive allocation. In the model screening part, we first calculate the F of the teacher model on the prediction set. 0.5 Score is used as an evaluation indicator of the degree to which the teacher model can pass on positive knowledge to students. Then, the F 0.5 Score is processed to obtain the model screening threshold. The screening process is as follows:
[0067] First, all teacher models F are calculated by formula 0.5 The mean and standard deviation of Score, where N represents the total number of teacher models, F 0.5,i represents the F of the i-th teacher model 0.5 Score:
[0068]
[0069]
[0070] In order to improve the generalization ability of the student model, this scheme uses the mean value minus 0.5 times the standard deviation as the teacher model screening threshold to ensure that while retaining as many teacher models as possible, F 0.5 The teacher model with a score that is too low. The threshold formula is as follows:
[0071]
[0072] This solution sets the number of initial teacher models to 3. After all teacher models are trained, the F 0.5 Score indicators, and remove models with indicators below the threshold. Other models participate in subsequent training and calculation.
[0073] In order to further enhance the performance of the student model, the present invention uses the exponential weighted allocation method to calculate the weight of the teacher model, so that the student model can learn richer and more comprehensive information from multiple teacher models. The weight of each teacher model w i The calculation formula is as follows:
[0074]
[0075] Considering the characteristic complexity of CSI data and the resource limitations and efficiency requirements in practical applications, the present invention uses three deep neural networks with residual connections as teacher models, and uses the algorithm of the present invention to adaptively adjust the weights. A shallow fully connected network with a simple structure is used as a student model to ensure efficient physical layer security authentication in a resource-limited environment.
[0076] The overall architecture of the solution is shown in the figure below. Figure 4 The relationship between the number of teacher models and the accuracy of student models is shown in Figure 5 shown.
Claims
1. A physical layer authentication scheme based on sparse autoencoder and knowledge distillation, characterized in that: The method comprises the following steps: Step 1: Data collection and preprocessing: Use Linux 802.11n CSITool to collect CSI data in real communication scenarios, then calculate the amplitude and phase of each subcarrier based on the subcarrier complex matrix, then use the Savitzky-Golay filter to denoise the CSI data, and finally normalize the data for subsequent analysis and processing; Step 2: Perform data enhancement and use a generative adversarial network based on Sensitivity-Specificity Loss to generate high-quality experimental data to improve the model's ability to recognize negative samples. Step 3: Perform data dimensionality reduction and reconstruct the input data using a sparse autoencoder based on the vulture optimization algorithm. The vulture optimization algorithm can automatically find the hyperparameters of the sparse autoencoder and improve the performance of the sparse autoencoder. Step 4: Use the knowledge distillation strategy to train the model. First, use a large amount of labeled data to train multiple teacher models and perform complex security authentication task training. Then, in order to verify the effectiveness of the teacher model in practical applications, the performance of each teacher model is evaluated on an independent test set and the model is screened to ensure that it can provide high-accuracy security authentication in a variety of scenarios. Finally, use the adaptive weighted multi-teacher knowledge distillation model AMTDF to distill the large teacher model into a lightweight student model, so that the student model can better simulate the output distribution of the teacher model and promote the effective transfer of complex knowledge.
2. According to the physical layer authentication scheme based on sparse autoencoder and knowledge distillation as described in claim 1, it is characterized in that: The process of generating high-quality experimental data by using a generative adversarial network based on Sensitivity-Specificity Loss in step 2 is characterized in that a fully connected network including two hidden layers is used as a generator, and a deep neural network based on a combination of Sensitivity-Specificity Loss and cross entropy loss is used as a discriminator in order to make full use of the sequence characteristics in the CSI data; Sensitivity-Specificity Loss is defined as follows: L ss =w·sensitivity+(1-w)·specificity In the formula In SensSpec-GAN, the total loss function L of the discriminator is D It can be defined by combining Sensitivity-SpecificityLoss and cross entropy loss, which can be expressed as follows: L D =α·L SSD +β·L CED Where α and β are weight coefficients used to balance the contribution of Sensitivity-Specificity Loss and cross entropy loss in the total loss of the discriminator. SSD and L CED Represent the specific values of Sensitivity-Specificity Loss and cross entropy loss respectively.
3. According to the physical layer authentication scheme based on sparse autoencoder and knowledge distillation as described in claim 1, it is characterized in that: The process of performing data dimensionality reduction using a sparse autoencoder based on a vulture optimization algorithm described in step 3 is characterized in that before training the sparse autoencoder, the vulture optimization algorithm is used to automatically adjust the sparsity penalty factor and other key hyperparameters to optimize the performance of the sparse autoencoder.
4. According to the physical layer authentication scheme based on sparse autoencoder and knowledge distillation as described in claim 1, it is characterized in that: The process of training the model using the knowledge distillation strategy described in step 4 is characterized by calculating the F of the teacher model on the prediction set 0.5 Score is used as an evaluation indicator of the degree to which the teacher model can transfer positive knowledge to students. Then, the F 0.5 Score is processed to obtain the model screening threshold. The screening process is as follows: First, all teacher models F are calculated by formula 0.5 The mean and standard deviation of Score, where N represents the total number of teacher models, F 0.5,i represents the F of the i-th teacher model 0.5 Score: Calculate the F corresponding to each model in the test set 0.5 Score indicator, and eliminate models with indicators below the threshold. Other models participate in subsequent training and calculation. At the same time, the exponential weighted allocation method is used to calculate the weight of the teacher model.
Citation Information
Patent Citations
Image recognition model compression method based on adversarial distillation technology
CN114170332A
Internet of Things indoor positioning method based on federated distillation
CN115358419A
Lightweight Internet of Things malicious traffic identification method based on knowledge distillation space-time neural network
CN116260642A
Intrusion detection method under group learning architecture fused with knowledge distillation
CN117812593A
Distributed Internet of Things equipment identification method and system based on federated learning
CN118301092A
Cited By
Neural network model construction method and device for gas turbine performance prediction
CN120930695A