Adaptive privacy protection knowledge distillation method based on multi-layer feature extraction
Patent Information
- Application Number
- CN202311548611.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-11-21
AI Technical Summary
然而,过多的噪声会影响学生模型的性能
[0014]In the distillation learning phase, the student model uses the soft labels output by the teacher model as the learning target to solve for the distillation loss. Then, it updates its own parameters through backpropagation based on the distillation loss. This training process forces the soft prediction vectors output by the student model to get closer and closer to the soft labels output by the teacher model, thus enabling the student model to learn more inter-class relationships. Since this phase uses knowledge from the teacher model, it may lead to sensitive data leakage issues related to the teacher's training set. Therefore, we use Dynamic Distillation Temperature (DT) in this phase. Dynamic Distillation Temperature (DT) achieves adaptive privacy protection by adjusting the distillation temperature of the teacher model. Here, the logistic unit is the layer preceding the classification function, and the soft label is derived from the logistic unit values through a classification function with a temperature coefficient. When the privacy requirements of the teacher model's training set are low, we can choose a lower distillation temperature. In this case, the distribution of soft labels will be sharper, meaning some classes have higher probability values. By using a lower distillation temperature, the student model can more easily focus on those classes with higher probability values, thereby improving performance. Conversely, when the privacy requirements of the teacher model's training set are high, we can use a higher distillation temperature. In this case, the distribution of soft labels will be flatter, meaning the probability values of each class are relatively balanced. By using a higher distillation temperature, we increase the difficulty of the student model's learning, as it requires more careful differentiation between different categories, thereby reducing the risk of privacy leakage of the teacher's training set. Adaptive privacy protection of the teacher's training set can be achieved during the distillation learning phase by flexibly adjusting the distillation temperature of the teacher model.
Smart Images

Figure CN117436128B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of privacy protection in knowledge distillation, specifically to an adaptive privacy-preserving knowledge distillation method based on multi-layer feature extraction. Background Technology
[0002] Knowledge distillation is a model compression technique that aims to transfer the "hidden knowledge" from a pre-trained teacher model to a student model, enabling the student model to better understand the relationships between categories and improve prediction accuracy. In traditional knowledge distillation, the output of the teacher model is treated as a soft label, which contains richer information about inter-class relationships compared to the true label. When training the student model, training typically involves minimizing the distillation loss with the soft label and the self-loss with the true label. However, traditional knowledge distillation does not utilize the feature knowledge extracted by the teacher model, which limits the student model's ability to understand and process complex data. Therefore, it is necessary to leverage the multi-layered feature knowledge extracted by the teacher model to further improve the performance of the student model.
[0003] Knowledge distillation can utilize differential privacy techniques to protect the privacy of the teacher model's training set. However, differential privacy achieves privacy by adding noise to the data; the more noise added, the higher the level of privacy protection. When the teacher model's training dataset contains highly sensitive information, more noise needs to be injected into the student model to prevent privacy leaks. However, excessive noise can negatively impact the student model's performance. Therefore, to improve the overall performance of the student model, adaptive privacy protection needs to be implemented for data of varying sensitivities within the teacher model's training set. Summary of the Invention
[0004] To learn multi-layer feature knowledge from the teacher model and to adaptively protect the privacy of the teacher training set, this invention provides an adaptive privacy-preserving knowledge distillation method based on multi-layer feature extraction.
[0005] The technical method adopted in this invention is: a multi-layer feature knowledge learning method, which enables the student model to further improve its performance based on knowledge distillation; secondly, this invention uses adaptive differential privacy gradient descent (ADP-SGD) and dynamic distillation temperature (DT) in multi-layer feature learning and distillation learning, respectively, to ensure the security of the teacher model training set during the knowledge distillation process.
[0006] The adaptive privacy-preserving knowledge distillation method based on multi-layer feature extraction has three learning stages. The first stage is feature learning, where the student model learns multi-layer feature knowledge from the teacher model to better extract data features from samples. During this learning process, adaptive differential privacy gradient descent (ADP-SGD) is used to adaptively preserve the privacy of the teacher training set. The second stage is self-learning, where the student model learns the label knowledge of samples to correct its erroneous predictions. The third stage is distillation learning, where the student model learns the soft label knowledge from the teacher model to better understand the relationships between different categories. During this learning process, dynamic distillation temperature (DT) is used to adaptively preserve the privacy of the teacher training set. A detailed description follows:
[0007] (1) The teacher model extracts feature knowledge of each class of samples in the training set.
[0008] The teacher model is pre-trained using public and sensitive data and does not update its parameters during knowledge distillation. In deep learning tasks, the features extracted by the teacher model's feature extractor are redundant. When these features are used as labels for the student model's feature learning, this redundancy increases the student model's training time. To address this issue, this paper averages the data features of each class of samples to obtain a generalized vector representing the general characteristics of that class. Using these generalized vectors as feature knowledge effectively improves the training speed of the student model's feature learning.
[0009] (2) In the feature learning stage, the student model learns the multi-layered feature knowledge of the teacher model.
[0010] In the feature learning phase, the student model uses the multi-layer feature knowledge extracted by the teacher model as the learning objective to solve for the multi-layer feature loss. Then, it updates its own parameters based on the multi-layer feature loss through backpropagation. This training process forces the multi-layer features extracted by the student model to become increasingly closer to those extracted by the teacher model, thereby improving the student model's performance in extracting features from sample data. However, the feature learning phase uses knowledge from the teacher model, which may lead to sensitive data leakage from the teacher's training set. Therefore, we use adaptive differential privacy protection in this phase and apply this technique to the gradient of the student model's backpropagation. Then, we use Adaptive Differential Privacy Gradient Descent (ADP-SGD) to update the student model's parameters. The feature learning phase employs a classification training approach. Assuming σ represents the standard deviation of a Gaussian distribution, adaptive differential privacy protection achieves adaptive privacy protection by adjusting the standard deviation σ of the Gaussian distribution. When a certain type of data requires stronger privacy protection, a larger standard deviation σ can be chosen to make the generated Gaussian noise more dispersed, thereby increasing the perturbation of the gradient for that type of data and reducing the risk of privacy leakage for that type of data. Conversely, if a certain type of data does not require strong privacy protection, a smaller standard deviation σ can be chosen to concentrate the generated Gaussian noise, thereby reducing the perturbation of the gradient and improving the overall performance of the student model. By flexibly adjusting the standard deviation σ of the Gaussian distribution for differential privacy, different Gaussian noises can be generated. Applying these Gaussian noises to the gradients during the backpropagation of the student model can achieve adaptive privacy protection of the teacher training set during the feature learning stage.
[0011] (3) During the self-learning phase, the student model learns the label knowledge of the samples.
[0012] During the self-learning phase, the student model uses the true label of the sample as the learning target to solve for the self-learning loss, and then updates its own parameters based on the backpropagation of the self-learning loss. This training process can force the student model's prediction vector to get closer and closer to the true label of the sample, thereby correcting the student model's wrong prediction of the sample.
[0013] (4) During the distillation learning phase, the student model learns the soft-label knowledge of the teacher model using dynamic distillation temperature.
[0014] In the distillation learning phase, the student model uses the soft labels output by the teacher model as the learning target to solve for the distillation loss. Then, it updates its own parameters through backpropagation based on the distillation loss. This training process forces the soft prediction vectors output by the student model to get closer and closer to the soft labels output by the teacher model, thus enabling the student model to learn more inter-class relationships. Since this phase uses knowledge from the teacher model, it may lead to sensitive data leakage issues related to the teacher's training set. Therefore, we use Dynamic Distillation Temperature (DT) in this phase. Dynamic Distillation Temperature (DT) achieves adaptive privacy protection by adjusting the distillation temperature of the teacher model. Here, the logistic unit is the layer preceding the classification function, and the soft label is derived from the logistic unit values through a classification function with a temperature coefficient. When the privacy requirements of the teacher model's training set are low, we can choose a lower distillation temperature. In this case, the distribution of soft labels will be sharper, meaning some classes have higher probability values. By using a lower distillation temperature, the student model can more easily focus on those classes with higher probability values, thereby improving performance. Conversely, when the privacy requirements of the teacher model's training set are high, we can use a higher distillation temperature. In this case, the distribution of soft labels will be flatter, meaning the probability values of each class are relatively balanced. By using a higher distillation temperature, we increase the difficulty of the student model's learning, as it requires more careful differentiation between different categories, thereby reducing the risk of privacy leakage of the teacher's training set. Adaptive privacy protection of the teacher's training set can be achieved during the distillation learning phase by flexibly adjusting the distillation temperature of the teacher model. Attached Figure Description
[0015] Figure 1 Gaussian distributions at different standard deviations;
[0016] Figure 2 Distribution of soft labels at different distillation temperatures;
[0017] Figure 3 Adaptive privacy-preserving knowledge distillation based on multi-layer feature extraction;
[0018] Figure 4 Accuracy of student models tested on the MNIST dataset; Detailed Implementation
[0019] (1) The teacher model extracts feature knowledge of each class of samples in the training set.
[0020] (1.1) The teacher model selects feature extractors at different levels. The neural network used as the teacher model is divided into three parts: a shallow network, an intermediate network, and a deep network. The teacher model selects a convolutional layer located in the middle of the shallow network as the shallow feature extractor F, a convolutional layer located in the middle of the intermediate network as the intermediate feature extractor M, and a convolutional layer located in the middle of the deep network as the deep feature extractor D.
[0021] (1.2) The teacher model uses feature extractors at different levels to extract data features. (1.2.1) Assume x p This represents the training set, which contains K classes and N. j This represents the number of samples in class j (1≤j≤K). This represents the i-th data point of the j-th sample. and Let represent the shallow feature vector, intermediate feature vector, and deep feature vector of the i-th data of the j-th sample extracted by the teacher model, respectively. (1.2.2) will The input is fed into the teacher model, which uses a shallow feature extractor F to extract the features. Shallow feature vectors The intermediate layer feature extractor M can be used to extract... intermediate layer feature vectors Using the deep feature extractor D, we can extract... deep feature vectors
[0022] (1.3) The teacher model obtains feature knowledge at different levels for each type of sample. (1.3.1) Assumption and These represent the shallow feature knowledge, intermediate feature knowledge, and deep feature knowledge of the j-th class of samples extracted by the teacher model, respectively. (1.3.2) The teacher model aggregates the shallow, intermediate, and deep feature vectors of the extracted j-th class samples according to their respective layers and calculates the mean to obtain the shallow feature knowledge of the j-th class samples. Intermediate layer feature knowledge and deep feature knowledge
[0023] (2) The student model learns multi-layered feature knowledge from the teacher model during the feature learning stage.
[0024] (2.1) Selecting feature extractors at different levels for the student model The neural network used as the student model is divided into three parts: shallow network, intermediate network, and deep network. The student model selects a convolutional layer located in the middle of the shallow network as the shallow feature extractor f, a convolutional layer located in the middle of the intermediate network as the intermediate feature extractor m, and a convolutional layer located in the middle of the deep network as the deep feature extractor d.
[0025] (2.2) The student model uses feature extractors at different levels to extract data features. (2.2.1) Assumption and These represent the shallow feature vector, intermediate feature vector, and deep feature vector of the i-th data of the j-th sample extracted by the student model, respectively. (2.2.2) will The input is fed into the student model, which uses a shallow feature extractor f to extract the features. The shallow feature vector is Using the intermediate layer feature extractor m, the following can be extracted: The intermediate layer feature vector is Using the deep feature extractor d, we can extract... The deep feature vector is
[0026] (2.3) The student model calculates the multi-level feature loss for each class of samples. The feature learning phase employs classification training, with the student model using its own extracted features. Feature knowledge extracted from teacher models The multi-level feature loss of the j-th class sample during the multi-level feature learning stage can be calculated.
[0027] (2.4) The student model calculates the gradient of backpropagation for each class of samples based on the multi-level feature loss of each class. feature Solve for the gradient g of the j-th class sample. j ;
[0028] (2.5) Adaptive differential privacy perturbs the gradient of each class of samples with noise. (2.5.1) Assume σ j This represents the standard deviation of the Gaussian distribution used when performing adaptive differential privacy protection on samples of class j. This indicates that the mean is 0 and the standard deviation is σ. j A Gaussian distribution, where noisej represents the distribution with a mean of 0 and a standard deviation of σ. j Gaussian noise generated by a Gaussian distribution. This indicates that the gradient g jThe new gradient generated after noise perturbation; (2.5.2) Gaussian noise Injected into gradient g j In this process, a noisy version of the gradient can be obtained.
[0029] (2.6) The student model uses the gradient after noise perturbation for backpropagation. Assume S θ Let θ represent the student model, where θ is the parameter of the student model, which utilizes a noisy version of the gradient. Perform backpropagation to update the parameters θ of the student model;
[0030] (2.7) Repeat steps (2.2)-(2.6) until feature learning has been performed for all K categories in the dataset;
[0031] (3) The student model learns the label knowledge of the samples during the self-learning phase.
[0032] (3.1) Predicted vector of student model output sample Assuming y represents the prediction vector of the student model, let x p Input into student model S θ In the middle, the student model will output x p The prediction vector y = S θ (x p );
[0033] (3.2) Calculate the self-learning loss of the student model The self-learning phase does not employ classification training. Let N represent the total number of samples in the dataset, and y... r (1≤r≤N) represents the predicted vector of the r-th sample output by the student model. The student model uses the predicted vector y to represent the true label of the r-th sample. r and real labels The self-learning loss during the self-learning phase can be calculated.
[0034] (3.3) The student model calculates the gradient based on the self-learning loss and performs backpropagation. The student model is based on the self-learning loss L self Solve for the gradient g and perform backpropagation to update the parameters θ of the student model;
[0035] (4) During the distillation learning phase, the student model learns the soft-label knowledge of the teacher model using dynamic distillation temperature.
[0036] (4.1) The teacher model and student model output the distilled sample prediction vectors. (4.1.1) In the distillation learning phase, classification training is not used. Assume T represents the distillation temperature of the student model, λ represents the hyperparameters, n represents the total number of samples in each batch, α represents the number of samples in the high-sensitivity category in each batch, and t represents the dynamic distillation temperature of the teacher model. z r w represents the logical unit value of the r-th sample output by the teacher model. r This represents the logical unit value of the r-th sample output by the student model. This represents the soft label of the r-th sample output by the teacher model. This represents the soft prediction vector of the r-th sample output by the student model; (4.1.2) Input the r-th sample into the teacher model, and the teacher model outputs the soft label of the r-th sample. (4.1.3) Input the r-th sample into the student model, and the student model outputs the soft prediction vector of the r-th sample.
[0037] (4.2) Student model calculates distillation loss Student models use soft labels and soft prediction vector The distillation loss during the distillation learning phase can be calculated.
[0038] (4.3) The student model calculates the gradient based on the distillation loss and performs backpropagation. The student model uses backpropagation to update its parameters g by solving for the gradient g based on the distillation loss Ldistill;
[0039] (5) Repeat steps (2)-(4) until L. self L feature L distill Both have stabilized, achieving adaptive privacy-preserving knowledge distillation based on multi-layer feature extraction.
Claims
1. An adaptive privacy-preserving knowledge distillation method based on multi-layer feature extraction, characterized in that: (1) The teacher model extracts feature knowledge of each type of sample in the training set. (1.1) The teacher model selects feature extractors at different levels. The neural network used as the teacher model is divided into three parts: a shallow network, an intermediate network, and a deep network; the teacher model selects a convolutional layer located in the middle of the shallow network as a shallow feature extractor. In the intermediate layer network, a convolutional layer located in the middle position is selected as the intermediate layer feature extractor. In the deep network, a convolutional layer located in the middle is selected as the deep feature extractor. ; (1.2) The teacher model uses feature extractors at different levels to extract data features. (1.2.1) Assumption This represents the training set, which contains... Categories Indicates the first (1 The number of samples in class ) (1 ) indicates the first Class of samples One data point, , and These represent the first and second parts extracted by the teacher model, respectively. Class Sample No. The data consists of shallow feature vectors, intermediate feature vectors, and deep feature vectors. (1.2.2) will The input is fed into the teacher model, which uses a shallow feature extractor. It can be extracted Shallow feature vectors Using intermediate layer feature extractors It can be extracted intermediate layer feature vectors Using a deep feature extractor It can be extracted deep feature vectors ; (1.3) The teacher model obtains feature knowledge at different levels for each type of sample. (1.3.1) Assumption , and These represent the first and second parts extracted by the teacher model, respectively. Shallow feature knowledge, intermediate feature knowledge, and deep feature knowledge of class samples; (1.3.2) The teacher model will extract the first... The shallow, intermediate, and deep feature vectors of a class of samples are aggregated and averaged according to their respective layers to obtain the first... Shallow feature knowledge of class samples = Intermediate layer feature knowledge = and deep feature knowledge = ; (2) In the feature learning stage, the student model learns the multi-layer feature knowledge of the teacher model. (2.1) Selecting feature extractors at different levels for the student model The neural network used as the student model is divided into three parts: a shallow network, an intermediate network, and a deep network; the student model selects a convolutional layer located in the middle of the shallow network as a shallow feature extractor. In the intermediate layer network, a convolutional layer located in the middle position is selected as the intermediate layer feature extractor. In the deep network, a convolutional layer located in the middle is selected as the deep feature extractor. ; (2.2) The student model uses feature extractors at different levels to extract data features. (2.2.1) Assumption , and These represent the first and second parts extracted by the student model, respectively. Class Sample No. The shallow feature vector, intermediate feature vector, and deep feature vector of each data point; (2.2.2) will The input is fed into the student model, which uses a shallow feature extractor. It can be extracted The shallow feature vector is Using intermediate layer feature extractors It can be extracted The intermediate layer feature vector is Using a deep feature extractor It can be extracted The deep feature vector is ; (2.3) The student model calculates the multi-level feature loss for each class of samples. The feature learning phase employs classification training, with the student model using its own extracted features. , , Feature knowledge extracted from teacher models , , The first one can be calculated. Multi-level feature loss of class samples in the multi-level feature learning stage ; (2.4) The student model calculates the gradient of backpropagation for each class of samples based on the multi-layer feature loss of each class of samples. The student model uses multi-layer feature loss. Solve for the first Gradient of class samples ; (2.5) Adaptive differential privacy perturbs the gradient of each class of samples with noise. (2.5.1) Assumption Indicates the first The standard deviation of the Gaussian distribution used when performing adaptive differential privacy protection on class samples. The mean is Standard deviation is Gaussian distribution, This indicates that the mean is used. Standard deviation is Gaussian noise generated by a Gaussian distribution. Indicates gradient The new gradient generated after noise perturbation; (2.5.2) Gaussian noise ( Injected into gradient In this process, a noisy version of the gradient can be obtained. ; (2.6) The student model uses the gradient after noise perturbation for backpropagation. Assumption Representing the student model, These are the parameters of the student model, which utilizes a noisy version of the gradient. Perform backpropagation to update the parameters of the student model. ; (2.7) Repeat steps (2.2)-(2.6) until the dataset contains... Feature learning was performed for each category; (3) During the self-learning phase, the student model learns the label knowledge of the samples. (3.1) Predicted vector of student model output sample Assumption This represents the prediction vector of the student model. Input into student model In the middle, the student model will output Prediction vector ; (3.2) Calculate the self-learning loss of the student model The self-learning phase does not employ classification training; it is assumed that... This represents the total number of samples in the dataset. ( ) represents the first output of the student model. The predicted vector for each sample. Indicates the first The true label of each sample is used by the student model to predict vectors. and real labels The self-learning loss during the self-learning phase can be calculated. ; (3.3) The student model calculates the gradient based on the self-learning loss and performs backpropagation. The student model is based on self-learning loss Solving the gradient Perform backpropagation to update the parameters of the student model. ; (4) During the distillation learning phase, the student model learns the soft-label knowledge of the teacher model using dynamic distillation temperature. (4.1) The teacher model and student model output the distilled sample prediction vectors. (4.1.1) During the distillation learning phase, classification training is not used. It is assumed that... This indicates the distillation temperature of the student model. Indicates hyperparameters, This indicates the total number of samples in each batch. The number of samples of the high-sensitivity category included in each batch. The dynamic distillation temperature of the teacher model is represented by, where , The first output of the teacher model The logical unit value of each sample, The first output of the student model represents the... The logical unit value of each sample, The output of the teacher model represents the first... Soft labels for each sample, The first output of the student model represents the... The soft prediction vector for each sample; (4.1.2) The first The nth sample is input into the teacher model, and the teacher model outputs the nth sample. Soft labels for each sample ; (4.1.3) The first The nth sample is input into the student model, and the student model outputs the nth sample. Soft prediction vector for each sample ; (4.2) Student model calculates distillation loss Student models use soft labels and soft prediction vector The distillation loss during the distillation learning phase can be calculated. ; (4.3) The student model calculates the gradient based on the distillation loss and performs backpropagation. Student model based on distillation loss Solving the gradient Perform backpropagation to update the parameters of the student model. ; (5) Repeat (2)-(4) until , , Both have stabilized, achieving adaptive privacy-preserving knowledge distillation based on multi-layer feature extraction.