Differential privacy stochastic gradient descent method based on VGP and SU

By introducing vertical gradient perturbation and selective update strategies in the differential privacy stochastic gradient descent method, the problems of privacy and utility balance and low computing efficiency are solved, higher model generalization ability and training efficiency are achieved, and high classification accuracy is maintained under strict privacy requirements.

CN119940461APending Publication Date: 2025-05-06QIQIHAR UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510014363.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing differential privacy stochastic gradient descent (DP-SGD) method is difficult to balance between privacy and utility, and has low computational efficiency, resulting in reduced model accuracy loss and convergence speed.

Method used

The differential privacy stochastic gradient descent method based on VGP and SU is used to decompose the gradient into parallel and vertical components through the vertical gradient perturbation (VGP) mechanism, and different degrees of perturbation are applied to it. Combined with the selective update strategy (SU) to monitor the loss changes, the model update frequency is dynamically adjusted.

Benefits of technology

While ensuring privacy, it improves the generalization ability and training efficiency of the model, reduces noise injection, improves information gain, and maintains high classification accuracy under strict privacy requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940461A_ABST
    Figure CN119940461A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of deep learning, and discloses a VGP and SU-based differential privacy stochastic gradient descent method, which is an improved scheme based on a DP-SGD algorithm, and combines a vertical gradient perturbation mechanism VGP and a selective update mechanism SU for use. A privacy budget is allocated to incremental gradient information using a validation set filter model update, thereby reducing noise injection and improving information gain. Meanwhile, in order to further optimize privacy budget distribution, a staged application strategy is adopted, VGP and SU are applied at the same time in the initial training stage, and only SU is used in the subsequent stage to cope with continuously changing gradient correlation and improve privacy efficiency; according to the method, the problem of balance between privacy and utility and the problem of low calculation efficiency of a differential privacy stochastic gradient (DP-SGD) method in the prior art are solved, and the method is suitable for a deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to a differential privacy stochastic gradient descent method based on VGP and SU. Background Art

[0002] With the advent of the digital age, deep learning technology has developed rapidly with its outstanding learning and feature extraction capabilities, and has achieved remarkable results in many fields such as image recognition, natural language processing, and health care. However, when the training data contains personal sensitive information, this sensitive information may be "remembered" by the deep learning model, causing the training data to be threatened by data reconstruction and member reasoning attacks. Therefore, the training process of deep learning models faces severe data privacy challenges. Privacy protection technologies for deep learning involve data encryption (including homomorphic encryption, symmetric encryption, asymmetric encryption, etc.), data anonymization, federated learning, secure multi-party computing, etc. Among them, differential privacy technology can protect personal sensitive information while ensuring the stability of model prediction results, and occupies an important position in privacy protection technologies for deep learning.

[0003] Traditional differential privacy deep learning protection technologies, such as the differential privacy stochastic gradient descent (DP-SGD) method, can protect sensitive information in training data to a certain extent, but they have obvious limitations, which are mainly reflected in the following two aspects: First, the balance between privacy and utility. The DP-SGD method implements differential privacy protection by adding random noise in the gradient calculation process, which will lead to a large loss of model accuracy; second, the problem of computational efficiency. The random noise introduced by the DP-SGD method in the gradient calculation process will interfere with the gradient direction, causing the model update to deviate from the optimal trajectory, thereby reducing the convergence speed. Moreover, the noise variance is inversely proportional to the privacy budget ε, which means that stronger privacy protection (smaller ε) will further aggravate the negative impact on the convergence speed. These limitations have seriously restricted the widespread use of the DP-SGD method in practical applications. Summary of the invention

[0004] The present invention intends to provide a differential privacy stochastic gradient descent method based on VGP and SU to solve the problems of balancing privacy and utility and computational efficiency in the prior art differential privacy stochastic gradient (DP-SGD) method.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] A differential privacy stochastic gradient descent method based on VGP perturbation and SU update includes the following steps:

[0007] S1, initialize the model parameter w0 and perform iterative optimization;

[0008] S2. In each iteration t, if t is less than the threshold s, the VGP process is called, otherwise:

[0009] S21. Extract a batch of samples B from the data set with a specified probability t ;

[0010] S22, for each sample x i , calculate and normalize the gradient g t (x i );

[0011] S23, calculating the mean of the normalized gradient of the sample and adding Gaussian noise;

[0012] S3, using the calculated perturbation gradient g t To update the model parameters w new ;

[0013] S4, call SU process;

[0014] S5. Return the final model parameter w after the iteration. t .

[0015] Furthermore, the VGP process in step S2 includes the following steps:

[0016] A1. Randomly extract a batch of samples B from the data set t , and calculate each sample xi relative to the current model parameter w t The gradient g t (x i );

[0017] A2, decompose the gradient into two components, an orthogonal component perpendicular to the baseline gradient and a parallel component parallel to the baseline gradient;

[0018] A3. Normalize these two components to ensure the consistency of scale;

[0019] A4, adding Gaussian noise with a specific distribution to the orthogonal components;

[0020] A5. Combine the normalized parallel components and the perturbed orthogonal components to obtain the final noise gradient for updating the model.

[0021] Furthermore, the calling SU process in step S4 includes the following steps:

[0022] B1. Extract a batch of samples from the data set with a specified probability:

[0023] B2. Calculate the current model parameters w new and the last iteration parameter w t-1The associated loss value J(w new ) and J(w t-1 );

[0024] B3. Calculate the loss change ΔE, trim it and add Gaussian noise to increase randomness;

[0025] B4. If the adjusted loss change ΔE is lower than a given threshold, the new model parameter w is accepted new And update the model, otherwise, the current parameter w t-1 Remain unchanged.

[0026] B5. Use the current gradient g t Update the reference vector and return the updated model parameters w t .

[0027] Principle and beneficial effects of the technical solution: The present invention is an improved solution DP-VGPSU based on the differential privacy stochastic gradient descent (DP-SGD) method. The present method first decomposes the gradient into parallel and vertical components through the vertical gradient perturbation (VGP) mechanism, and applies different degrees of perturbation to it, thereby reducing the risk of information leakage and increasing the generalization ability of the model. Secondly, the loss change is monitored by the selective update strategy (SU), and the model update frequency is dynamically adjusted, which can not only fully learn in the early stage of training, but also reduce unnecessary calculations in the stable stage. The present invention combines the vertical gradient perturbation mechanism (VGP) and the selective update mechanism (SU), uses the validation set to filter the model update, and allocates the privacy budget to the incremental gradient information, thereby reducing noise injection and improving information gain. At the same time, in order to further optimize the privacy budget allocation, a phased application strategy is adopted, and VGP and SU are applied simultaneously in the initial training stage, while only SU is used in the subsequent stages to cope with the changing gradient correlation and improve privacy efficiency. Under the condition of using the same data set and privacy budget, the classification accuracy of the method of the present invention is always better than other existing methods, and it is more obvious within a smaller privacy budget range, highlighting the ability to maintain accuracy under strict privacy requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is the working flow chart of the method DP-VGPSU of the present invention;

[0029] Figure 2 It is the VGP algorithm diagram;

[0030] Figure 3 It is a schematic diagram of incremental information;

[0031] Figure 4 It is the SU algorithm diagram;

[0032] Figure 5 DP-VGPSU algorithm diagram of the present invention;

[0033] Figure 6 This is a comparison chart of the classification accuracy of different algorithms on the CIFAR-10 dataset;

[0034] Figure 7 This is a comparison chart of the classification accuracy of different algorithms on the FMNIST dataset;

[0035] Figure 8 Comparison chart of test loss of different algorithms on MNIST dataset; DETAILED DESCRIPTION

[0036] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments:

[0037] like Figure 1 As shown in the figure, the DP-VGPSU workflow diagram obtained by improving the differential privacy stochastic gradient descent (DP-SGD) method in the present invention. The improved DP-VGPSU combines the vertical gradient perturbation mechanism VGP and the selective update mechanism SU, uses the validation set to filter the model update, and allocates the privacy budget to the incremental gradient information, thereby reducing noise injection and improving information gain.

[0038] First, to address the balance between privacy and utility in the existing differential privacy stochastic gradient descent (DP-SGD) method, this paper adopts the vertical gradient perturbation mechanism VGP, and its specific algorithm process is as follows: Figure 2 As shown in , in the early stage of model training, the vertical gradient perturbation mechanism uses the baseline gradient obtained in the previous step of training to decompose the current gradient based on the baseline gradient to obtain the orthogonal component (perpendicular to the baseline gradient) and the parallel component (parallel to the baseline gradient). The incremental information is as follows Figure 3 As shown. Due to the nature of vector decomposition, the size of the orthogonal gradient is smaller than the original full gradient. Then Gaussian noise is added to the orthogonal component, and because of its small size, the added noise is relatively small. In order to make full use of all the information of the original sample, the vertical gradient perturbation mechanism cleverly reconstructs the new noisy gradient by iteratively reusing the noisy gradient added in the previous step. This maximizes the integrity of the data information while ensuring privacy. The specific process is as follows:

[0039] 1. Randomly extract a batch of samples Bt from the data set and calculate the gradient gt(xi) of each sample xi relative to the current model parameter wt;

[0040] 2. Decompose the gradient into two components, an orthogonal component perpendicular to the baseline gradient and a parallel component parallel to the baseline gradient;

[0041] 3. Normalize these two components to ensure the consistency of scale;

[0042] 4. Add Gaussian noise with a specific distribution to the orthogonal components;

[0043] 5. Combine the normalized parallel components and the perturbed orthogonal components to obtain the final noise gradient for updating the model.

[0044] To solve the problem of computational efficiency, the present invention adopts a selective update mechanism SU, and its specific algorithm process is as follows: Figure 4 As shown, the selective update mechanism decides whether to update the parameters by evaluating the contribution of the gradient to the model performance in each iteration. In deep learning, when some gradient changes are not conducive to model updating and may reduce model performance, updating the parameters at this time will not only introduce unnecessary noise interference but also lead to reduced model performance. Therefore, the present invention uses a selective update mechanism to significantly reduce the negative impact of invalid updates, improve the efficiency of training and the generalization ability of the model. The specific process is as follows:

[0045] 1. Extract a batch of samples from the data set with a specified probability:

[0046] 2. Calculate the current model parameters w respectively new and the last iteration parameter w t-1 The associated loss value J(w new ) and J(w t-1 );

[0047] 3. Calculate the loss change ΔE, trim it and add Gaussian noise to increase randomness;

[0048] 4. If the adjusted loss change ΔE is below a given threshold, the new model parameters w are accepted new And update the model, otherwise, the current parameter w t-1 Remain unchanged.

[0049] 5. Use the current gradient g t Update the reference vector and return the updated model parameters w t .

[0050] The improved scheme DP-VGPSU of the present invention is obtained by combining the gradient perturbation mechanism VGP and the selective update mechanism SU. The specific algorithm process is as follows: Figure 5 As shown, the technical solution of the present invention can effectively improve the efficiency of model training and enhance the privacy protection ability of the model. The specific process is as follows:

[0051] 1. Initialize the model parameter w0 and optimize it iteratively;

[0052] 2. In each iteration t, if t is less than the threshold s, the VGP process is called, otherwise:

[0053] (1) Extract a batch of samples B from the data set with a specified probability t ;

[0054] (2) For each sample x i , calculate and normalize the gradient g t (x i );

[0055] (3) Calculate the mean of the normalized gradient of the sample and add Gaussian noise;

[0056] 3. Using the calculated perturbation gradient g t To update the model parameters w new ;

[0057] 4. Call the SU process;

[0058] 5. After the iteration, the final model parameter w is returned t .

[0059] Comparative experiments with existing methods:

[0060] In order to verify the accuracy of the method of the present invention, four existing differential privacy stochastic gradient descent methods (DPIS, DPSGD-HF, DPSGD-TS, DPSD) were used to conduct comparative experiments with the method of the present invention (DP-VGPSU). The same data set and privacy budget were set to perform model training and prediction on these five methods. The data sets are MNIST, FMNIST, CIFAR-10, and IMDB. The prediction results are shown in Tables 1 to 4:

[0061] Table 1 Classification accuracy on the MNIST dataset

[0062]

[0063]

[0064] Table 2 Classification accuracy on the FMNIST dataset

[0065]

[0066] Table 3 Classification accuracy on the CIFAR-10 dataset

[0067]

[0068] Table 4 Classification accuracy on the IMDB dataset

[0069]

[0070] As shown in Tables 1 to 4, it can be seen that the classification accuracy of the DP-VGPSU method of the present invention is always better than other existing methods under the conditions of the same data set and privacy budget setting, and it is more obvious within a smaller privacy budget range, highlighting its ability to maintain accuracy under strict privacy requirements.

[0071] At the same time Figure 6 As shown in Figure 1, the classification accuracy comparison of DP-VGPSU, DPSGD-HF, DPSGD-TS and DPSGD methods on the CIFAR-10 dataset is shown. Figure 6 It can be seen that the classification accuracy of the proposed method DP-VGPSU on the CIFAR-10 dataset is higher than that of the other three methods.

[0072] like Figure 7 As shown in Figure 1, the classification accuracy comparison of DP-VGPSU, DPSGD-HF, DPSGD-TS and DPSGD methods on the FMNIST dataset is shown. Figure 7 It can be seen that the classification accuracy of the proposed method DP-VGPSU on the FMNIST data set is higher than that of the other three methods.

[0073] like Figure 8 As shown in Figure 1, the test loss comparison of DP-VGPSU, DPSGD-HF, DPSGD-TS and DPSGD methods on the MNIST dataset is shown. Figure 8 It can be seen that the test loss of the proposed method DP-VGPSU on the MNIST dataset is smaller than that of the other three methods.

[0074] The above is only an embodiment of the present invention, and the common knowledge such as the known specific technical solutions or characteristics in the solution is not described in detail here. It should be pointed out that for those skilled in the art, several modifications and improvements can be made without departing from the technical solution of the present invention, which should also be regarded as the protection scope of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A differential privacy stochastic gradient descent method based on VGP and SU, characterized in that: The following steps are involved: S1, initialize the model parameter w0 and perform iterative optimization; S2. In each iteration t, if t is less than the threshold s, the VGP process is called, otherwise: S21. Extract a batch of samples B from the data set with a specified probability t ; S22, for each sample x i , calculate and normalize the gradient g t (x i ); S23, calculating the mean of the normalized gradient of the sample and adding Gaussian noise; S3, using the calculated perturbation gradient g t To update the model parameters w new ; S4, call SU process; S5. Return the final model parameter w after the iteration. t .

2. A differential privacy stochastic gradient descent method based on VGP and SU according to claim 1, characterized in that: The VGP process in step S2 includes the following steps: A1. Randomly extract a batch of samples B from the data set t , and calculate each sample xi relative to the current model parameter w t The gradient g t (x i ); A2, decompose the gradient into two components, an orthogonal component perpendicular to the baseline gradient and a parallel component parallel to the baseline gradient; A3. Normalize these two components to ensure the consistency of scale; A4, adding Gaussian noise with a specific distribution to the orthogonal components; A5. Combine the normalized parallel components and the perturbed orthogonal components to obtain the final noise gradient for updating the model.

3. According to claim 1, a differential privacy stochastic gradient descent method based on VGP and SU is characterized in that: The calling SU process in step S4 includes the following steps: B1. Extract a batch of samples from the data set with a specified probability: B2. Calculate the current model parameters w new and the last iteration parameter w t-1 The associated loss value J(w new ) and J(w t-1 ); B3. Calculate the loss change ΔE, trim it and add Gaussian noise to increase randomness; B4. If the adjusted loss change ΔE is lower than a given threshold, the new model parameter w is accepted new And update the model, otherwise, the current parameter w t-1 remain unchanged; B5. Use the current gradient g t Update the reference vector and return the updated model parameters w t .