Training Method and Device for a Data Generation System Based on Differential Privacy
By introducing a gradient descent method of differential privacy in GAN, combining self-encoding network and discriminator, the shortcomings of generation effects and privacy security in GAN training methods are solved, and more secure and effective data generation is achieved.
Patent Information
- Application Number
- CN202111082998.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-06
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-05-06
AI Technical Summary
The existing GAN training methods have shortcomings in the generation effect and privacy security of the generative model, and it is difficult to ensure the privacy and security of data and the improvement of the generation effect.
The training method of a data generation system based on differential privacy is adopted. Through the combination of the self-encoding network and the discriminator, the gradient descent of differential privacy is used to introduce noise into the self-encoding network and the discriminator, and the model parameters are adjusted to achieve safer and more effective data generation.
By introducing differential privacy, it is difficult to back-reflect or identify the information of training samples based on public models, providing privacy protection while improving the effectiveness and security of data generation.
Smart Images

Figure CN113642731B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application number 202010373419.7, titled "Training Method and Device for Data Generation System Based on Differential Privacy", filed on May 6, 2020. Technical Field
[0002] One or more embodiments of this specification relate to the field of computer technology, and in particular, to a training method and device for a data generation system based on differential privacy executed by a computer. Background Art
[0003] With the development of computer technology, there is a great demand for automatic data synthesis. For example, in the scenario of image recognition, a large number of images need to be automatically generated or synthesized for machine learning; in scenarios such as intelligent customer service, dialogue texts need to be automatically generated. In one case, when presenting research results based on user sample data, for the purpose of protecting user privacy, it is necessary to synthesize some simulated user sample data to replace the real user data for presentation. In other cases, it may also be necessary to automatically generate synthetic data in other formats such as audio.
[0004] Therefore, attempts are made to train some generative models through machine learning to automatically generate data. For example, in one approach, a generative adversarial network (GAN, Generative Adversarial Networks) is trained, and the generative model therein is used for data synthesis. However, in the conventional GAN training method, on the one hand, the generation effect of the generative model needs to be further improved, and on the other hand, it is vulnerable to attacks and it is difficult to ensure the privacy and security of data.
[0005] Therefore, it is hoped that there can be an improved solution that can obtain a more secure and more effective data generation system. Summary of the Invention
[0006] One or more embodiments of this specification describe a training method for a data generation system based on differential privacy to obtain a data generation system that protects privacy and is more effective.
[0007] According to a first aspect, there is provided a training method for a data generation system based on differential privacy, the data generation system including an autoencoder network and a discriminator, the method including:
[0008] Input a first real sample into the autoencoder network to obtain a first reconstructed sample;
[0009] Determine a sample reconstruction loss according to the comparison between the first real sample and the first reconstructed sample;
[0010] Generate a first synthetic sample through the autoencoder network;
[0011] Input the first real sample into the discriminator to obtain a first probability that it belongs to a real sample; and input the first synthetic sample into the discriminator to obtain a second probability that it belongs to a real sample;
[0012] For the first parameter corresponding to the discriminator, in a differentially private manner, add noise to the gradient obtained with the goal of reducing the first prediction loss, and adjust the first parameter according to the obtained first noisy gradient, where the first prediction loss is negatively correlated with the first probability and positively correlated with the second probability;
[0013] For the second parameter corresponding to the autoencoder network, in a differentially private manner, add noise to the gradient obtained with the goal of reducing the second prediction loss, and adjust the second parameter according to the obtained second noisy gradient, where the second prediction loss is positively correlated with the sample reconstruction loss, positively correlated with the first probability, and negatively correlated with the second probability.
[0014] According to one embodiment, the autoencoder network includes an encoder, a generator, and a decoder; in such a case, inputting the first real sample into the autoencoder network to obtain a first restored sample specifically includes: inputting a first original vector corresponding to the first real sample into the encoder to obtain a first feature vector reduced to a first feature space; inputting the first feature vector into the decoder to obtain the first restored sample; generating a first synthetic sample through the autoencoder network specifically includes: generating a second feature vector in the first feature space through the generator; inputting the second feature vector into the decoder to obtain the first synthetic data.
[0015] Further, in one embodiment, the encoder can be implemented as a first multi-layer perceptron, and the number of neurons in each layer decreases layer by layer; the decoder can be implemented as a second multi-layer perceptron, and the number of neurons in each layer increases layer by layer.
[0016] According to one embodiment, determine the sample reconstruction loss in the following manner: determine the vector distance between a first original vector corresponding to the first real sample and a first restored vector corresponding to the first restored sample; determine the sample reconstruction loss as being positively correlated with the vector distance.
[0017] In one embodiment, noise is added to the gradient obtained with the goal of reducing the first prediction loss, and the first parameter is adjusted according to the obtained first noise gradient. Specifically, it includes: for the first parameter, determining a first original gradient that reduces the first prediction loss; based on a preset first clipping threshold, clipping the first original gradient to obtain a first clipped gradient; using a first Gaussian distribution determined based on the first clipping threshold to determine a first Gaussian noise for implementing differential privacy; and superimposing the first Gaussian noise on the first clipped gradient to obtain the first noise gradient.
[0018] In one embodiment, noise is added to the gradient obtained with the goal of reducing the second prediction loss, and the second parameter is adjusted according to the obtained second noise gradient. Specifically, it includes: for the second parameter, determining a second original gradient that reduces the second prediction loss; based on a preset second clipping threshold, clipping the second original gradient to obtain a second clipped gradient; using a second Gaussian distribution determined based on the second clipping threshold to determine a second Gaussian noise for implementing differential privacy; and superimposing the second Gaussian noise on the second clipped gradient to obtain the second noise gradient.
[0019] Further, the second parameter can be divided into an encoder parameter, a generator parameter, and a decoder parameter; in one embodiment, through gradient backpropagation, a third original gradient corresponding to the decoder parameter, a fourth original gradient corresponding to the encoder parameter, and a fifth original gradient corresponding to the generator parameter can be determined respectively; using the method of differential privacy, noise is added to the third original gradient, the fourth original gradient, and the fifth original gradient respectively to obtain corresponding third noise gradient, fourth noise gradient, and fifth noise gradient; using the third noise gradient to adjust the decoder parameter; using the fourth noise gradient to adjust the encoder parameter; and using the fifth noise gradient to adjust the generator parameter.
[0020] In another embodiment, after respectively determining a third original gradient corresponding to the decoder parameter, a fourth original gradient corresponding to the encoder parameter, and a fifth original gradient corresponding to the generator parameter through gradient backpropagation, using the method of differential privacy, noise is added to the third original gradient to obtain a corresponding third noise gradient; using the third noise gradient to adjust the decoder parameter; using the fourth original gradient to adjust the encoder parameter; and using the fifth original gradient to adjust the generator parameter.
[0021] In various embodiments, the first real sample can be a picture sample, an audio sample, a text sample, or a business object sample.
[0022] According to a second aspect, there is provided a training apparatus for a differential privacy-based data generation system, the data generation system including an autoencoder network and a discriminator, the apparatus including:
[0023] A restored sample acquisition unit configured to input a first real sample into the autoencoder network to obtain a first restored sample;
[0024] A reconstruction loss determination unit configured to determine a sample reconstruction loss according to a comparison between the first real sample and the first restored sample;
[0025] A synthesized sample acquisition unit configured to generate a first synthesized sample through the autoencoder network;
[0026] A probability acquisition unit configured to input the first real sample into the discriminator to obtain a first probability that it belongs to a real sample; and input the first synthesized sample into the discriminator to obtain a second probability that it belongs to a real sample;
[0027] A first parameter adjustment unit configured to, for a first parameter corresponding to the discriminator, add noise to a gradient obtained with the goal of reducing a first prediction loss in a differential privacy manner, and adjust the first parameter according to the obtained first noise gradient, where the first prediction loss is negatively correlated with the first probability and positively correlated with the second probability;
[0028] A second parameter adjustment unit configured to, for a second parameter corresponding to the autoencoder network, add noise to a gradient obtained with the goal of reducing a second prediction loss in a differential privacy manner, and adjust the second parameter according to the obtained second noise gradient, where the second prediction loss is positively correlated with the sample reconstruction loss, positively correlated with the first probability, and negatively correlated with the second probability.
[0029] According to a third aspect, there is provided a computer-readable storage medium having stored thereon a computer program which, when executed on a computer, causes the computer to execute the method of the first aspect.
[0030] According to a fourth aspect, there is provided a computing device including a memory and a processor, characterized in that the memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.
[0031] Through the methods and devices provided in the embodiments of this specification, a generative model in a conventional GAN is implemented through an autoencoder network. This autoencoder network can be assisted in training through an encoding process that restores real samples, thereby obtaining synthetic data that highly simulates real samples. Moreover, during the training process, through the differential privacy gradient descent method, differential privacy is introduced into the autoencoder network and the discriminator respectively, resulting in a data generation system with differential privacy characteristics. Due to the introduction of differential privacy, it is difficult to reverse-infer or identify the information of training samples based on the publicly available model, providing privacy protection for the model. In this way, a more effective and safer data generation system is obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0033] Figure 1 Shows a schematic structural diagram of a data generation system according to the technical concept of this specification;
[0034] Figure 2 Shows a flowchart of a training method for a differential privacy-based data generation system according to an embodiment;
[0035] Figure 3 Shows a schematic structural diagram of an encoder and a decoder according to an embodiment;
[0036] Figure 4 Shows a schematic block diagram of a training device for a data generation system according to an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The solutions provided in this specification will be described below with reference to the accompanying drawings.
[0038] Figure 1 Shows a schematic structural diagram of a data generation system according to the technical concept of this specification. As Figure 1As shown, the data generation system as a whole includes an autoencoder network 100 and a discriminator 200. The autoencoder network 100 may include an encoder 110, a generator 120, and a decoder 130. The encoder 110 is used to encode the high-dimensional feature vector of the input real sample data x into a sample vector E(x) in the low-dimensional representation space. The generator 120 is used to generate a noise vector G(z) in the above low-dimensional representation space based on the noise z. The decoder 130 is used to decode the corresponding sample data based on the vector in the low-dimensional representation space. When the low-dimensional sample vector E(x) corresponding to the real sample data x is input into the decoder 130, the decoder outputs the restored sample data x'; when the noise vector G(z) is input into the decoder 130, the decoder outputs the synthesized sample data s.
[0039] The discriminator 200 is used to discriminate whether the input sample data is real sample data or synthesized sample data. When the above real sample data x is input into the discriminator 200, the discriminator can output the probability P1 that it is real data; when the above synthesized data s is input into the discriminator 200, the discriminator can output the probability P2 that it is real data.
[0040] The above generator 120, decoder 130, and discriminator 200 together constitute a generative adversarial network GAN. Specifically, the training objective of the discriminator is to try to distinguish between real samples and synthesized samples, that is, it is hoped that the above probability P1 is as large as possible and the probability P2 is as small as possible. The training objective of the generator together with the decoder is to generate synthesized sample data that is as realistic as possible, making it difficult for the discriminator to distinguish. Therefore, the training objective of the generator and the decoder is to make the restored sample data x' as close as possible to the real sample data x, and at the same time make the above probability P1 as small as possible and the probability P2 as large as possible. In this way, through the adversarial training of the decoder and the discriminator, the ability of the decoder to generate synthesized data is gradually improved.
[0041] Furthermore, in order to enhance the privacy security of the model, differential privacy can be introduced into the above GAN network, especially in the decoder 130 and the discriminator 200. Specifically, during the adversarial training process, differential privacy-based gradient descent can be used to add noise to the gradient, thereby obtaining a differential privacy-based decoder and a differential privacy-based discriminator. In this way, it can be avoided that when the model is attacked, the training samples are deduced from the trained model, protecting the security of privacy data.
[0042] The following describes the specific implementation process of the above concept.
[0043] Figure 2The flowchart shows a training method for a differential privacy-based data generation system according to an embodiment. It can be understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. The following combines Figure 1 the architecture of the data generation system shown and Figure 2 the method flow shown to describe the training process of the differential privacy-based data generation system.
[0044] First, in step 21, the first real sample x is input into the autoencoder network to obtain the first restored sample x'.
[0045] In different embodiments, the above-mentioned first real sample x can be sample data in various different forms. For example, in the picture synthesis scenario, the first real sample can be a picture; in the text question-answering scenario, the above-mentioned first real sample can be a piece of text; in the speech synthesis scenario, the above-mentioned first real sample can be a segment of audio. In other examples, the first real sample can also be some business object samples, such as user samples, merchant samples, interaction event samples, and so on.
[0046] Generally, the first real sample x can be represented by the vector F(x), and this vector F(x) is called the first original vector. For example, when the first real sample x is a picture, the first original vector F(x) corresponds to the vector composed of pixel features in the picture; when the first real sample x is audio, the first original vector F(x) corresponds to the vector composed of audio spectrum features; in other examples, the first original vector can be correspondingly obtained to represent the first real sample.
[0047] When the first original vector corresponding to the first real sample is input into the autoencoder network, the autoencoder network can perform encoding and decoding processing on this first original vector and output the first restored sample.
[0048] Specifically, in one embodiment, the autoencoder network adopts the Figure 1 structure shown, which includes an encoder 110, a generator 120, and a decoder 130. In such a case, in step 21, the first original vector F(x) corresponding to the first real sample x is input into the encoder 110, and the encoder 110 performs dimensionality reduction processing on this first original vector F(x) to obtain the first feature vector E(x) in the reduced-dimensional representation space K. This first feature vector E(x) is further input into the decoder 130. The structure of the decoder 130 is symmetric to that of the encoder 110, and its algorithm and model parameters are correspondingly associated with those in the encoder 130 (for example, it is its inverse operation). Therefore, the decoder 130 can restore the first real sample x according to this first feature vector E(x) and output the first restored sample x'.
[0049] Figure 3Shows a structural schematic diagram of an encoder and a decoder according to an embodiment. As Figure 3 shown, the encoder 110 and the decoder 130 can each be implemented as a multi-layer perceptron, which contains multiple neural network layers. The difference is that in the encoder 110, the number of neurons in each layer decreases layer by layer, that is, the dimensions of each layer decrease layer by layer, so as to compress the dimensions of the input first original vector F(x) layer by layer, and the first feature vector E(x) in the representation space K is output at the output layer, also known as the representation vector. The dimension d of the representation space K is much smaller than the dimension D of the input first original vector, so as to realize the dimensionality reduction of the input original vector. For example, a first original vector of several hundred dimensions can be compressed into a coded vector of dozens of dimensions or even several dimensions.
[0050] In the decoder 130, the number of neurons in each layer increases layer by layer, that is, the dimensions of each layer increase layer by layer, so as to restore the dimensions of the low-dimensional first feature vector E(x) layer by layer, and a vector with the same dimension as the first original vector F(x) is obtained at the output layer as the restoration vector of the first restored sample x'.
[0051] It can be understood that the representation vector (such as the first feature vector E(x)) in the representation space K performs dimensionality reduction on the input original vector (such as the first original vector F(x)). The smaller the information loss of this dimensionality reduction operation, or in other words, the higher the information content of the representation vector in the representation space K, the easier it is for the decoder to restore the input real sample, that is, the higher the similarity between the restored sample and the real sample. This property can be used to assist in training the autoencoder network later.
[0052] It should be understood that although the exemplary structures of the encoder and the decoder are described above, their specific implementation manners can be various. For example, when processing picture sample data, the encoder can also correspondingly include several convolutional layers, and the decoder includes several deconvolutional layers, and so on. The specific designs of the encoder and the decoder can have various variants depending on the form of the sample data, and are not limited here.
[0053] In the above manner, the autoencoder network restores the input first real sample to obtain the first restored sample. Then, in step 22, according to the comparison between the first real sample and the first restored sample, the sample reconstruction loss Lr is determined.
[0054] In one embodiment, the first original vector F(x) corresponding to the first real sample x and the first restoration vector corresponding to the first restored sample can be compared to obtain the vector distance between the two vectors, for example, Euclidean distance, cosine distance, etc. Thus, the sample reconstruction loss Lr can be determined to be positively correlated with this vector distance. That is to say, the smaller the vector distance between the first original vector and the first restoration vector, the smaller the data difference and the smaller the sample reconstruction loss.
[0055] In another embodiment, the first real sample and the first restored sample can be compared to obtain the similarity between the two. For example, the similarity can be determined according to the dot product result between the first original vector and the first restored vector. Thus, the sample reconstruction loss Lr can also be determined to be negatively correlated with the above similarity. That is, the greater the similarity, the smaller the sample reconstruction loss.
[0056] The sample reconstruction loss Lr determined above can be used to measure the reconstruction ability of the autoencoder network, especially the decoder therein, for samples, and thus is used to train the autoencoder network.
[0057] On the other hand, in step 23, a first synthetic sample is generated through the autoencoder network.
[0058] In one embodiment, the autoencoder network adopts Figure 1 the structure shown, which includes an encoder 110, a generator 120, and a decoder 130. In such a case, in step 23, through the generator 120, a second feature vector G(z) that simulates the real representation vector is generated in the foregoing representation space K; then, the second feature vector G(z) is input into the decoder 130 to obtain the first synthetic data s.
[0059] In one embodiment, the generator 120 obtains the data distribution of the representation vectors of multiple real samples output by the encoder 110, and samples in this data distribution space with a certain probability, thereby generating the second feature vector G(z). In another embodiment, a noise signal is input to the generator 120, and the generator 120 generates the second feature vector G(z) in the above representation space K based on the noise signal.
[0060] The second feature vector G(z) generated in the above manner can be used to simulate the representation vector of the real sample in the representation space K. Therefore, when the second feature vector G(z) is input into the decoder 130, the decoder 130 can decode it in the same way as it processes the foregoing real representation vector E(x), thereby obtaining a synthetic sample s in the same data form as the real sample.
[0061] It should be understood that the above step 23 and the foregoing steps 21-22 can be executed in any reasonable relative order, such as in parallel, before or after them.
[0062] After that, in step 24, the first real sample x and the first synthetic sample s are respectively input into the discriminator, so as to obtain the first probability P1 that the first real sample belongs to the real sample, and the second probability P2 that the first synthetic sample s belongs to the real sample.
[0063] It should be understood that the discriminator is used to distinguish whether the input sample data is real or synthetic. Specifically, the discriminator gives the discrimination result by outputting a prediction probability. Usually, the discriminator outputs the probability that the sample data is a real sample. In such a case, the above-mentioned first probability P1 is the output probability of the discriminator after inputting the first real sample x; the above-mentioned second probability P2 is the output probability of the discriminator after inputting the first synthetic sample s.
[0064] In another example, the discriminator can also output the probability that the sample data is a synthetic sample. In such a case, the above-mentioned first probability P1 can be understood as 1 - P1', where P1' is the output probability of the discriminator for the first real sample x; the above-mentioned second probability P2 can be understood as 1 - P2', where P2' is the output probability of the discriminator for the first synthetic sample s.
[0065] Based on the sample reconstruction loss Lr obtained in step 22, and the first probability P1 and the second probability P2 obtained in step 24, the first prediction loss L1 for training the discriminator and the second prediction loss L2 for training the autoencoder network can be determined respectively.
[0066] It can be understood that the training objective of the discriminator is to try to distinguish between real samples and synthetic samples. Therefore, for the discriminator, it is desired that the above-mentioned first probability P1 is as large as possible and the second probability P2 is as small as possible. Therefore, the first prediction loss L1 can be set to be negatively correlated with the first probability P1 and positively correlated with the second probability P2. In this way, the direction in which the first prediction loss L1 decreases is the direction of increasing the first probability P1 and decreasing the second probability P2.
[0067] More specifically, in one embodiment, the first prediction loss can be set as:
[0068]
[0069] where i is the real sample, P1 is the first probability corresponding to each real sample, j is the synthetic sample, and P2 is the second probability corresponding to each synthetic sample.
[0070] On the other hand, the training objective of the autoencoder network is that for real samples, it is hoped to reconstruct more similar restored samples, and it is hoped that the discriminator cannot distinguish between real samples and synthetic samples generated by the decoder. Therefore, for the autoencoder network, it is hoped that the aforementioned sample reconstruction loss Lr is as small as possible, and it is hoped that the above first probability P1 is as small as possible and the second probability P2 is as large as possible. Therefore, the second prediction loss L2 can be set to be positively correlated with the sample reconstruction loss and the first probability P1, and negatively correlated with the second probability P2. In this way, the direction in which the second prediction loss L2 decreases is the direction of decreasing the sample reconstruction loss, decreasing the first probability P1, and increasing the second probability P2.
[0071] More specifically, in one embodiment, the second prediction loss can be set as:
[0072]
[0073] In this way, through the above method, the first prediction loss for the discriminator and the second prediction loss for the autoencoder network are obtained. From the definitions of the above first prediction loss L1 and second prediction loss L2, it can be seen that the training objectives of the autoencoder network and the discriminator are adversarial. Next, based on the first and second prediction losses, the parameter gradients that reduce the losses can be determined, so as to train the discriminator and the autoencoder network respectively.
[0074] Innovatively, in the embodiments of this specification, during the training process, the method of differential privacy is used to add noise to the gradients, and the data generation system is trained according to the gradients containing noise. That is, in step 25, for the first parameter corresponding to the discriminator, the method of differential privacy is used to add noise to the gradient obtained with the goal of reducing the first prediction loss L1, and the first parameter is adjusted according to the obtained first noise gradient; in step 26, for the second parameter corresponding to the autoencoder network, the method of differential privacy is used to add noise to the gradient obtained with the goal of reducing the second prediction loss L2, and the second parameter is adjusted according to the obtained second noise gradient. In this way, the characteristics of differential privacy are introduced into the discriminator and the autoencoder network respectively.
[0075] Differential privacy is a means in cryptography, aiming to provide a way to maximize the accuracy of data queries while minimizing the chance of identifying its records when querying a statistical database. Let there be a random algorithm M, and PM be the set of all possible outputs of M. For any two neighboring datasets D and D' and any subset SM of PM, if the random algorithm M satisfies: Pr[M(D) ∈ SM] <= e ε×Pr[M(D') ∈ SM], the algorithm M is said to provide ε-differential privacy protection, where the parameter ε is called the privacy protection budget, which is used to balance the degree of privacy protection and accuracy. ε can usually be preset in advance. The closer ε is to 0, e ε the closer it is to 1, the closer the processing results of the randomized algorithm for two neighboring datasets D and D' are, and the stronger the privacy protection degree is.
[0076] The implementation methods of differential privacy include, noise mechanisms, exponential mechanisms, etc. In order to introduce differential privacy into the data generation system, according to the embodiments of this specification, a noise mechanism is used here to implement differential privacy by adding noise to the parameter gradient. According to the noise mechanism, the noise can be embodied as Laplace noise, Gaussian noise, and so on. According to one embodiment, in step 25, differential privacy is introduced into the discriminator by adding Gaussian noise to the gradient determined based on the first prediction loss. The specific process may include the following steps.
[0077] First, for the first parameter corresponding to the discriminator, the first original gradient that makes the first prediction loss L1 decrease can be determined according to the aforementioned first prediction loss L1; then, based on a preset clipping threshold, the first original gradient is clipped to obtain the first clipped gradient; then, a first Gaussian noise for implementing differential privacy is determined by using a Gaussian distribution determined based on the first clipping threshold, where the variance of the Gaussian distribution is positively correlated with the square of the first clipping threshold; then, the first Gaussian noise thus obtained is superimposed on the aforementioned first clipped gradient to obtain a first noise gradient for updating the first parameter of the discriminator.
[0078] More specifically, as an example, assume that for the training set X composed of the first real sample x and the first synthetic sample s, the first original gradient obtained for the discriminator is:
[0079]
[0080] where, L1(θ D , X) represents the aforementioned first prediction loss, and θ D is the parameter in the discriminator, that is, the first parameter.
[0081] As described above, adding noise that implements differential privacy to the original gradient can be achieved by means such as Laplace noise, Gaussian noise, etc. In one embodiment, taking Gaussian noise as an example, the original gradient can be clipped based on a preset clipping threshold to obtain a clipped gradient, and then based on this clipping threshold and a predetermined noise scaling factor (a hyperparameter set in advance), Gaussian noise for implementing differential privacy can be determined, and then the clipped gradient and the Gaussian noise can be fused (such as summed) to obtain a gradient containing noise. It can be understood that in this way, on the one hand, the original gradient is clipped, and on the other hand, the clipped gradients are superimposed, so as to perform differential privacy processing on the gradient that satisfies Gaussian noise.
[0082] For example, the first original gradient is clipped to:
[0083]
[0084] Wherein, represents the clipped gradient, that is, the first clipped gradient, C1 represents the first clipping threshold, ||g D (X)|| 2 represents the second-order norm of g D (X). That is to say, when the original gradient is less than or equal to the clipping threshold C1, the original gradient is retained, and when the original gradient is greater than the clipping threshold C1, the original gradient is clipped to the corresponding size according to the ratio greater than the clipping threshold C1.
[0085] Add the first Gaussian noise to the first clipped gradient to obtain the first noisy gradient containing noise, for example:
[0086]
[0087] Wherein, represents the first noisy gradient; represents a first Gaussian noise whose probability density conforms to a Gaussian distribution with a mean of 0 and a variance of σ 2 C1 2 I; σ represents the above-mentioned noise scaling factor, which is a hyperparameter set in advance and can be set as needed; C1 is the above-mentioned first clipping threshold; I represents the indicator function, which can take 0 or 1. For example, it can be set to take 1 in the even rounds of multiple rounds of training and 0 in the odd rounds.
[0088] Then, the first noisy gradient after adding Gaussian noise can be used, with the goal of minimizing the aforementioned prediction loss L1, to adjust the first parameter θ D of the discriminator to:
[0089]
[0090] Among them, η represents the learning step size, or learning rate, which is a pre-set hyperparameter, such as 0.5, 0.3, etc. When adding Gaussian noise to the gradient to satisfy differential privacy, the adjustment of the model parameters of the above discriminator satisfies differential privacy.
[0091] On the other hand, in step 26, for the autoencoder network, in a similar manner, by adding noise to the gradient, the parameters of the autoencoder network can be adjusted in a differentially private manner. Specifically, in one embodiment, for the second parameter θ of the autoencoder network A , determine the second original gradient g that makes the aforementioned second prediction loss L2 decrease A (X), for example:
[0092]
[0093] Then, based on the preset second clipping threshold C2, clip the second original gradient to obtain the second clipped gradient The clipping method is similar to the above formula (4), where the second clipping threshold C2 is set independently of the first clipping threshold C1 and can be the same or different. Then, using the second Gaussian distribution determined based on the second clipping threshold, determine the second Gaussian noise for implementing differential privacy Superimpose the second Gaussian noise on the second clipped gradient to obtain the second noise gradient Thus, according to the second noise gradient, the corresponding second parameter of the autoencoder network can be adjusted.
[0094] The above describes the method of adding Gaussian noise to the second original gradient of the autoencoder network and then adjusting the second parameter. Further, in one embodiment, as Figure 1 shown, the autoencoder network further includes an encoder 110, a generator 120, and a decoder 130. Correspondingly, the above second parameter can be further divided into encoder parameters, generator parameters, and decoder parameters, and each part of the parameters corresponds to the original parameter gradients of each part. When adding noise to the second original gradient, noise can be added to the original parameter gradients of each part, or only to the original parameter gradients of some parts, such as the original parameter gradients corresponding to the decoder.
[0095] Specifically, in one embodiment, in step 26, through gradient backpropagation, the original parameter gradients of each parameter part in the autoencoder network can be determined respectively, including the third original gradient corresponding to the decoder parameters, the fourth original gradient corresponding to the encoder parameters, and the fifth original gradient corresponding to the generator parameters.
[0096] Then, in the manner of differential privacy, noise is added to the third original gradient, the fourth original gradient, and the fifth original gradient respectively to obtain the corresponding third noisy gradient, fourth noisy gradient, and fifth noisy gradient. Among them, the manner of adding noise can refer to the process of adding Gaussian noise described above. Thus, the decoder parameters can be adjusted using the third noisy gradient; the encoder parameters can be adjusted using the fourth noisy gradient; and the generator parameters can be adjusted using the fifth noisy gradient. In this way, the differential privacy feature is introduced into the autoencoder network.
[0097] According to another embodiment, in step 26, after respectively determining the third original gradient corresponding to the decoder parameters, the fourth original gradient corresponding to the encoder parameters, and the fifth original gradient corresponding to the generator parameters through gradient backpropagation, only for the third original gradient, noise is added thereto in the manner of differential privacy to obtain the corresponding third noisy gradient. Then, the decoder parameters are adjusted using the third noisy gradient, thereby introducing the differential privacy feature into the decoder. For the encoder and the generator, the corresponding original parameter gradients can be used for updating, that is, the encoder parameters are adjusted using the fourth original gradient; and the generator parameters are adjusted using the fifth original gradient.
[0098] It should be understood that the decoder is the core module in the autoencoder network. The real samples are restored through this decoder, and the synthetic samples are generated through this decoder. Therefore, introducing differential privacy into the decoder makes the entire autoencoder network have the differential privacy feature, which can also achieve the effect of making the entire data generation system have the differential privacy feature.
[0099] It should be noted that in actual operation, the training of the discriminator in step 25 and the training of the autoencoder network in step 26 can be alternately iterated. For example, after the discriminator is iteratively updated m times using a sample set including real samples and generated samples, the autoencoder network is then iteratively updated n times, and this is repeatedly executed. The update order and iteration manner of the discriminator and the autoencoder network are not limited herein.
[0100] After repeatedly updating the discriminator and the autoencoder network in the above manner until a predetermined end condition is reached (such as iterating a predetermined number of times, the parameters converge, etc.), the trained data generation system can be obtained. When using this data generation system to generate sample data, only need to use the generator therein to generate a noise vector and decode it with the decoder to obtain synthetic sample data simulating real samples.
[0101] Reviewing the above process, the generative model in a conventional GAN is implemented through an autoencoder network, which can be assisted in training by the encoding process of restoring real samples, so as to obtain synthetic data that highly simulates real samples. Moreover, during the training process, differential privacy is introduced into the autoencoder network and the discriminator respectively by means of differential privacy gradient descent, resulting in a data generation system with differential privacy characteristics. Due to the introduction of differential privacy, it is difficult to reverse-infer or identify the information of training samples based on the publicly available model, providing privacy protection for the model. In this way, a more effective and secure data generation system is obtained.
[0102] According to an embodiment of another aspect, there is also provided a training device for a data generation system based on differential privacy. The data generation system includes an autoencoder network and a discriminator, and the training device can be deployed in any device, equipment, platform, or device cluster with computing and processing capabilities. Figure 4 The schematic block diagram of a training device for a data generation system according to an embodiment is shown. As Figure 4 shown, the training device 400 includes:
[0103] A restored sample acquisition unit 41, configured to input a first real sample into the autoencoder network to obtain a first restored sample;
[0104] A reconstruction loss determination unit 42, configured to determine a sample reconstruction loss according to the comparison between the first real sample and the first restored sample;
[0105] A synthetic sample acquisition unit 43, configured to generate a first synthetic sample through the autoencoder network;
[0106] A probability acquisition unit 44, configured to input the first real sample into the discriminator to obtain a first probability that it belongs to a real sample; and input the first synthetic sample into the discriminator to obtain a second probability that it belongs to a real sample;
[0107] A first parameter adjustment unit 45, configured to, for the first parameters corresponding to the discriminator, add noise to the gradient obtained with the goal of reducing the first prediction loss in a differential privacy manner, and adjust the first parameters according to the obtained first noise gradient, where the first prediction loss is negatively correlated with the first probability and positively correlated with the second probability;
[0108] A second parameter adjustment unit 46, configured to, for the second parameters corresponding to the autoencoder network, add noise to the gradient obtained with the goal of reducing the second prediction loss in a differential privacy manner, and adjust the second parameters according to the obtained second noise gradient, where the second prediction loss is positively correlated with the sample reconstruction loss, positively correlated with the first probability, and negatively correlated with the second probability.
[0109] According to one embodiment, the autoencoder network includes an encoder, a generator, and a decoder. In such a case, the restored sample acquisition unit 41 may be configured to: input the first original vector corresponding to the first real sample into the encoder to obtain a first feature vector reduced to the first feature space; input the first feature vector into the decoder to obtain the first restored sample; the synthetic sample acquisition unit 43 may be configured to: generate a second feature vector in the first feature space through the generator; input the second feature vector into the decoder to obtain the first synthetic data.
[0110] Furthermore, in one embodiment, the encoder may be implemented as a first multi-layer perceptron, and the number of neurons in each layer thereof decreases layer by layer; the decoder may be implemented as a second multi-layer perceptron, and the number of neurons in each layer thereof increases layer by layer.
[0111] According to one embodiment, the reconstruction loss determination unit 42 is specifically configured to: determine the vector distance between the first original vector corresponding to the first real sample and the first restored vector corresponding to the first restored sample; determine the sample reconstruction loss to be positively correlated with the vector distance.
[0112] In one embodiment, the first parameter adjustment unit 45 is specifically configured to: for the first parameter, determine a first original gradient that reduces the first prediction loss; based on a preset first clipping threshold, clip the first original gradient to obtain a first clipped gradient; use a first Gaussian distribution determined based on the first clipping threshold to determine a first Gaussian noise for implementing differential privacy; superimpose the first Gaussian noise on the first clipped gradient to obtain the first noise gradient.
[0113] Similarly, the second parameter adjustment unit 46 may be specifically configured to: for the second parameter, determine a second original gradient that reduces the second prediction loss; based on a preset second clipping threshold, clip the second original gradient to obtain a second clipped gradient; use a second Gaussian distribution determined based on the second clipping threshold to determine a second Gaussian noise for implementing differential privacy; superimpose the second Gaussian noise on the second clipped gradient to obtain the second noise gradient.
[0114] More specifically, in one embodiment, the second parameter specifically includes an encoder parameter, a generator parameter, and a decoder parameter. In one example, the second parameter adjustment unit 46 is specifically configured to: respectively determine a third original gradient corresponding to the decoder parameter, a fourth original gradient corresponding to the encoder parameter, and a fifth original gradient corresponding to the generator parameter through gradient backpropagation; add noise to the third original gradient, the fourth original gradient, and the fifth original gradient respectively in a differential privacy manner to obtain corresponding third noise gradients, fourth noise gradients, and fifth noise gradients; use the third noise gradient to adjust the decoder parameter; use the fourth noise gradient to adjust the encoder parameter; and use the fifth noise gradient to adjust the generator parameter.
[0115] In another example, the second parameter adjustment unit 46 is specifically configured to: respectively determine a third original gradient corresponding to the decoder parameter, a fourth original gradient corresponding to the encoder parameter, and a fifth original gradient corresponding to the generator parameter through gradient backpropagation; add noise to the third original gradient in a differential privacy manner to obtain a corresponding third noise gradient; use the third noise gradient to adjust the decoder parameter; use the fourth original gradient to adjust the encoder parameter; and use the fifth original gradient to adjust the generator parameter.
[0116] In various different embodiments, the first true sample can be a picture sample, an audio sample, a text sample, or a business object sample.
[0117] It is worth noting that Figure 4 the shown device 400 is a device embodiment corresponding to Figure 2 the method embodiment shown. Figure 2 The corresponding descriptions in the method embodiment shown also apply to the device 400 and will not be repeated here.
[0118] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method described in conjunction with Figure 2 .
[0119] According to an embodiment of still another aspect, there is also provided a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in conjunction with Figure 2 is implemented.
[0120] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0121] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the technical concept of this specification. It should be understood that the above is only the specific embodiments of the technical concept of this specification and is not used to limit the protection scope of the technical concept of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions in the embodiments of this specification should be included within the protection scope of the technical concept of this specification.
Claims
1. A training method for a differential privacy-based data generation system, the data generation system including an autoencoder network and a discriminator, the method comprises: Inputting a first real sample into the autoencoder network to obtain a first reconstructed sample; The first real sample includes one of the following: a picture sample, an audio sample, a text sample; Determining a sample reconstruction loss according to the comparison between the first real sample and the first reconstructed sample; Generating a first synthetic sample based on noise through the autoencoder network; Inputting the first real sample into the discriminator to obtain a first probability that it belongs to a real sample; and inputting the first synthetic sample into the discriminator to obtain a second probability that it belongs to a real sample; For the first parameter corresponding to the discriminator, in a differential privacy manner, adding noise to the gradient obtained with the goal of reducing the first prediction loss, and adjusting the first parameter according to the obtained first noisy gradient, where the first prediction loss is negatively correlated with the first probability and positively correlated with the second probability; For the second parameter corresponding to the autoencoder network, in a differential privacy manner, adding noise to the gradient obtained with the goal of reducing the second prediction loss, and adjusting the second parameter according to the obtained second noisy gradient, where the second prediction loss is positively correlated with the sample reconstruction loss, positively correlated with the first probability, and negatively correlated with the second probability.
2. The method according to claim 1, wherein, The autoencoder network includes an encoder, a generator, and a decoder; The encoder is implemented as a first multi-layer perceptron, and the number of neurons in each layer decreases layer by layer; the decoder is implemented as a second multi-layer perceptron, and the number of neurons in each layer increases layer by layer.
3. The method according to claim 1, wherein, Determining the sample reconstruction loss includes: Determining the vector distance between a first original vector corresponding to the first real sample and a first reconstructed vector corresponding to the first reconstructed sample; Determining the sample reconstruction loss as being positively correlated with the vector distance.
4. The method according to claim 1, wherein, Adding noise to the gradient obtained with the goal of reducing the first prediction loss and adjusting the first parameter according to the obtained first noisy gradient includes: For the first parameter, determining a first original gradient that reduces the first prediction loss; Based on a preset first clipping threshold, clipping the first original gradient to obtain a first clipped gradient; Using a first Gaussian distribution determined based on the first clipping threshold to determine a first Gaussian noise for implementing differential privacy; Superimposing the first Gaussian noise and the first clipped gradient to obtain the first noisy gradient.
5. The method according to claim 1, wherein, Adding noise to the gradient obtained with the goal of reducing the second prediction loss and adjusting the second parameter according to the obtained second noisy gradient includes: For the second parameter, determining a second original gradient that reduces the second prediction loss; Based on a preset second clipping threshold, clipping the second original gradient to obtain a second clipped gradient; Determine the second Gaussian noise for implementing differential privacy by using the second Gaussian distribution determined based on the second clipping threshold; Superimpose the second Gaussian noise and the second clipped gradient to obtain the second noisy gradient.
6. The method according to claim 2, wherein, the second parameters include encoder parameters, generator parameters, and decoder parameters; Adding noise to the gradient obtained with the goal of reducing the second prediction loss and adjusting the second parameters according to the obtained second noisy gradient includes: Through gradient backpropagation, respectively determine the third original gradient corresponding to the decoder parameters, the fourth original gradient corresponding to the encoder parameters, and the fifth original gradient corresponding to the generator parameters; Using the method of differential privacy, add noise to the third original gradient, the fourth original gradient, and the fifth original gradient respectively to obtain the corresponding third noisy gradient, fourth noisy gradient, and fifth noisy gradient; Use the third noisy gradient to adjust the decoder parameters; use the fourth noisy gradient to adjust the encoder parameters; use the fifth noisy gradient to adjust the generator parameters.
7. The method according to claim 2, wherein, the second parameters include encoder parameters, generator parameters, and decoder parameters; Adding noise to the gradient obtained with the goal of reducing the second prediction loss and adjusting the second parameters according to the obtained second noisy gradient includes: Through gradient backpropagation, respectively determine the third original gradient corresponding to the decoder parameters, the fourth original gradient corresponding to the encoder parameters, and the fifth original gradient corresponding to the generator parameters; Using the method of differential privacy, add noise to the third original gradient to obtain the corresponding third noisy gradient; Use the third noisy gradient to adjust the decoder parameters; use the fourth original gradient to adjust the encoder parameters; use the fifth original gradient to adjust the generator parameters.
8. A training device for a data generation system based on differential privacy, the data generation system includes an autoencoder network and a discriminator, the device comprises: A restored sample acquisition unit configured to input a first real sample into the autoencoder network to obtain a first restored sample; The first real sample includes one of the following: a picture sample, an audio sample, a text sample; A reconstruction loss determination unit configured to determine a sample reconstruction loss according to the comparison between the first real sample and the first restored sample; A synthetic sample acquisition unit configured to generate a first synthetic sample based on noise through the autoencoder network; A probability acquisition unit configured to input the first real sample into the discriminator to obtain a first probability that it belongs to a real sample; and input the first synthetic sample into the discriminator to obtain a second probability that it belongs to a real sample; The first parameter adjustment unit is configured to, for the first parameter corresponding to the discriminator, add noise to the gradient obtained with the goal of reducing the first prediction loss in a differentially private manner, and adjust the first parameter according to the obtained first noisy gradient, where the first prediction loss is negatively correlated with the first probability and positively correlated with the second probability; The second parameter adjustment unit is configured to, for the second parameter corresponding to the autoencoder network, add noise to the gradient obtained with the goal of reducing the second prediction loss in a differentially private manner, and adjust the second parameter according to the obtained second noisy gradient, where the second prediction loss is positively correlated with the sample reconstruction loss, positively correlated with the first probability, and negatively correlated with the second probability.
9. A computing device, comprising a memory and a processor, wherein, executable code is stored in the memory, and when the processor executes the executable code, the method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Deeply differential privacy protection method based on generative adversarial network
CN107368752A
Capsule type endoscope image generation method and device and computer storage medium
CN110458904A
Video generation method combining variational auto-encoder and generative adversarial network
CN110572696A