An Optimal Method for Differential Privacy Disposal of Deep Learning in a Certain Field

By adjusting the privacy budget through dynamic optimization, the balance between privacy protection requirements and model efficiency in deep learning model training is solved, and effective privacy data protection and model training effect are achieved.

CN119760783BActive Publication Date: 2025-06-27JINAN SANZE INFORMATION SECURITY EVALUATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510251936.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-27
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

How to reasonably allocate privacy budgets during deep learning model training to ensure that privacy protection needs are met without affecting the training efficiency and accuracy of the model.

Method used

By obtaining the domain privacy data set, building a deep learning loss function, and calculating contribution gradients, dynamically optimizing the privacy budget to ensure that the privacy protection needs at different stages are met.

Benefits of technology

It realizes effective privacy data protection in deep learning model training, avoids privacy leakage, and ensures the training efficiency and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119760783B_ABST
    Figure CN119760783B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data differential privacy processing, and discloses a preferred method for data differential privacy processing in domain deep learning. The method includes: constructing a domain privacy data set into a deep learning loss function; initializing a privacy budget for generating domain privacy data; calculating the contribution gradient of the domain privacy data during the training process of the deep learning loss function, and dynamically and preferably adjusting the privacy budget of the domain privacy data based on the contribution gradient; and using the dynamically and preferably adjusted privacy budget to protect the privacy of the domain privacy data. Based on the gradient of the deep learning model parameters during the training process, the present invention calculates the contribution gradients of different domain privacy data, and then dynamically and preferably adjusts the privacy budget based on the contribution gradient and the differential privacy change rate of the domain privacy data. Under the condition of ensuring the model training effect, a privacy budget for protecting the privacy of the domain privacy data is obtained, realizing the privacy protection of the domain privacy data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data differential privacy processing, and particularly to an optimal method for data differential privacy processing in domain deep learning. Background Art

[0002] With the rapid development of deep learning technology, artificial intelligence (AI) has achieved remarkable application results in multiple fields, especially in the fields of healthcare, finance, retail, social networks, etc. A large amount of sensitive data is used to train deep learning models to improve the prediction accuracy and generalization ability of the models. However, with the popularization of artificial intelligence, data privacy issues have become increasingly serious. Especially in application scenarios involving personal privacy information, the risk of data leakage and privacy protection have become one of the main bottlenecks in technology applications. To protect user privacy, differential privacy, as a powerful data privacy protection technology, has been proposed and widely applied. Differential privacy adds noise to the data set, minimizing the impact of individual data on the final result, thus avoiding the problem of data leakage. Specifically, differential privacy ensures that it is impossible to distinguish whether a single data individual participated in the training process in the output result, avoiding the possibility of inferring sensitive information externally through the model. The application of differential privacy not only effectively protects personal data but also enables multi-party data to be jointly learned and shared while ensuring privacy, thus promoting the application of deep learning in multiple fields. However, the effect of differential privacy often depends on the privacy budget, that is, how much privacy protection resources are allocated throughout the training process. Excessive use of the privacy budget will lead to a decrease in the effectiveness of the model, while insufficient budget may lead to privacy leakage. Therefore, how to reasonably allocate the privacy budget to ensure that the privacy protection requirements in different stages are met without affecting the training efficiency and accuracy of the model has become an urgent problem to be solved. Summary of the Invention

[0003] In view of this, the present invention provides an optimal method for data differential privacy processing in domain deep learning, which optimally configures the privacy budget based on the training effect of deep learning model parameters to achieve privacy data protection.

[0004] To achieve the above object, an optimal method for data differential privacy processing in domain deep learning provided by the present invention includes the following steps:

[0005] S1: Obtain a domain privacy data set and construct it into a deep learning loss function. The domain privacy data set consists of domain privacy data and the predicted classification results of the domain privacy data.

[0006] S2: Initialize the privacy budget for generating domain privacy data.

[0007] S3: Calculate the contribution gradient of the domain privacy data during the training process of the deep learning loss function, and dynamically and preferably adjust the privacy budget of the domain privacy data based on the contribution gradient.

[0008] S4: Use the privacy budget after dynamic and preferred adjustment to protect the domain privacy data, and train the deep learning loss function based on the domain privacy data after privacy protection until the deep learning loss function converges, obtain the optimal deep learning model parameters after domain privacy protection, and construct a deep learning model based on the optimal deep learning model parameters.

[0009] As a further improved method of the present invention:

[0010] Optionally, the personal sensitive domain includes the medical field, the financial field, the consumer field, and the social communication field. The domain privacy data set includes N groups of domain privacy data, where N represents the number of domain privacy data in the domain privacy data set. The representation form of the nth group of domain privacy data is:

[0011] 。

[0012] Where:

[0013] represents the nth group of domain privacy data; successively represent the privacy data in the medical field, the financial field, the consumer field, and the social communication field in;

[0014] The domain privacy data The predicted classification result of is 。

[0015] Optionally, based on the domain privacy data set, the expression of the constructed deep learning loss function is :

[0016] 。

[0017] Where:

[0018] represents the deep learning model parameters, T represents the transpose, represents the L2 norm;

[0019] represents the predicted classification result of the domain privacy data output by the deep learning model constructed based on the deep learning model parameters .

[0020] Optionally, initialize and generate the deep learning model parameters , the deep learning loss function including and excluding the domain private data is calculated as the initial contribution deviation of the domain private data. The initial contribution deviation of the nth group of domain private data is:

[0021] .

[0022] Where:

[0023] represents the initial contribution deviation of the nth group of domain private data, represents the L1 norm, represents the deep learning loss function excluding the nth group of domain private data.

[0024] Based on the initial contribution deviation of the domain private data, the privacy budget for generating the domain private data is initialized. The privacy budget of the nth group of domain private data generated by initialization is:

[0025] .

[0026] Where:

[0027] represents the privacy budget of the nth group of domain private data generated by initialization, represents the total number of preset privacy budgets.

[0028] Optionally, the deep learning model parameters are trained iteratively during the training process of the deep learning loss function. The maximum number of training iterations of the deep learning model parameters is Max, and based on the change value of the deep learning loss function before and after the training iteration, the training gradient of the deep learning model parameters after each training iteration is generated. The deep learning model parameters after the tth training iteration are , , the training gradient of the deep learning model parameters is . Based on the training gradient, the contribution gradient of each group of domain private data during the training process of the deep learning loss function is calculated. The contribution gradient of the nth group of domain private data during the tth training iteration is:

[0029] .

[0030] Where:

[0031] represents the contribution gradient of the nth group of domain private data during the tth training iteration, represents the preset control coefficient, represents the privacy budget dynamically and preferably adjusted after the (t - 1)th training iteration of the nth group of domain private data.

[0032] Optionally, the privacy budget of the domain privacy data is dynamically and preferably adjusted based on the contribution gradient, where the privacy budget The formula for dynamic and preferred adjustment is:

[0033] .

[0034] Where:

[0035] represents the differential privacy change rate of the privacy budget , and represents a preset contribution gradient threshold.

[0036] Optionally, the privacy budget after the dynamic and preferred adjustment is used to protect the domain privacy data. The formula for protecting the domain privacy data with the privacy budget is:

[0037] .

[0038] Where:

[0039] represents a random number in a normal distribution with a mean of 0 and a standard deviation of , and e represents a random noise intensity control parameter;

[0040] represents the protection result of the nth group of domain privacy data based on the privacy budget .

[0041] Optionally, the deep learning loss function is trained based on the domain privacy data after privacy protection to obtain the training iteration result of the deep learning model parameters until the training gradient of the deep learning model parameters is less than a preset gradient threshold. If the training gradient of the deep learning model parameters is less than the preset gradient threshold, the deep learning loss function converges. The training iteration formula of the deep learning model parameters is:

[0042] ;

[0043] .

[0044] Where:

[0045] represents the learning rate of the (t + 1)-th training iteration, represents the deep learning model parameters after the (t + 1)-th training iteration.

[0046] To solve the above problems, the present invention provides an electronic device, which includes:

[0047] A memory that stores at least one instruction;

[0048] A communication interface that enables communication of an electronic device; and

[0049] A processor that executes the instructions stored in the memory to implement the preferred method for differential privacy disposal of data in the above-mentioned field of deep learning.

[0050] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in an electronic device to implement the preferred method for differential privacy disposal of data in the above-mentioned field of deep learning.

[0051] Compared with the prior art, the present invention proposes a preferred method for differential privacy disposal of data in the field of deep learning, and this technology has the following advantages:

[0052] First of all, this solution proposes a privacy protection method for training data in the training of a deep learning model in a field. Based on the differential privacy method, the impact of whether domain privacy data is included on the deep learning loss function is quantified to obtain the initial privacy budget of the domain privacy data. The greater the impact of the domain privacy data on the deep learning loss function, the higher the initial privacy budget, and the lower the degree of protection for the domain privacy data. This avoids overly affecting the model training effect of the deep learning model, and protecting the domain privacy data can effectively ensure that even if an individual in the domain privacy data set is removed or added, the query result will not change significantly, thereby protecting the personal privacy information in the domain privacy data set from being leaked and avoiding privacy leakage.

[0053] At the same time, this solution proposes a dynamic optimal adjustment method for the privacy budget. Based on the gradient of the deep learning model parameters during the training process, the contribution gradient of different domain privacy data is calculated. The higher the contribution gradient, the greater the role of the domain privacy data during the training process. Then, based on the contribution gradient and the differential privacy change rate of the domain privacy data, the privacy budget is dynamically and optimally adjusted. Under the condition of ensuring the model training effect, the privacy budget for protecting the domain privacy data is obtained. A random number is generated using the privacy budget to perform differential privacy disposal on the domain privacy data, and the training iteration of gradient descent is performed on the deep learning model parameters in combination with the change of the privacy budget. The greater the increase in the privacy budget, the lower the degree of privacy protection for the current domain privacy data, and the privacy budget needs to be increased and the change step size of the deep learning model parameters needs to be reduced. Description of the Drawings

[0054] Figure 1Schematic diagram of a preferred method for differential privacy processing of domain deep learning provided by an embodiment of the present invention.

[0055] The implementation, functional features, and advantages of the present invention will be further described in conjunction with embodiments with reference to the accompanying drawings. Specific embodiments

[0056] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0057] The embodiments of the present application provide a preferred method for differential privacy processing of domain deep learning. The execution subject of the preferred method for differential privacy processing of domain deep learning includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiments of the present application. In other words, the preferred method for differential privacy processing of domain deep learning can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.

[0058] Referring to Figure 1 , Embodiment 1 of the present invention is:

[0059] A preferred method for differential privacy processing of domain deep learning includes the following steps:

[0060] S1: Obtain a domain privacy data set, and construct a deep learning loss function from the domain privacy data set, where the domain privacy data set consists of domain privacy data and the predicted classification results of the domain privacy data.

[0061] The personal sensitive domain includes the medical field, the financial field, the consumer field, and the social communication field. The domain privacy data set includes N groups of domain privacy data, where N represents the number of domain privacy data in the domain privacy data set. The representation form of the nth group of domain privacy data is:

[0062] .

[0063] Where:

[0064] represents the nth group of domain privacy data; successively represent the privacy data that conforms to the medical field, the financial field, the consumer field, and the social communication field in

[0065] The domain privacy data The predicted classification result of is .

[0066] As an embodiment of the present invention, the predicted classification result of the domain privacy data is in the form of a classification matrix, including the predicted classification results of multiple personal sensitive domains. Among them, the predicted classification result in the medical domain is the risk assessment result of common major diseases, the predicted classification result in the financial domain is the credit risk assessment result, the predicted classification result in the consumption domain is the consumption behavior prediction result, and the predicted classification result in the social communication domain is the social behavior prediction result. The consumption behavior prediction result is the purchase probability of different types of goods, and the social behavior prediction result is the attention probability of different social sectors.

[0067] Specifically, the privacy data in the medical domain includes health status, diagnosis results, treatment plans, genomic data, and drug use history, etc. The privacy data in the financial domain includes bank account information, credit card transaction data, and insurance data information, etc. The privacy data in the consumption domain includes shopping history data, and the privacy data in the social communication domain includes social dynamics and social interest data.

[0068] Based on the domain privacy data set, the expression of the deep learning loss function constructed is :

[0069] 。

[0070] Among them:

[0071] represents the deep learning model parameters, T represents the transpose, represents the L2 norm;

[0072] represents the predicted classification result of the domain privacy data output by the deep learning model constructed based on the deep learning model parameters .

[0073] S2: Initialize the privacy budget for generating domain privacy data.

[0074] Initialize the deep learning model parameters , and calculate the deep learning loss function including domain privacy data and not including domain privacy data as the initial contribution deviation of the domain privacy data. The initial contribution deviation of the nth group of domain privacy data is:

[0075] .

[0076] Among them:

[0077] represents the initial contribution deviation of the nth group of domain privacy data, represents the L1 norm, Denote the deep learning loss function that does not include the nth group of domain privacy data; specifically, the expression is:

[0078] .

[0079] Where:

[0080] denotes the set of domain privacy data that does not include the nth group of domain privacy data, x denotes any domain privacy data in the set of domain privacy data , and y denotes the predicted classification result of the domain privacy data x.

[0081] Based on the initial contribution deviation of the domain privacy data, initialize and generate the privacy budget for the domain privacy data. The privacy budget for the nth group of domain privacy data generated by initialization is:

[0082] .

[0083] Where:

[0084] denotes the privacy budget for the nth group of domain privacy data generated by initialization, denotes the total number of preset privacy budgets.

[0085] S3: Calculate the contribution gradient of the domain privacy data during the training process of the deep learning loss function, and dynamically and preferably adjust the privacy budget of the domain privacy data based on the contribution gradient.

[0086] The deep learning model parameters are trained iteratively during the training process of the deep learning loss function. The maximum number of training iterations of the deep learning model parameters is Max, and based on the change value of the deep learning loss function before and after the training iteration, generate the training gradient of the deep learning model parameters after each training iteration. Among them, the deep learning model parameters after the tth training iteration are , , the training gradient of the deep learning model parameters is . Based on the training gradient, calculate the contribution gradient of each group of domain privacy data during the training process of the deep learning loss function. The contribution gradient of the nth group of domain privacy data during the tth training iteration is:

[0087] .

[0088] Where:

[0089] denotes the contribution gradient of the nth group of domain privacy data during the tth training iteration, denotes the preset control coefficient, Represents the privacy budget dynamically and preferably adjusted after the (t - 1)-th training iteration for the n-th group of domain privacy data.

[0090] Specifically, the training gradient g t has the following calculation expression:

[0091] ;

[0092] .

[0093] Where:

[0094] represents the deep learning loss function with the deep learning model parameters as variables, represents the protection result of the domain privacy data based on the privacy budget .

[0095] Dynamically and preferably adjust the privacy budget for the domain privacy data based on the contribution gradient, where the dynamic and preferred adjustment formula for the privacy budget is:

[0096] .

[0097] Where:

[0098] represents the differential privacy change rate of the privacy budget , and g represents a preset contribution gradient threshold.

[0099] Specifically, the calculation formula for the differential privacy change rate is:

[0100] ;

[0101] .

[0102] Where:

[0103] represents the set of domain privacy data that does not include , , and u represents the predicted classification result of the i-th domain privacy data;

[0104] represents the deep learning loss function that does not include , represents the protection result of the n-th group of domain privacy data based on the privacy budget .

[0105] S4: Use the privacy budget after dynamic optimal adjustment to perform privacy protection on the domain privacy data, and train the deep learning loss function based on the domain privacy data after privacy protection until the deep learning loss function converges, obtaining the optimal deep learning model parameters after domain privacy protection, and constructing a deep learning model based on the optimal deep learning model parameters.

[0106] Perform privacy protection on the domain privacy data using the privacy budget after the dynamic optimal adjustment, and the privacy budget for the domain privacy data The formula for performing privacy protection is:

[0107] .

[0108] Where:

[0109] represents a random number in a normal distribution with a mean of 0 and a standard deviation of , and e represents the random noise intensity control parameter;

[0110] represents the protection result of the nth group of domain privacy data based on the privacy budget .

[0111] Train the deep learning loss function based on the domain privacy data after privacy protection to obtain the training iteration result of the deep learning model parameters until the training gradient of the deep learning model parameters is less than a preset gradient threshold. If the training gradient of the deep learning model parameters is less than the preset gradient threshold, then the deep learning loss function converges. The training iteration formula for the deep learning model parameters is:

[0112] ;

[0113] .

[0114] Where:

[0115] represents the learning rate of the (t + 1)-th training iteration, represents the deep learning model parameters after the (t + 1)-th training iteration.

[0116] It should be understood that the above embodiments are for illustrative purposes only and are not limited by this structure in the scope of the patent application.

[0117] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments. And the term "including", "comprising" or any other variant thereof in this article is intended to cover a non-exclusive inclusion, so that a process, apparatus, article or method including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, apparatus, article or method including the element.

[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0119] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for optimizing data differential privacy treatment for domain deep learning, characterized in that: The method comprises: S1: Obtain a domain privacy dataset, and construct the domain privacy dataset as a deep learning loss function, wherein the domain privacy dataset consists of domain privacy data and predicted classification results of the domain privacy data, and the domain privacy data includes privacy data in a variety of personal sensitive fields; The deep learning loss function uses the deep learning model parameters as variables and the domain privacy data set as input to train the deep learning model parameters until the deep learning loss function converges, thereby obtaining the optimal deep learning model parameters that enable the prediction and classification results of the deep learning model to be optimal; The deep learning model takes user data as input and the user's predicted classification result as output, wherein the user data is the user's data information in various personal sensitive areas; S2: Initialize the privacy budget for generating domain privacy data; S3: Calculate the contribution gradient of domain privacy data in the deep learning loss function training process, and dynamically optimize and adjust the privacy budget of domain privacy data based on the contribution gradient; The deep learning model parameters are trained iteratively during the deep learning loss function training process, the maximum training iteration of the deep learning model parameters is Max, and based on the change value of the deep learning loss function before and after the training iteration, the training gradient of the deep learning model parameters after each training iteration is generated, where the deep learning model parameters after the tth training iteration are θ t , t∈[0,Max], deep learning model parameters θ t The training gradient is g t , based on the training gradient, the contribution gradient of each group of domain privacy data in the deep learning loss function training process is calculated, The contribution gradient of the nth group of domain privacy data in the tth training iteration is: in: represents the contribution gradient of the nth group of domain privacy data in the tth training iteration process, ρ represents the preset control coefficient, represents the privacy budget of the nth group of domain privacy data dynamically optimized and adjusted after the t-1th training iteration; The privacy budget of the domain privacy data is dynamically and optimally adjusted based on the contribution gradient, wherein the privacy budget The dynamic optimization adjustment formula is: in: Representing Privacy Budget The differential privacy change rate of g represents the preset contribution gradient threshold; S4: Use the privacy budget adjusted by dynamic optimization to protect the privacy of domain privacy data, train the deep learning loss function based on the privacy-protected domain privacy data until the deep learning loss function converges, obtain the optimal deep learning model parameters after domain privacy protection, and build a deep learning model based on the optimal deep learning model parameters.

2. The optimal method for handling data differential privacy in domain deep learning as claimed in claim 1, characterized in that: The personal sensitive fields include the medical field, the financial field, the consumer field and the social communication field. The domain privacy data set includes N groups of domain privacy data, where N represents the number of domain privacy data in the domain privacy data set, and the representation form of the nth group of domain privacy data is: in: x n Represents the nth group of domain privacy data; Sequentially represents x n Privacy data in the medical, financial, consumer and social communication fields; The field privacy data x n The predicted classification result is y n .

3. The optimal method for handling data differential privacy in domain deep learning as claimed in claim 2, characterized in that: Based on the domain privacy dataset, the deep learning loss function expression constructed is L(θ): in: θ represents the deep learning model parameters, T represents transpose, and ||·||2 represents the L2 norm; θ T x n represents the domain privacy data x output by the deep learning model built based on the deep learning model parameter θ n The predicted classification results.

4. The optimal method for handling data differential privacy in domain deep learning as claimed in claim 3, characterized in that: Initialize and generate the deep learning model parameter θ0, calculate and obtain the deep learning loss function including domain privacy data and excluding domain privacy data as the initial contribution deviation of the domain privacy data, and the initial contribution deviation of the nth group of domain privacy data is: in: represents the initial contribution deviation of the nth group of domain privacy data, ||·||1 represents the L1 norm, and L(θ0;n) represents the deep learning loss function that does not include the nth group of domain privacy data; Based on the initial contribution deviation of the domain privacy data, the privacy budget of the domain privacy data is initialized and generated. The privacy budget of the nth group of domain privacy data generated by the initialization is: in: represents the privacy budget of the nth group of domain privacy data generated by initialization, and ε represents the total preset privacy budget.

5. The optimal method for handling data differential privacy in domain deep learning as claimed in claim 1, characterized in that: The privacy budget adjusted by the dynamic optimization is used to protect the privacy of the domain privacy data. For domain privacy data x n The formula for privacy protection is: in: The mean is 0 and the standard deviation is A random number in a normal distribution, e represents the random noise intensity control parameter; Based on the privacy budget The protection result of the privacy data of the nth group of domains.

6. The optimal method for handling data differential privacy in domain deep learning as claimed in claim 1, characterized in that: The deep learning loss function is trained based on the privacy-protected domain privacy data to obtain the training iteration results of the deep learning model parameters until the training gradient of the deep learning model parameters is less than the preset gradient threshold. If the training gradient of the deep learning model parameters is less than the preset gradient threshold, the deep learning loss function converges. The training iteration formula of the deep learning model parameters is: in: β t represents the learning rate of the t+1th training iteration, θ t+1 Represents the deep learning model parameters after the t+1th training iteration.

Citation Information

Patent Citations

  • Differential privacy protection deep learning algorithm for adaptively allocating dynamic privacy budget

    CN113642715A

  • Federal learning method for adaptive differential privacy, client, server, storage medium and product

    CN119494421A