A utility-optimized local differential privacy approach for estimating the frequency of two-dimensional data
By setting a privacy budget and perturbation mechanism in local differential privacy technology, distinguishing the sensitive and non-sensitive parts of the data for processing, the problem of insufficient sensitive data protection or excessive non-sensitive data protection in the prior art is solved, and finer-grained privacy protection and higher data frequency estimation accuracy are achieved.
Patent Information
- Application Number
- CN202210738857.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-06-28
AI Technical Summary
When existing local differential privacy technologies protect user privacy, they can easily lead to insufficient sensitive data protection or excessive non-sensitive data protection, affecting data utility.
A utility optimization local differential privacy method is proposed. By setting up a server-side privacy budget and perturbation mechanism, distinguishing the sensitive and non-sensitive parts of the data, and performing targeted perturbation processing to achieve the estimation of the frequency of two-dimensional data.
While ensuring that the degree of privacy protection of sensitive data is not reduced, the protection effect of non-sensitive data is improved, the finer-grained privacy protection is achieved, and the accuracy and application breadth of data frequency estimation are improved.
Smart Images

Figure CN115168893B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to information security technology and relates to a utility-optimized local differential privacy method for estimating two-dimensional data frequency. Background Art
[0002] Due to the rapid development of big data technology, the application scenarios of technologies such as machine learning and data mining have continued to expand, and at the same time, it has also brought about the issue of user data security. Users' personal information, home addresses, etc. may be reported to the collector through devices and various apps. Since the implementation of the "Data Security Law", the "Personal Information Protection Law of the People's Republic of China" was officially implemented on November 1, 2021, and my country's data security and personal information protection have entered a new era.
[0003] Privacy protection technology has gradually developed from the initial anonymization and secure multi-party computing to the current differential privacy. At present, differential privacy (DP) has become a standard for the release of private data. Due to the lack of trusted third parties in real life, local differential privacy (LDP) came into being. With the further refinement of the degree of privacy protection, utility-optimized local differential privacy (ULDP) and distinguishable input local differential privacy (ID-LDP) have gradually been proposed to further protect the privacy security of users.
[0004] LDP has been widely studied and used to estimate the distribution of private data. LDP does not require a trusted third party. It assumes that all personal data is equally sensitive, which actually leads to overprotection and has a certain impact on utility. For example, when asking users whether they have certain diseases, answers such as "AIDS" and "cancer" are more sensitive than "cold". At this time, if the perturbation mechanism of LDP is used to perturb the data, it is easy to cause insufficient protection of sensitive data or excessive protection of non-sensitive data. Therefore, in order to solve this problem, Murakami et al. proposed the concept of utility optimized local differential privacy (ULDP), which only provides ε-LDP privacy protection for sensitive data. ULDP is actually an optimization of LDP. It perturbs sensitive data and non-sensitive data separately, which is more suitable for situations with more non-sensitive data. It can ensure the security of sensitive data while solving the problem of LDP's actual data utility. If the data privacy protection levels are further divided, Gu et al. proposed the concept of Input-Discriminative Protection for Local Differential Privacy (ID-LDP), which allows users to customize the protection level of their privacy data and achieve more fine-grained protection. Summary of the invention
[0005] Purpose of the invention: The present invention provides a utility-optimized local differential privacy method for estimating two-dimensional data frequency, which extends the one-dimensional data frequency estimation under the existing ULDP to two dimensions for later wider applications.
[0006] In order to achieve the above-mentioned object of the invention, the technical solution provided by the present invention is as follows.
[0007] A utility-optimized local differential privacy method for estimating the frequency of two-dimensional data includes the following steps:
[0008] S1. Set the value of the server-side privacy budget;
[0009] The total privacy budget of the two-dimensional attribute is ε, and the privacy budget that needs to be satisfied separately for all sensitive parts is ε. x ;
[0010] S2. Each user reports a two-dimensional data, denoted as (x, y), where the input domain of x is X and the output domain is denoted as M; the input domain of y includes the sensitive part Y s and the non-sensitive part Y n , the output domain of y corresponds to N p and N I , and there are the following definitions:
[0011] definition For any (y0, x0) ≠ (y′0, x′0), the following inequality is satisfied:
[0012]
[0013] Where ε is the privacy budget. This definition ensures that (x, y) satisfies ε-LDP as a whole.
[0014] definition For any y0≠y′0, the following conditions are satisfied:
[0015] and
[0016] definition For any x≠x′, we have the following inequality:
[0017]
[0018] S3. Set the following disturbance mechanism, and the calculation process is as follows:
[0019] For (x, y)∈(X, Y S )part:
[0020] For (x, y)∈(X, Y N )part:
[0021]
[0022] in:
[0023]
[0024]
[0025]
[0026] S4. Report the disturbed data in step (S3) to the server, and the server estimates the frequency of the two-dimensional data based on the disturbance values reported by all users.
[0027] Further, step (S3) includes the following process:
[0028] Attribute x is considered fully sensitive, attribute y includes sensitive and non-sensitive parts, the input domain of attribute x is X, the output domain is M, and the sensitive input domain of y is Y S , the non-sensitive input domain is Y N , the protection output domain is N P , the reversible output domain is N I ;
[0029] Attribute x is considered to be fully sensitive, that is, it is fully perturbed. For the sensitive part of attribute y, it remains unchanged with a high probability, otherwise it is perturbed to any value in the protection output domain. For the insensitive part, it is perturbed to the protection output domain with a certain probability, otherwise it is perturbed to the reversible output domain; thus, the attacker cannot determine the initial value of the value in the protection output domain with a high confidence, and the reversible output domain can be reversed to the insensitive value. Among them, the summary of the perturbation of the sensitive part of attribute y follows the value range of probability a in step (S3), and the summary of the perturbation of the insensitive part of attribute y follows the value range of probability b in step (S3).
[0030] Furthermore, in step (S4), the server receives the disturbed value and estimates the frequency of the two-dimensional data as follows:
[0031] For the output domain in (M, N P ) part, set represents the true frequency of attribute values x=t, y=j, represents the empirical frequency under this value, Represents the estimated value of frequency, and the estimation formula is:
[0032]
[0033] The frequency of x=i, y=j is estimated by the above formula, and the estimated value of the frequency is It is calculated by dividing the frequency by the total number of users, that is,
[0034]
[0035] For the output domain in (M, N I ), the estimated calculation formula is as follows:
[0036]
[0037]
[0038] Beneficial effect: The present invention sets two different protection levels, sensitive and non-sensitive, for one-dimensional data on the server side, and gives the total privacy budget ε that the sensitive part of the two-dimensional data needs to meet, as well as the value of the privacy budget ε that the other dimensional data needs to meet separately. x ; The present invention ensures that in this model, the privacy protection level of fully sensitive data is not reduced and can still reach ε x -LDP, and the invention also provides overall protection for two-dimensional data, and the perturbation process of the non-sensitive part of one-dimensional data is still reversible. The present invention extends the existing one-dimensional data frequency estimation under ULDP to two dimensions, and its privacy data is more secure, the processing process and data calculation model and system are more stable, and the application is more extensive. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is the overall flow chart of the present invention;
[0040] Figure 2 Optimizing the local differential privacy model graph for the two-dimensional utility of the present invention;
[0041] Figure 3 Schematic diagram of disturbance in the present invention. DETAILED DESCRIPTION
[0042] In order to explain the technical solution disclosed by the present invention in detail, further elaboration is made below in conjunction with the accompanying drawings and specific embodiments of the specification.
[0043] First, combining the existing ULDP and ID-LDP applications, the data security over-protection and sensitive data under-protection issues, the present invention includes defining a two-dimensional ULDP model, and then designing a sensitive data perturbation mechanism based on the defined model. The method of the present invention is used to estimate the joint distribution of two-dimensional data, which can be applied to solving problems such as conditional probability and naive Bayes classification.
[0044] In the method of the present invention, first, each user reports the values of two attributes to form a two-dimensional data, and the utility optimization local differential privacy model of the two-dimensional data is defined. Then, the server regards one of the attributes as fully sensitive, and divides the other attribute into sensitive and non-sensitive parts, and sets a privacy budget value for the overall two-dimensional attribute, as well as a privacy budget value that needs to be satisfied by the fully sensitive attribute alone. The user perturbs his own two-dimensional data locally according to the value of the privacy budget provided by the server, and submits the perturbed noise value to the data collector. Finally, after the server collects the perturbed data of all users, it counts the data to estimate the frequency of the two-dimensional data.
[0045] Specifically, combined with Figure 1 , a utility-optimized local differential privacy method for estimating the frequency of two-dimensional data, the steps are as follows:
[0046] (1) The steps required by the user are as follows:
[0047] (1.1) Each user reports a two-dimensional data, denoted as (x, y), and the total number of users is denoted as N.
[0048] In order to simplify the solution steps and facilitate practical applications, usually, the output domains of x and y are considered equal to the input domains. Under the premise of considering only one attribute, the present invention regards x as fully sensitive and y as sensitive and non-sensitive. Further, let the input domain of attribute x be X = {x1, x2}, the output domain M = {x1, x2}, and the sensitive input domain of y be Y S ={y1, y2}, the non-sensitive input domain is Y N ={y3, y4}, the protected output domain is N P ={y1, y2}, the reversible output domain is N I ={y3,y4}.
[0049] (1.2) The present invention provides the following disturbance mechanism. Figure 2 and Figure 3 , the specific disturbance process and calculation can be expressed as follows:
[0050] For (x, y)∈(X, Y S ) part is as follows:
[0051]
[0052] For (x, y)∈(X, Y N ) part is as follows:
[0053]
[0054] in:
[0055]
[0056]
[0057]
[0058] Furthermore, attribute x is considered to be fully sensitive and is perturbed. For the sensitive part of attribute y, it remains unchanged with a high probability, otherwise it is perturbed to any value in the protected output domain. For the non-sensitive part, it is perturbed to the protected output domain with a certain probability, otherwise it is perturbed to the reversible output domain. In this way, the attacker cannot determine the initial value of the value in the protected output domain with a high degree of confidence, and the reversible output domain can be reversed to the non-sensitive value.
[0059] (1.3) Finally, we get a perturbed noise value Report it to the server.
[0060] (2) For the server side, the steps required are as follows:
[0061] (2.1) The server first sets a privacy budget value ε for two-dimensional data and a privacy budget value ε that all sensitive attributes need to satisfy. x . And divide the data into sensitive parts and non-sensitive parts.
[0062] (2.2) After the server collects all the data reported by users, it performs statistical analysis on the data.
[0063] For the output domain in (M, N P ) part, set represents the true frequency of attribute values x=i, y=j, represents the empirical frequency under this value, It represents the estimated value of frequency, and according to the disturbance process, it has the following formula:
[0064]
[0065]
[0066] Therefore, the estimated frequency of x=i, y=j can be estimated, and the estimated value of the joint distribution can also be calculated by dividing by the total number of users, as shown below:
[0067]
[0068] For the output domain in (M, N I ) part, the following non-homogeneous linear equations can be established:
[0069]
[0070] In order to solve this non-homogeneous linear equation system, the augmented matrix is subjected to elementary row transformation,
[0071]
[0072] This linear system of equations has a unique solution. Therefore, the estimation method of the joint distribution can be summarized as follows:
[0073]
[0074]
[0075] Based on the above calculation process, the following are the experimental results of the present invention.
[0076] The experiment uses the Nursery dataset from UCI, which contains 12,960 instances, 8 features, and a value domain size of 5 for the class label. This embodiment uses the second feature and class label for the experiment, and the data scale is 5*5. The total privacy budget values set in the experiment are 0.5, 1.5, 2, 3, 4, 5, and 6, respectively, and the privacy budget of one attribute is ε, 0.5, 2, and 5, respectively. Since this dataset is an evaluation of nurseries, the present invention selects the first two categories of class labels as sensitive values and the remaining three categories as non-sensitive values. The second feature is considered to be fully sensitive. The results use Euclidean distance as the measurement standard, and the final two-dimensional frequency error between each feature and the class label is shown in Table 1.
[0077] According to the experimental results, as the privacy budget continues to increase, the protection level continues to decrease, and the error of the present invention continues to decrease, and gradually approaches 0. When the privacy budget takes a minimum value of 0.5, the error of the present invention does not exceed 0.4, and the accuracy can be guaranteed.
[0078] Table 1 Error between estimated frequency and true frequency of two-dimensional data
[0079]
[0080] Finally, it should be noted that the existing inventions have only considered the utility optimization of local differential privacy under one-dimensional attributes. The present invention proposes a utility optimization local differential privacy model under two-dimensional data. And provides the corresponding model definition. Further, based on the given model definition, a utility optimization local differential privacy perturbation mechanism for two-dimensional data frequency estimation is designed, including providing a corresponding estimation formula. Experiments and data show that the method described in the present invention can achieve more fine-grained protection, and reasonably protect sensitive data on the basis of ensuring the security of privacy data, thereby improving the utility of estimating privacy data.
[0081] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A utility-optimized local differential privacy method for estimating the frequency of two-dimensional data, characterized in that: The following steps are involved: S1. Set the value of the server-side privacy budget; The total privacy budget of the two-dimensional attribute is ε, and the privacy budget that needs to be satisfied separately for all sensitive parts is ε. x ; S2. Each user reports a two-dimensional data, denoted as (x, y), where the input domain of x is X and the output domain is denoted as M; the input domain of y includes the sensitive part Y s and the non-sensitive part Y N , the output domain of y corresponds to N p and N I , and there are the following definitions: definition For any (y0,x0)≠(y'0,x'0), the following inequality is satisfied: Where ε is the privacy budget. This definition ensures that (x, y) satisfies ε-LDP as a whole. definition For any y0≠y'0, the following conditions are met: and definition For any x≠x', we have the following inequality: S3. Set the following disturbance mechanism, including the following process: Consider attribute x as fully sensitive, attribute y includes two parts, sensitive and non-sensitive, the input domain of attribute x is X, the output domain is M, and the sensitive input domain of y is Y S , the non-sensitive input domain is Y N , the protection output domain is N P , the reversible output domain is N I ; Attribute x is considered to be fully sensitive and perturbed. For the sensitive part of attribute y, it remains unchanged with probability a, otherwise it is perturbed to any value in the protection output domain. For the insensitive part, it is perturbed to the protection output domain with probability b, otherwise it is perturbed to the reversible output domain. For (x,y)∈(X,Y S )part: For (x,y)∈(X,Y N )part: in: S4. Report the disturbed data in step S3 to the server, and the server estimates the frequency of the two-dimensional data according to the disturbance values reported by all users.
2. The utility-optimized local differential privacy method for estimating two-dimensional data frequency according to claim 1, characterized in that: In step S4, the server receives the disturbed value and estimates the frequency of the two-dimensional data as follows: For the output domain in (M,N P ) part, set represents the true frequency of attribute values x=i, y=j, represents the empirical frequency under this value, Represents the estimated value of frequency, and the estimation formula is: The frequency of x=i, y=j is estimated by the above formula, and the estimated value of the frequency is It is calculated by dividing the frequency by the total number of users, that is, For the output domain in (M,N I ), the estimated calculation formula is as follows:
Citation Information
Patent Citations
Differential privacy two-dimensional spatial data publishing method based on step mechanism
CN111723168A
Data collection method based on personalized local differential privacy
CN113297621A