Personalized differential privacy protection method, device and system based on small sample data
By grouping users and determining the best disturbance method, combined with weighted aggregation, the problem of personalized privacy protection needs under small sample data is solved, and statistical accuracy and user participation are improved.
Patent Information
- Application Number
- CN202211334103.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-28
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2042-10-28
AI Technical Summary
Existing differential privacy protection technology based on random responses is difficult to achieve personalized privacy needs under small sample data, resulting in low statistical accuracy and insufficient user participation.
Based on the expectation criterion of minimum mean square error, by grouping users, the optimal perturbation method for each group is determined, and the data aggregation is adopted in a weighted aggregation method, considering the user's personalized privacy needs.
It improves the accuracy of statistical estimation under small sample data, enhances users' enthusiasm for participating in data collection and sharing, and realizes personalized privacy protection.
Smart Images

Figure CN115630398B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of privacy data protection technology, and in particular to a personalized differential privacy protection method and device based on small sample data. Background Art
[0002] Randomized Response (RR) is a mainstream perturbation mechanism in localized differential privacy (LDP) based on data distortion. Its model is simple, intuitive, and easy to implement, and its perturbation level can be directly quantified. It has excellent performance in estimating statistical properties, thus attracting widespread attention. RR protects the privacy of data providers (or respondents) by using a probabilistic approach to answer questions, ensuring strong deniability for sensitive questions and privacy protection. It has been implemented in Google Chrome's privacy protection tools and Apple's system. RR also fully considers the possibility of data collectors stealing or leaking user privacy during data collection. In this model, respondents can independently perform privacy processing on their individual data, preventing even data collectors from obtaining the exact original private data, greatly motivating them to participate in data collection. Therefore, unlike centralized privacy protection mechanisms that assume a trusted third party, the RR-based localized differential privacy protection mechanism eliminates the need for a trusted third party and eliminates the potential for privacy leaks and attacks from untrusted third-party data collectors.
[0003] However, the random response-based LDP technique perturbs individual data points both positively and negatively, ultimately aggregating a large number of perturbations to offset the added positive and negative noise, resulting in valid statistical results. Due to the random nature of the noise, ensuring unbiased statistical results requires massive datasets to achieve statistical accuracy that meets data availability.
[0004] On the other hand, in reality, different individuals have different privacy protection needs. If all users' data is rigidly protected at the same level, users with high privacy needs will receive insufficient protection, while users with low privacy needs will receive excessive protection. This will not only cause users to oppose the openness and sharing of data, but also reduce the accuracy of statistical estimates to a certain extent. Summary of the Invention
[0005] The purpose of this invention is to provide a local personalized privacy-preserving data aggregation method suitable for small sample data based on random responses. This method can, to a certain extent, address the problem of low statistical accuracy caused by small sample data, fully consider the personalized privacy needs of local end users, and provide an optimal perturbation method (perturbation probability, transmission parameter quantity, and perturbation output form) based on the minimum mean square error expectation criterion. Furthermore, appropriate weighting factors are constructed based on the minimum mean square error expectation to perform weighted aggregation, thereby improving the accuracy of statistical estimation.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] The first aspect provides a personalized differential privacy protection method based on small sample data, including:
[0008] The data aggregator receives the privacy protection level sent by the user;
[0009] Divide users into groups based on their privacy protection levels, and group users with the same privacy protection level into the same group;
[0010] Based on the user's privacy protection level and the minimum mean square error (MMSE) expectation criterion, the optimal perturbation method for the user's private data in each group is determined. This optimal perturbation method is sent to the corresponding users in the group. Each user in the group uses the corresponding optimal perturbation method to perturb their private data, obtaining perturbed data and sending it to the data aggregator. The optimal perturbation method includes perturbation probability, transmission parameter quantity, and perturbation output form.
[0011] A weighted aggregation method is used to aggregate the perturbation data from different groups at different privacy protection levels to obtain a statistical estimate of the user's private data.
[0012] In one embodiment, based on the user's privacy protection level and the desired criterion of minimum mean square error, determining the optimal perturbation method for the user's private data in each group includes:
[0013] Based on maximum likelihood estimation, the probability distribution of private data is estimated, and the expectation of the minimum mean square error of the estimated distribution is obtained;
[0014] Based on the expectation of minimum mean square error, whether the perturbed data contains real data, and the probability that the perturbed data contains a specific value, the relationship between the privacy protection level, the perturbation probability, and the transmission parameter quantity is obtained according to the definition of differential privacy.
[0015] Based on the expectation of the minimum mean square error and the privacy protection level of each group, an objective function is constructed and optimized to determine the optimal perturbation method for user privacy data in each group.
[0016] In one embodiment, estimating the probability distribution of private data based on maximum likelihood estimation and obtaining the expectation of the minimum mean square error of the estimated distribution includes:
[0017] Based on the maximum likelihood estimation, the probability of private data is estimated, where the i-th private data x i The true probability Estimated value of The calculation formula is:
[0018]
[0019] in, N τ The disturbed data contains x i The number of data items, G τ is the τth group, N τ For group G τ The number of users in the corresponding τ perturbation data, real data x=x i When the perturbation data contains x i The probability is p τ , that is, the perturbation probability. On the contrary, the perturbation data contains x j The probability is q τ , x j with x i For different data, s τ is the number of data contained in a disturbance data, i.e., the transmission parameter;
[0020] It should be noted that the perturbation data is in units of pieces. After a type of private data is perturbed, a piece of perturbation data is obtained. This perturbation data can contain only one type of data or multiple types of data, depending on the transmission parameter quantity and the perturbation output form.
[0021] Based on the estimated value of the true probability of each type of private data, we can obtain the estimated value of the true probability distribution of all private data. We can further obtain the expectation of the minimum mean square error of the estimated distribution, which is specifically:
[0022]
[0023] k is the value of the private data.
[0024] In one embodiment, based on the expectation of minimum mean square error, whether the perturbed data contains real data, and the probability of the perturbed data containing a specific value, the relationship between the privacy protection level, the perturbation probability, and the transmission parameter is obtained according to the definition of differential privacy, including:
[0025] According to whether the disturbance data contains real data, it is obtained that if the disturbance data contains real data, it contains specific s τ The probability P of α1 privacy values u,x , when the perturbation data does not contain the real data, the perturbation data contains specific s τ The probability of a value
[0026]
[0027]
[0028] Among them, the real data x=x i When the perturbation data contains x i The probability is p τ , that is, the perturbation probability, s τ is the number of data contained in a piece of disturbed data, and k is the value of the private data;
[0029] According to the expectation of minimum mean square error, whether the perturbed data contains real data, and the probability that the perturbed data contains a specific value, the relationship between the privacy protection level, the perturbation probability, and the transmission parameter is obtained according to the definition of differential privacy:
[0030]
[0031] Among them, |·| is the absolute value operation, ∈ τ is the privacy protection level of the τth group.
[0032] In one embodiment, an objective function is constructed based on the expectation of the minimum mean square error and the privacy protection level of each group, and is implemented in the following manner:
[0033]
[0034]
[0035] 0<p τ <1
[0036] s τ ∈{1, 2, ...., k}
[0037] Then the optimal transmission parameter s τ and the perturbation probability p τ satisfy:
[0038]
[0039]
[0040] in, It is a floor operation.
[0041] In one embodiment, a weighted aggregation method is used to aggregate disturbance data from different groups at different privacy protection levels to obtain a statistical estimate of the user's private data, including:
[0042] Construct weighting factors based on the expectation of minimum mean square error:
[0043]
[0044] Among them, w τ is the weighting factor of the τth group, l is the counting symbol, ranging from 1 to m, E[MSE] τ is the expectation of the minimum mean square error of the τth group;
[0045] Based on the weighting factor, the estimated values of the true probability of each type of private data in the m groups are weighted and aggregated to obtain the statistical estimate of the user's private data. Then the i-th type of private data x i The estimated probability of:
[0046]
[0047] in, The user's private data x i The statistical estimate of is the estimated value of the true probability of the private data xi of the τth group user.
[0048] Based on the same inventive concept, the second aspect of the present invention provides a personalized differential privacy protection device based on small sample data, wherein the device is a data aggregator, comprising:
[0049] A privacy protection level receiving module is used to receive the privacy protection level sent by the user;
[0050] The group division module is used to divide users into groups according to their privacy protection levels, and divide users with the same privacy protection level into the same group;
[0051] The perturbation method determination module is used to determine the optimal perturbation method for the private data of users in each group based on the user's privacy protection level and the expected minimum mean square error criterion. The optimal perturbation method is sent to the corresponding users in the group. Each user in the group uses the corresponding optimal perturbation method to perform privacy protection operations on their respective private data, obtain perturbed data, and send it to the data aggregator. The optimal perturbation method includes the perturbation probability, the amount of transmission parameters, and the perturbation output form.
[0052] The weighted aggregation module is used to aggregate the disturbance data from different groups at different privacy protection levels using a weighted aggregation method to obtain a statistical estimate of the user's private data.
[0053] Based on the same inventive concept, the third aspect of the present invention provides a personalized differential privacy protection system based on small sample data, including the personalized differential privacy protection device based on small sample data as described in the second aspect and a user terminal, wherein the user terminal is used to send the privacy protection level to the data aggregator, perform perturbation processing on the corresponding private data according to the optimal perturbation method sent by the data aggregator, obtain perturbation data, and send it to the data aggregator.
[0054] Based on the same inventive concept, the fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the method described in the first aspect when the program is executed.
[0055] Based on the same inventive concept, the fifth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.
[0056] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows:
[0057] The present invention proposes a personalized differential privacy protection method based on small sample data, which divides users into groups according to their privacy protection levels, and divides users with the same privacy protection level into the same group; and based on the user's privacy protection level and the expected criterion of minimum mean square error, determines the optimal perturbation method for the user's privacy data in each group, fully considering the personalized privacy needs of local users, while achieving personalized privacy protection, and to a certain extent improving the enthusiasm and initiative of local users in participating in data collection and sharing; based on the expected criterion of minimum mean square error, derives and gives the optimal perturbation method under different privacy protection levels, which can improve the problem of low statistical estimation accuracy under small sample data; further, adopts a weighted aggregation method to aggregate the perturbation data from different groups under different privacy protection levels, and can further improve the accuracy of privacy distribution estimation compared with direct aggregation based on the weighted aggregation of the expected minimum mean square error;
[0058] The perturbation output for subgroups with different privacy requirements is no longer a one-to-one perturbation output, but a one-to-many perturbation output, which is equivalent to increasing the sample size and thus obtaining more perturbation data. When sample data is scarce, the personalized random response mechanism described in this invention can achieve higher usability statistical accuracy due to the increased amount of perturbation data. This is a practical method for statistics and analysis with strong practical significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0060] Figure 1 This is a flow chart of a personalized differential privacy protection method based on small sample data provided by an embodiment of the present invention;
[0061] Figure 2 This is an overall framework diagram of the personalized differential privacy protection method provided by an embodiment of the present invention;
[0062] Figure 3 This is a structural block diagram of a personalized differential privacy protection device based on small sample data provided by an embodiment of the present invention;
[0063] Figure 4 A schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present invention;
[0064] Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The inventor's previous patent, a personalized privacy protection method and device based on random responses, utilizes a personalized random response model. The probability of outputting true values for highly weighted sensitive values is lower than the probability of outputting true values for less weighted sensitive values. This approach can, to a certain extent, address the problem of existing random response-based technologies overprotecting some private data and underprotecting others during privacy collection. However, this solution employs a one-to-one perturbation output method. While it improves estimation accuracy to a certain extent, it cannot address the problem of low estimation accuracy with small sample sizes (Issue 1). Furthermore, this solution provides personalized privacy protection for private data, determining the level of privacy protection based on the sensitivity of the private data. This approach protects privacy based on the characteristics of the private data itself, not the user's personalized needs. Therefore, it is not a privacy protection solution specifically tailored to individual privacy needs. In reality, privacy needs vary from person to person, and even for the same private data, different individuals may have different privacy needs. Therefore, addressing the privacy protection issue (Issue 2) under the personalized privacy needs of individuals is of great practical significance.
[0066] Therefore, based on the above discussion, the present invention targets small sample data and considers the individual's personalized privacy needs. It proposes a novel personalized random response mechanism based on the expectation of minimum mean square error, and constructs a suitable weighting factor based on the expectation of minimum mean square error, and then adopts weighted aggregation to obtain a better statistical estimate. Specifically, for problem 1, the technical solution of the present invention is no longer a one-to-one perturbation output for subgroups (groups) with different privacy needs, but a one-to-one and one-to-many perturbation output. This increases the sample size and can obtain more perturbation data, thereby solving the problem of low estimation accuracy under a small sample size. For problem 2, compared with the personalized privacy protection method based on random response that protects according to the sensitivity of privacy data, the present invention fully considers the personalized privacy needs of local users, groups users according to the privacy protection level, and divides users with the same privacy protection level into the same group. Based on the user's privacy protection level and the expected criterion of minimum mean square error, the optimal perturbation method for the user's privacy data in each group is determined, thereby achieving truly personalized privacy protection. In addition, it also improves the enthusiasm and initiative of local users in participating in data collection and sharing to a certain extent.
[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0068] Example 1
[0069] The embodiment of the present invention provides a personalized differential privacy protection method based on small sample data, including:
[0070] The data aggregator receives the privacy protection level sent by the user;
[0071] Divide users into groups based on their privacy protection levels, and group users with the same privacy protection level into the same group;
[0072] Based on the user's privacy protection level and the minimum mean square error (MMSE) expectation criterion, the optimal perturbation method for the user's private data in each group is determined. This optimal perturbation method is sent to the corresponding users in the group. Each user in the group uses the corresponding optimal perturbation method to perturb their private data, obtaining perturbed data and sending it to the data aggregator. The optimal perturbation method includes perturbation probability, transmission parameter quantity, and perturbation output form.
[0073] A weighted aggregation method is used to aggregate the perturbation data from different groups at different privacy protection levels to obtain a statistical estimate of the user's private data.
[0074] See Figure 1 , is a flowchart of a personalized differential privacy protection method based on small sample data provided by an embodiment of the present invention.
[0075] Specifically, the user's privacy protection level can be determined according to the user's own needs. After the data aggregator receives the privacy protection level sent by the user, it divides the user into different subgroups, i.e., groups, according to the privacy protection level. Users with the same privacy protection level are divided into the same group.
[0076] For each group, the data aggregator (or data collector) determines the optimal perturbation method (perturbation probability, transmission parameter quantity, and perturbation output form) for the private data of each subgroup (group) based on the expected minimum mean square error criterion. This is the optimal perturbation method at the corresponding privacy protection level for the group. The local users in each subgroup use the optimal perturbation method to perform privacy protection operations (i.e., perturbation processing) on their private data and send the perturbed data to the data aggregator.
[0077] The data aggregator aggregates the disturbed data from subgroups (groups) at different privacy protection levels, constructs appropriate weighting factors based on the expectation of minimum mean square error, and then uses weighted aggregation to obtain better statistical estimates.
[0078] This scheme is a localized data collection method based on random responses. Data providers (local users) can perform personalized privacy protection based on their privacy needs. Privacy protection needs are measured by the differential privacy parameter ∈. Users are divided into different subgroups based on their individual privacy needs. Users in the same subgroup have the same privacy needs, while users in different subgroups have different privacy needs. Based on the expected minimum mean square error (MMSE) criterion, the optimal perturbation method is determined for each subgroup, including the optimal perturbation probability, optimal transmission parameter quantity, and optimal output form. Appropriate weighting factors are constructed based on the expected minimum mean square error (MMSE) for weighted aggregation to improve the accuracy of statistical estimation. This ensures high statistical accuracy while achieving personalized privacy protection.
[0079] In one embodiment, based on the user's privacy protection level and the desired criterion of minimum mean square error, determining the optimal perturbation method for the user's private data in each group includes:
[0080] Based on maximum likelihood estimation, the probability distribution of private data is estimated, and the expectation of the minimum mean square error of the estimated distribution is obtained;
[0081] Based on the expectation of minimum mean square error, whether the perturbed data contains real data, and the probability that the perturbed data contains a specific value, the relationship between the privacy protection level, the perturbation probability, and the transmission parameter quantity is obtained according to the definition of differential privacy.
[0082] Based on the expectation of the minimum mean square error and the privacy protection level of each group, an objective function is constructed and optimized to determine the optimal perturbation method for user privacy data in each group.
[0083] Estimated distribution refers to the estimated probability distribution of private data.
[0084] In one embodiment, estimating the probability of private data based on maximum likelihood estimation and obtaining the expectation of the minimum mean square error of the estimated distribution includes:
[0085] Based on the maximum likelihood estimation, the probability of private data is estimated, where the true probability of the i-th private data xi is Estimated value of The calculation formula is:
[0086]
[0087] in, There are Nτ pieces of perturbation data containing x i The number of data items, Gτ is the τth group, Nτ is the number of users in group Gτ, corresponding to N τ perturbation data, real data x=x i When the perturbation data contains x i The probability is p τ , that is, the perturbation probability. On the contrary, the perturbation data contains x j The probability is q τ , x j with x i For different data, s τ is the number of data contained in a disturbance data, i.e., the transmission parameter;
[0088] It should be noted that the perturbation data is in units of pieces. After a type of private data is perturbed, a piece of perturbation data is obtained. This perturbation data can contain only one type of data or multiple types of data, depending on the transmission parameter quantity and the perturbation output form.
[0089] Based on the estimated value of the true probability of each type of private data, we can obtain the estimated value of the true probability distribution of all private data. We can further obtain the expectation of the minimum mean square error of the estimated distribution, which is specifically:
[0090]
[0091] k is the value of the private data.
[0092] This estimated distribution refers to the previously estimated group G τ An estimate of the true probability distribution of all private data contained in .
[0093] In the specific implementation process, Figure 1 As shown in the figure: Assume that the local user has m personalized privacy protection levels, namely ∈1, ∈2, ..., ∈ m , which can be divided into m subgroups G1, G2, ..., G m , the number of users in the subgroups (groups) is N1, N2, ..., N m The total sample size is Where m ≥ 2 and is an integer.
[0094] Without loss of generality, assume that each user has only one kind of private data x∈X={x1,x2,...,x k}, where k is the number of different private data, x i is the i-th type of private data. Let private data x i The true probability is P i, after the privacy protection operation, the estimated probability is The mean square error (MSE) is defined to measure the accuracy of statistical estimation:
[0095]
[0096] in, is the estimated value of the private data distribution, P x =[P1, P2, ..., P k ] is the true distribution of private data, E represents the expected operation, is the two-norm operation.
[0097] Estimation of privacy distribution for subgroups (groups): For privacy protection level ∈ τ The true distribution of private data of the group Gτ is defined as Then its estimated distribution is defined as in For group G τ Privacy data in China i The true probability The estimated value of , where τ = 1, 2, ..., m, i = 1, 2, ..., k. For the group G τ In the private data x, the disturbance probability p is introduced τ and the transmission parameter s τ , the perturbation method is described as: Assume that the probability of a non-uniform coin facing up is p τ , there are two cases where disturbance occurs:
[0098] (1) If the front side is facing up, two steps are performed: first, the original value x is retained; second, the value x is removed from the data set {x1, x2, ..., x k}\x randomly select s τ -1 private data, forming a line containing s τ The perturbation data of each data.
[0099] (2) If the back side is facing up, remove x from the data set {x1, x2, ..., x k}\x randomly select s τ private data, forming a line containing s τ The perturbation data of each data.
[0100] For example, the private data X is the patient's condition, and there are k=5 different values, namely X={x1, x2, ..., x k} = {lung cancer, liver cancer, heart disease, coronary heart disease, chronic gastritis}. Assume that the transmission parameter s τ =3, disturbance probability pτ = 0.6, the real private data of user u is x = "liver cancer". At this time, the data set after removing x = "liver cancer" is {x1, x2, ..., x k}\x = {lung cancer, heart disease, coronary heart disease, chronic gastritis}. A coin is tossed, and the probability of it landing heads is 0.6. If heads, keep {"liver cancer"}. Then randomly select two values from the set {lung cancer, heart disease, coronary heart disease, chronic gastritis}, assuming they are {lung cancer, coronary heart disease}. Combine "liver cancer" with the perturbed data {lung cancer, lung cancer, coronary heart disease} and output it, including u's true private data "liver cancer". If tails, randomly select three values from the set {lung cancer, heart disease, coronary heart disease, chronic gastritis} and output it as perturbed data. Assume the output is {lung cancer, coronary heart disease, chronic gastritis}, excluding u's true private data "liver cancer".
[0101] For Group G τ , the number of users is N τ , using the above perturbation method, we will get Nτ pieces of perturbation data, and N τ =|G τ |. If the real data of user u is x=x i , then the perturbation data contains x i The probability is p τ On the contrary, the perturbation data contains x j The probability q of (i≠j) τ for
[0102]
[0103] in, is the symbol for permutations and combinations, ! is the symbol for finding the hierarchy, n! = n × (n-1) × (n-2) × … × 2 × 1.
[0104] Based on maximum likelihood estimation, we can get the i-th privacy data x i The true probability The estimated value of . And further, the probability of all private data can be calculated, thus obtaining the expected expression of the minimum mean square error.
[0105] In one embodiment, based on the expectation of minimum mean square error, whether the perturbed data contains real data, and the probability of the perturbed data containing a specific value, the relationship between the privacy protection level, the perturbation probability, and the transmission parameter is obtained according to the definition of differential privacy, including:
[0106] According to whether the disturbance data contains real data, it is obtained that if the disturbance data contains real data, it contains specific s τ -Probability P of 1 privacy value u,x, when the perturbation data does not contain the real data, the perturbation data contains specific s τ The probability of a value
[0107]
[0108]
[0109] Among them, the real data x=x i When the perturbation data contains x i The probability is p τ , that is, the perturbation probability, s τ is the number of data contained in a piece of disturbed data, and k is the value of the private data;
[0110] It should be noted that the perturbation data is in units of pieces. After a type of private data is perturbed, a piece of perturbation data is obtained. This perturbation data can contain only one type of data or multiple types of data, depending on the transmission parameter quantity and the perturbation output form.
[0111] According to the expectation of minimum mean square error, whether the perturbed data contains real data, and the probability that the perturbed data contains a specific value, the relationship between the privacy protection level, the perturbation probability, and the transmission parameter is obtained according to the definition of differential privacy:
[0112]
[0113] Among them, |·| is the absolute value operation, ∈ τ is the privacy protection level of the τth group.
[0114] In one embodiment, an objective function is constructed based on the expectation of the minimum mean square error and the privacy protection level of each group, and is implemented as follows:
[0115]
[0116]
[0117] 0<p τ <1
[0118] s τ ∈{1, 2, ...., k}
[0119] Then the optimal transmission parameter s τ and the perturbation probability p τ satisfy:
[0120]
[0121]
[0122] in, It is a floor operation.
[0123] In one embodiment, a weighted aggregation method is used to aggregate disturbance data from different groups at different privacy protection levels to obtain a statistical estimate of the user's private data, including:
[0124] Construct weighting factors based on the expectation of minimum mean square error:
[0125]
[0126] Among them, w τ is the weighting factor of the τth group, l is the counting symbol, ranging from 1 to m, ElMSE] τ is the expectation of the minimum mean square error of the τth group;
[0127] Based on the weighting factor, the estimated values of the true probability of each type of private data in the m groups are weighted and aggregated to obtain the statistical estimate of the user's private data. Then the i-th type of private data x i The estimated probability of:
[0128]
[0129] in, The user's private data x i The statistical estimate of is the privacy data x of the τth group user i An estimate of the true probability of .
[0130] Because each group has different privacy requirements (privacy protection levels) and adopts different perturbation methods, that is, the amount of perturbation data submitted is different, each subgroup contributes differently to the accuracy of estimating a certain private data in the entire group. The more accurate the estimate in the subgroup (group), the greater its contribution, the larger the response weighting factor, and the better the weighted accuracy.
[0131] Specifically, for m subgroups, the private data x i The estimated distribution of : Use Figure 2 The weighted aggregation shown is used to obtain the final private data x i Statistical estimates of
[0132] It should be noted that, in practice, a user's private data can be one-dimensional (the example above is one-dimensional data) or multi-dimensional, such as {gender; marital status; medical condition} (i.e., three dimensions), where "gender" = {male, female}; "marital status" = {married, unmarried}; "medical condition" = {lung cancer, liver cancer, heart disease, coronary heart disease, chronic gastritis}, or provide the collection and sharing of multi-dimensional data for other situations.
[0133] During the specific implementation process, the solution can be expanded by concatenating multidimensional data with equivalent single-dimensional data according to the specific application scenario to complete the processing of multidimensional data, or each single-dimensional data can be processed separately.
[0134] For example, define the privacy data multidimensional attributes X1, X2, ..., X v , each attribute X i There is k i Different values.
[0135] The simplest way to handle this is: for each attribute X i K i The seed value adopts the personalized random response perturbation mechanism proposed in this application, that is, each attribute is individually perturbed using the optimal perturbation parameters to meet the needs of statistical analysis of the privacy distribution of each attribute.
[0136] In another processing method, v attributes can be combined into X1×X2×…×X v , and then construct a new joint attribute, the number of its different values is k = k1×k2×…×k v Then, the personalized random response mechanism proposed in this application is adopted, that is, v attributes are cascaded into one attribute for processing to meet the needs of mining association rules between attributes. In actual production, the selection can be made according to the specific application scenario.
[0137] The advantages and beneficial technical effects of the present invention include:
[0138] (1) Fully consider the personalized privacy needs of local users, and while achieving personalized privacy protection, to a certain extent improve the enthusiasm and initiative of local users in participating in data collection and sharing;
[0139] (2) Based on the minimum mean square error expectation criterion, the optimal perturbation method under different privacy protection levels is derived and given, which is expected to improve the problem of low statistical estimation accuracy under small sample data;
[0140] (3) Weighted aggregation based on the minimum mean square error expectation further improves the accuracy of privacy distribution estimation compared to direct aggregation;
[0141] (4) The perturbation output of subgroups with different privacy requirements is no longer a one-to-one perturbation output, but a one-to-many perturbation output, which is equivalent to increasing the sample size, so more perturbation data can be obtained;
[0142] (5) When sample data is scarce, the personalized random response mechanism described in the present invention can obtain statistical accuracy with higher availability because the amount of disturbance data increases. It is a practical method for statistics and analysis and has strong practical significance.
[0143] Example 2
[0144] Based on the same inventive concept, this embodiment provides a personalized differential privacy protection device based on small sample data. The device is a data aggregator. Figure 3 ,include:
[0145] The privacy protection level receiving module 201 is used to receive the privacy protection level sent by the user;
[0146] A grouping module 202 is configured to group users according to their privacy protection levels, and group users with the same privacy protection level into the same group;
[0147] The perturbation mode determination module 203 is used to determine the optimal perturbation mode for the private data of users in each group based on the user's privacy protection level and the desired minimum mean square error criterion, and send the optimal perturbation mode to the corresponding users in the group. Each user in the group uses the corresponding optimal perturbation mode to perform privacy protection operations on their respective private data, obtain perturbed data, and send it to the data aggregator. The optimal perturbation mode includes perturbation probability, transmission parameter amount, and perturbation output form.
[0148] The weighted aggregation module 204 is configured to aggregate the disturbance data from different groups at different privacy protection levels using a weighted aggregation method to obtain a statistical estimate of the user's private data.
[0149] Since the device described in Example 2 of the present invention is used to implement the personalized differential privacy protection method based on small sample data in Example 1 of the present invention, the specific structure and variations of the device are well understood by those skilled in the art based on the method described in Example 1 of the present invention, and therefore will not be described in detail here. All devices used in the method of Example 1 of the present invention fall within the scope of protection of the present invention.
[0150] Example 3
[0151] Based on the same inventive concept, the present invention provides a personalized differential privacy protection system based on small sample data, including the personalized differential privacy protection device based on small sample data in Example 2 and a user terminal, wherein the user terminal is used to send the privacy protection level to the data aggregator, perform perturbation processing on the corresponding private data according to the optimal perturbation method sent by the data aggregator, obtain perturbation data, and send it to the data aggregator.
[0152] See Figure 2 , is the implementation block diagram of local end users and data aggregators, where each group corresponds to a privacy protection level. For this group, the data it has is the group data under the corresponding privacy protection level. Different groups correspond to different optimal perturbation methods (optimal perturbation probability and transmission parameter amount). Users in each group perform perturbation processing on their data to obtain the data that satisfies ∈ τ The perturbation data of the LDP is then aggregated by the data aggregator, specifically, based on the statistical estimate and weighting coefficient (weighting factor) of each group to obtain the final statistical estimate.
[0153] Since the system described in Example 3 of the present invention is the system used to implement the personalized differential privacy protection method based on small sample data in Example 1 of the present invention, the specific structure and variations of this system are well understood by those skilled in the art based on the method described in Example 1 of the present invention, and therefore will not be described in detail here. All systems used in the method of Example 1 of the present invention fall within the scope of protection of the present invention.
[0154] Example 4
[0155] Based on the same inventive concept, see Figure 4 The present invention further provides a computer-readable storage medium 300 on which a computer program 311 is stored. When the program is executed, the method described in the first embodiment is implemented.
[0156] Since the computer-readable storage medium described in Example 4 of the present invention is the computer-readable storage medium used to implement the personalized differential privacy protection method based on small sample data in Example 1 of the present invention, the specific structure and variations of the computer-readable storage medium are well understood by those skilled in the art based on the method described in Example 1 of the present invention, and therefore will not be described in detail here. All computer-readable storage media used in the method of Example 1 of the present invention fall within the scope of protection of the present invention.
[0157] Example 5
[0158] Based on the same inventive concept, the present application also provides a computer device, such as Figure 5As shown, it includes a memory 401, a processor 402 and a computer program 403 stored in the memory and executable on the processor. When the processor executes the above program, the method in the first embodiment is implemented.
[0159] Since the computer device described in Example 5 of the present invention is the computer device used to implement the personalized differential privacy protection method based on small sample data in Example 1 of the present invention, the specific structure and variations of the computer device are well understood by those skilled in the art based on the method described in Example 1 of the present invention, and therefore will not be described in detail here. All computer devices used in the method of Example 1 of the present invention fall within the scope of protection of the present invention.
[0160] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0161] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0162] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0163] Obviously, those skilled in the art may make various changes and modifications to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if such changes and modifications of the embodiments of the present invention fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A personalized differential privacy protection method based on small sample data, characterized by: include: The data aggregator receives the privacy protection level sent by the user; Divide users into groups based on their privacy protection levels, and group users with the same privacy protection level into the same group; Based on the user's privacy protection level and the minimum mean square error (MMSE) expectation criterion, the optimal perturbation method for the user's private data in each group is determined. This optimal perturbation method is sent to the corresponding users in the group. Each user in the group uses the corresponding optimal perturbation method to perturb their private data, obtaining perturbed data and sending it to the data aggregator. The optimal perturbation method includes perturbation probability, transmission parameter quantity, and perturbation output form. Using weighted aggregation, we aggregate the perturbation data from different groups at different privacy protection levels to obtain a statistical estimate of the user's private data. The optimal perturbation method for user privacy data in each group is determined based on the user's privacy protection level and the expected minimum mean square error criterion, including: Based on maximum likelihood estimation, the probability distribution of private data is estimated, and the expectation of the minimum mean square error of the estimated distribution is obtained; According to the expectation of the minimum mean square error, whether the perturbed data contains real data, and the probability that the perturbed data contains a predetermined value, the relationship between the privacy protection level, the perturbation probability, and the transmission parameter quantity is obtained according to the definition of differential privacy; Based on the minimum mean square error expectation and the privacy protection level of each group, an objective function is constructed and optimized to determine the best perturbation method for user privacy data in each group. Based on maximum likelihood estimation, the probability distribution of private data is estimated, and the expectation of the minimum mean square error of the estimated distribution is obtained, including: Based on the maximum likelihood estimation, the probability of private data is estimated, where the i-th private data x i The true probability Estimated value of The calculation formula is: in, N ′ The disturbed data contains x i The number of data items, G ′ is the τth group, N ′ For group G ′ The number of users in the corresponding ′ perturbation data, real data x=x i When the perturbation data contains x i The probability is p τ , that is, the perturbation probability. On the contrary, the perturbation data contains x j The probability is q τ , x j with x i For different data, s τ is the number of data contained in a disturbance data, i.e., the transmission parameter; Based on the estimated value of the true probability of each type of private data, we obtain the estimated value of the true probability distribution of all private data, and further obtain the expectation of the minimum mean square error of the estimated distribution, which is specifically: k is the value of the private data.
2. The personalized differential privacy protection method based on small sample data according to claim 1, characterized in that: Based on the expectation of minimum mean square error, whether the perturbed data contains real data, and the probability that the perturbed data contains a predetermined value, the relationship between the privacy protection level, the perturbation probability, and the transmission parameter quantity is obtained according to the definition of differential privacy, including: According to whether the disturbance data contains real data, the disturbance data contains real data, and the disturbance data contains predetermined s τ -Probability P of 1 privacy value u,x , when the disturbance data does not contain the real data, the disturbance data contains the predetermined s τ The probability of a value Among them, the real data x=x i When the perturbation data contains x i The probability is p τ , that is, the perturbation probability, s τ is the number of data contained in a piece of disturbed data, and k is the value of the private data; According to the expectation of minimum mean square error, whether the perturbed data contains real data, and the probability that the perturbed data contains a predetermined value, the relationship between the privacy protection level, the perturbation probability, and the transmission parameter is obtained according to the definition of differential privacy: Among them, |·| is the absolute value operation, ∈ τ is the privacy protection level of the τth group.
3. The personalized differential privacy protection method based on small sample data according to claim 1, characterized in that: Based on the expectation of the minimum mean square error and the privacy protection level of each group, the objective function is constructed and implemented in the following way: 0<p τ <1 s τ ∈{1,2,....,k} Then the optimal transmission parameter s τ and the perturbation probability p τ satisfy: in, It is a floor operation.
4. The personalized differential privacy protection method based on small sample data according to claim 1, characterized in that: Using weighted aggregation, we aggregate perturbation data from different groups at different privacy protection levels to obtain statistical estimates of the user's private data, including: Construct weighting factors based on the expectation of minimum mean square error: Among them, w τ is the weighting factor of the τth group, l is the counting symbol, ranging from 1 to m, E[MSE] τ is the expectation of the minimum mean square error of the τth group; Based on the weighting factor, the estimated values of the true probability of each type of private data in the m groups are weighted and aggregated to obtain the statistical estimate of the user's private data. Then the i-th type of private data x i The estimated probability of: in, The user's private data x i The statistical estimate of is the privacy data x of the τth group user i An estimate of the true probability of .
5. A personalized differential privacy protection device based on small sample data, characterized in that: The device is a data aggregator and includes: A privacy protection level receiving module is used to receive the privacy protection level sent by the user; The group division module is used to divide users into groups according to their privacy protection levels, and divide users with the same privacy protection level into the same group; The perturbation method determination module is used to determine the optimal perturbation method for the private data of users in each group based on the user's privacy protection level and the expected minimum mean square error criterion. The optimal perturbation method is sent to the corresponding users in the group. Each user in the group uses the corresponding optimal perturbation method to perform privacy protection operations on their respective private data, obtain perturbed data, and send it to the data aggregator. The optimal perturbation method includes the perturbation probability, the amount of transmission parameters, and the perturbation output form. The weighted aggregation module is used to aggregate the perturbation data from different groups at different privacy protection levels using a weighted aggregation method to obtain a statistical estimate of the user's private data; The optimal perturbation method for user privacy data in each group is determined based on the user's privacy protection level and the expected minimum mean square error criterion, including: Based on maximum likelihood estimation, the probability distribution of private data is estimated, and the expectation of the minimum mean square error of the estimated distribution is obtained; According to the expectation of the minimum mean square error, whether the perturbed data contains real data, and the probability that the perturbed data contains a predetermined value, the relationship between the privacy protection level, the perturbation probability, and the transmission parameter quantity is obtained according to the definition of differential privacy; Based on the minimum mean square error expectation and the privacy protection level of each group, an objective function is constructed and optimized to determine the best perturbation method for user privacy data in each group. Based on maximum likelihood estimation, the probability distribution of private data is estimated, and the expectation of the minimum mean square error of the estimated distribution is obtained, including: Based on the maximum likelihood estimation, the probability of private data is estimated, where the i-th private data x i The true probability Estimated value of The calculation formula is: in, N τ The disturbed data contains x i The number of data items, G τ is the τth group, N τ For group G τ The number of users in the corresponding τ perturbation data, real data x=x i When the perturbation data contains x i The probability is p τ , that is, the perturbation probability. On the contrary, the perturbation data contains x j The probability is q τ , x j with x i For different data, s τ is the number of data contained in a disturbance data, i.e., the transmission parameter; Based on the estimated value of the true probability of each type of private data, we obtain the estimated value of the true probability distribution of all private data, and further obtain the expectation of the minimum mean square error of the estimated distribution, which is specifically: k is the value of the private data.
6. A personalized differential privacy protection system based on small sample data, characterized by: It includes a personalized differential privacy protection device based on small sample data as described in claim 5 and a user terminal, wherein the user terminal is used to send a privacy protection level to a data aggregator, perturb its private data according to the optimal perturbation method sent by the data aggregator, obtain perturbation data, and send it to the data aggregator.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed, the method according to any one of claims 1 to 4 is implemented.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Identification method of differential privacy DNA motif based on random sampling and phantom compression
CN108664807A
Activity time sequence track mining method based on local differential privacy
CN110569286A