Adaptive privacy protection method, device and system based on minimum mean square error criterion

Through an adaptive privacy protection method based on the minimum mean square error criterion, users are grouped according to their privacy needs and the optimal perturbation method and probability are selected. Combined with multiple perturbation strategies, the problem of privacy requirement mismatch in existing technologies is solved, and the accuracy of statistical estimation and data availability are improved.

CN115879152BActive Publication Date: 2025-10-10HUBEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211578970.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-05
Publication Date
2025-10-10
Estimated Expiration
2042-12-05

AI Technical Summary

Technical Problem

In the existing technology, the localized differential privacy protection method of random response cannot be adaptively adjusted according to the user's personalized privacy needs, resulting in insufficient protection for users with high privacy needs and overprotection for users with low needs, affecting the accuracy of statistical estimation and the enthusiasm of users to participate in data sharing.

Method used

An adaptive privacy protection method based on the minimum mean square error criterion is adopted. By clustering, adaptively selecting perturbation methods and perturbation probabilities, and combining multiple perturbation strategies, personalized privacy protection and data aggregation are achieved, thereby improving the accuracy of statistical estimation.

Benefits of technology

It achieves personalized protection based on user privacy needs, improves the accuracy of statistical estimation and data availability, enhances user participation enthusiasm, and increases the sample size equivalently without leaking additional privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115879152B_ABST
    Figure CN115879152B_ABST
Patent Text Reader

Abstract

The application provides an adaptive privacy protection method, device and system based on a minimum mean square error criterion, adaptive selection of an optimal perturbation method for data perturbation and adaptive selection of an optimal perturbation probability for output of perturbed data are included, the method not only realizes personalized privacy protection, and higher data utility can be obtained through weighted aggregation. Wherein, based on the minimum mean square error, the adaptive boundaries of two kinds of classic localized differential privacy technologies, basic RAPPOR technology and k-RR technology, are derived, and the participants at the local end adaptively select one of the two kinds of localized differential privacy technologies as the optimal data perturbation method based on the adaptive boundaries, and adaptively select the optimal perturbation probability according to the privacy demand to perturb and output. In addition, the application gives a multiple perturbation data expansion strategy, which equivalently increases the sample size of certain subgroups with high privacy demand without revealing additional privacy, thereby further improving the accuracy of statistical estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of privacy data protection technology, and in particular to an adaptive privacy protection method, device and system based on a minimum mean square error criterion. Background Art

[0002] Randomized Response (RR) is a mainstream perturbation mechanism in localized differential privacy (LDP) based on data distortion. Its model is simple, intuitive, and easy to implement, and its perturbation level can be directly quantified. It has excellent performance in estimating statistical properties, thus attracting widespread attention. RR protects the privacy of data providers (or respondents) by using a probabilistic approach to answer questions, ensuring strong deniability for sensitive questions and privacy protection. It has been implemented in Google Chrome's privacy protection tools and Apple's system. RR also fully considers the possibility of data collectors stealing or leaking user privacy during data collection. In this model, respondents can independently perform privacy processing on their individual data, preventing even data collectors from obtaining the exact original private data, greatly motivating them to participate in data collection. Therefore, unlike centralized privacy protection mechanisms that assume a trusted third party, the RR-based localized differential privacy protection mechanism eliminates the need for a trusted third party and eliminates the potential for privacy leaks and attacks from untrusted third-party data collectors.

[0003] However, in reality, different individuals have different privacy protection needs. If all users' data is rigidly protected at the same level, users with high privacy needs will be underprotected, while those with low privacy needs will be overprotected. This will not only lead to user opposition to data openness and sharing, but also reduce the accuracy of statistical estimates to a certain extent. Summary of the Invention

[0004] The purpose of the present invention is to fully consider the personalized privacy needs of local users in the localized data collection of random responses, and to provide an adaptive privacy protection method based on the minimum mean square error criterion, which includes adaptive perturbation method selection and adaptive perturbation probability selection, and constructs appropriate weighting factors based on the minimum mean square error to perform weighted aggregation to improve the accuracy of statistical estimation. At the same time, a data expansion strategy with multiple perturbations is adopted to equivalently increase the sample size of certain subgroups without leaking additional privacy, thereby further improving the availability of data.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] The first aspect provides an adaptive privacy protection method based on the minimum mean square error criterion, including:

[0007] The data aggregator receives the privacy protection level sent by the local end user;

[0008] Local users are grouped according to their privacy protection levels, and users with the same privacy protection level are grouped into the same subgroup.

[0009] Based on the localized differential privacy technology and the privacy protection level, the optimal perturbation probability is determined. Based on the minimum mean square error criterion, the adaptive boundaries of two classic localized differential privacy technologies are determined. The optimal data perturbation method is selected based on the adaptive boundaries, and the adaptive results are sent to users in the corresponding subgroups. This allows users in each subgroup to use the corresponding optimal data perturbation method to perturb their private data and perform privacy protection operations using the optimal perturbation probability. The perturbed data is obtained and sent to the data aggregator. The adaptive results include the optimal data perturbation method and the optimal perturbation probability.

[0010] A weighting factor is constructed based on the minimum mean square error, and the perturbed data sent from each subgroup under different privacy protection levels are aggregated to obtain a statistical estimate of the local user's private data.

[0011] In one embodiment, two classic localized differential privacy techniques include the basic RAPPOR technique or the k-RR technique. Determining the optimal perturbation probability based on the localized differential privacy technique and the privacy protection level includes:

[0012] When the localized differential privacy technology used is the basic RAPPOR technology, at the privacy protection level ∈, the optimal perturbation probability for each bit of the binary-encoded private data is:

[0013]

[0014] Among them,∈Privacy protection level;

[0015] When the localized differential privacy technology used is the k-RR technology, at the privacy protection level ∈, the optimal perturbation probability for each bit of the binary-encoded private data is:

[0016]

[0017] The above formula means that the private data is kept at its original value with probability p, and is perturbed to output any of the other k-1 types with probability (1-p) / , where k is the number of different private data.

[0018] In one embodiment, two classic localized differential privacy techniques include the basic RAPPOR technique or the k-RR technique. Based on the minimum mean square error criterion, the adaptive boundaries of the two classic localized differential privacy techniques are determined, including:

[0019] The first estimation error of the privacy distribution when using the basic RAPPOR technique is calculated based on the maximum likelihood estimation criterion:

[0020]

[0021] The second estimation error of the privacy distribution when using the k-RR technique is calculated based on the maximum likelihood estimation criterion:

[0022]

[0023] Where n represents the amount of data or the number of users, ∈ is the privacy protection level, also known as the privacy budget, and x i is the i-th type of private data, private data x i The true probability is P i , k is the number of different private data, is the first estimation error, is the second estimation error;

[0024] The adaptive boundaries of two classic localized differential privacy techniques are determined based on the first estimation error and the second estimation error.

[0025] In one embodiment, determining the adaptive boundaries of two classic localized differential privacy techniques based on the first estimation error and the second estimation error includes:

[0026] Build Function Then the value of ΔMSE at zero point is:

[0027]

[0028] Among them, the expressions of u and v are:

[0029]

[0030]

[0031] Will ∈ * As the optimal adaptive boundary of basic RAPPOR technique and k-RR technique under the minimum MSE criterion.

[0032] In one embodiment, a weighting factor is constructed based on the minimum mean square error, and the perturbed data sent from each subgroup at different privacy protection levels is aggregated to obtain a statistical estimate of the local user's private data, including:

[0033] Construct weighting factors based on the expectation of minimum mean square error:

[0034]

[0035] w τ is the weighting factor of the vth subgroup and satisfies l is a counting symbol, ranging from 1 to m, MSE τ is the τth subgroup at the privacy protection level ∈ τ The mean square error value of the estimated distribution under;

[0036] Based on the constructed weighting factors, the perturbed data in the m subgroups are weighted and aggregated to obtain a statistical estimate of the local user's private data:

[0037]

[0038] in, is the estimated distribution of private data for the τth subgroup, m is the total number of subgroups, It is a statistical estimate of local user privacy data.

[0039] In one embodiment, the method further includes: expanding the data in the subgroup using a multi-perturbation data expansion strategy to equivalently increase the number of privacy subgroups with high privacy requirements.

[0040] Based on the same inventive concept, the second aspect of the present invention provides an adaptive privacy protection device based on the minimum mean square error criterion, comprising:

[0041] Privacy protection level receiving module: the data aggregator receives the privacy protection level sent by the local user;

[0042] The group segmentation module is used to group local users according to their privacy protection level, and divide users with the same privacy protection level into the same subgroup;

[0043] The adaptive result generation module is used to determine the optimal perturbation probability based on the localized differential privacy technology and the privacy protection level. Based on the minimum mean square error criterion, it determines the adaptive boundaries of two classic localized differential privacy technologies, selects the optimal data perturbation method based on the adaptive boundaries, and sends the adaptive results to users in the corresponding subgroups, so that users in each subgroup use the corresponding optimal data perturbation method to perturb their private data and perform privacy protection operations using the optimal perturbation probability. The perturbed data is obtained and sent to the data aggregator. The adaptive results include the optimal data perturbation method and the optimal perturbation probability.

[0044] The weighted aggregation module is used to construct a weighting factor based on the minimum mean square error, aggregate the perturbed data sent from each subgroup at different privacy protection levels, and obtain a statistical estimate of the local user's private data.

[0045] Based on the same inventive concept, the third aspect of the present invention provides an adaptive privacy protection system based on the minimum mean square error criterion, including the adaptive privacy protection device based on the minimum mean square error criterion described in the second aspect and a local user terminal, wherein the local user terminal is used to send the privacy protection level to the data aggregator, select the best perturbation method to perturb the private data according to the adaptive result sent by the data aggregator, and use the best perturbation probability to perform the privacy protection operation to obtain the perturbed data, and send it to the data aggregator.

[0046] Based on the same inventive concept, the fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the method described in the first aspect when the program is executed.

[0047] Based on the same inventive concept, the fifth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.

[0048] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows:

[0049] This paper considers individual privacy needs and proposes an adaptive privacy protection method based on minimum mean square error (MMSE). This adaptation involves two levels: an adaptive perturbation method and an adaptive perturbation probability. First, the MSE is used to measure the accuracy (i.e., usability) of the privacy distribution estimate, and the optimal adaptive bound is derived from the perspective of MSE. Second, a data collection protocol for localized adaptive privacy protection is designed: local participants adaptively select the optimal LDP algorithm based on the adaptive bound according to the corresponding personalized privacy protection level, and perturb the data using the optimal adaptive probability, which is then uploaded to a third party (data aggregator). Finally, weighted aggregation is introduced for efficient data aggregation to obtain highly usable statistical analysis.

[0050] Furthermore, a multi-perturbation data expansion strategy is introduced, which equivalently increases the sample size of certain subgroups without leaking additional privacy, which can further improve the availability of data. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0052] Figure 1 is a flow chart of a data collection framework of an adaptive privacy protection method provided by an embodiment of the present invention;

[0053] Figure 2 1 is a framework diagram of an adaptive privacy protection method for dividing four subgroups (∈1=0.1, ∈2=0.5, ∈3=1.0, ∈4=1.5) in an embodiment of the present invention;

[0054] Figure 3 This is a framework diagram of an adaptive privacy protection method (∈1=0.1, ∈2=0.5, ∈3=1.0, ∈4=1.5) that divides data into four subgroups and combines multiple perturbations in an embodiment of the present invention;

[0055] Figure 4 This is a structural block diagram of an adaptive privacy protection device based on the minimum mean square error criterion provided by an embodiment of the present invention;

[0056] Figure 5 A schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present invention;

[0057] Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0058] The purpose of the present invention is to fully consider the personalized privacy needs of local users in the localized data collection of random responses, and to provide an adaptive privacy protection method based on the minimum mean square error criterion, which includes adaptive perturbation method selection and adaptive perturbation probability selection, and constructs appropriate weighting factors based on the minimum mean square error to perform weighted aggregation to improve the accuracy of statistical estimation. At the same time, a data expansion strategy with multiple perturbations is adopted to equivalently increase the sample size of certain subgroups without leaking additional privacy, thereby further improving the availability of data.

[0059] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solution: in localized data collection, participants can perform personalized privacy protection processing according to their own personalized privacy needs, where the privacy protection needs are measured by the differential privacy parameter ∈, and are divided into different subgroups according to their personalized privacy needs. The privacy needs of users in the same subgroup are the same, and the privacy needs of users in different subgroups are different. According to the privacy needs of each subgroup, based on the minimum mean square error criterion, an optimal perturbation method (basic RAPPOR technology or k-RR technology) is adaptively selected, and the optimal perturbation probability is adaptively selected for data perturbation output. Based on the minimum mean square error, a suitable weighting factor is constructed to perform weighted aggregation to improve the accuracy of statistical estimation, while achieving personalized privacy protection and ensuring high statistical accuracy. At the same time, a data expansion strategy of multiple perturbations is adopted to equivalently increase the sample size without leaking additional privacy, further improving the availability of data.

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0061] Example 1

[0062] The embodiment of the present invention provides an adaptive privacy protection method based on the minimum mean square error criterion, including:

[0063] The data aggregator receives the privacy protection level sent by the local end user;

[0064] Local users are grouped according to their privacy protection levels, and users with the same privacy protection level are grouped into the same subgroup.

[0065] Based on the localized differential privacy technology and the privacy protection level, the optimal perturbation probability is determined. Based on the minimum mean square error criterion, the adaptive boundaries of two classic localized differential privacy technologies are determined. The optimal data perturbation method is selected based on the adaptive boundaries, and the adaptive results are sent to users in the corresponding subgroups. This allows users in each subgroup to use the corresponding optimal data perturbation method to perturb their private data and perform privacy protection operations using the optimal perturbation probability. The perturbed data is obtained and sent to the data aggregator. The adaptive results include the optimal data perturbation method and the optimal perturbation probability.

[0066] A weighting factor is constructed based on the minimum mean square error, and the perturbed data sent from each subgroup under different privacy protection levels are aggregated to obtain a statistical estimate of the local user's private data.

[0067] In a specific application scenario, the privacy protection method proposed in the present invention includes the following steps:

[0068] Step (1): The local user determines the personalized privacy protection level and sends it to the data collector or aggregator.

[0069] Step (2): The data collector or aggregator groups local users according to the privacy protection level, and users with the same privacy protection level are divided into a subgroup.

[0070] Step (3): The data aggregator determines the adaptive boundary based on the minimum mean square error criterion according to the privacy protection level in step (1), so that each subgroup of users can adaptively select the best perturbation method (basic RAPPOR technology or k-RR technology) and the best perturbation probability, and then sends the adaptive selection result to the local end user.

[0071] Step (4): Users in each subgroup perceive their private data using the best perturbation method, perform privacy protection operations on them using the best perturbation probability, and send the perturbed data to the data aggregator.

[0072] Step (5): The data aggregator aggregates the perturbed data from subgroups at different privacy protection levels, constructs appropriate weighting factors based on the minimum mean square error, and then uses weighted aggregation to obtain better statistical estimates.

[0073] In traditional research work based on localized differential privacy, it is assumed that the privacy protection parameters are completely determined by the data collector or aggregator and then distributed to all participants. However, due to different privacy preferences, it is unreasonable to require all participants to adopt the same privacy protection strength during the data collection or aggregation process. The present invention proposes an adaptive privacy protection method based on the minimum mean square error (MSE) criterion. The adaptation includes adaptively selecting the best perturbation method for data perturbation and adaptively selecting the best perturbation probability to output perturbation data. This method not only achieves personalized privacy protection, but also obtains higher data utility through weighted aggregation. Among them, based on the minimum MSE, the adaptive boundaries of two classic localized differential privacy technologies - basic RAPPOR technology and k-RR technology are derived. Based on this adaptive boundary, the local participants adaptively select one of the above two localized differential privacy technologies as the best data perturbation method, and adaptively select the best perturbation probability for perturbation output according to privacy requirements. In addition, the present invention proposes a data expansion strategy with multiple perturbations, which equivalently increases the sample size of a subpopulation without leaking additional privacy, thereby further improving the accuracy of statistical estimation. It is a practical method for statistics and analysis with strong practical significance.

[0074] In this implementation, the adaptive perturbation method and adaptive perturbation probability are derived based on the minimum mean square error (MMSE) criterion, which can improve the accuracy of the estimation to a certain extent. Weighted aggregation, also based on the MSE, further improves the accuracy of the privacy distribution estimation compared to direct aggregation.

[0075] In one embodiment, two classic localized differential privacy techniques include the basic RAPPOR technique or the k-RR technique. Determining the optimal perturbation probability based on the localized differential privacy technique and the privacy protection level includes:

[0076] When the localized differential privacy technology used is the basic RAPPOR technology, at the privacy protection level ∈, the optimal perturbation probability for each bit of the binary-encoded private data is:

[0077]

[0078] Among them,∈Privacy protection level;

[0079] When the localized differential privacy technology used is the k-RR technology, at the privacy protection level ∈, the optimal perturbation probability for each bit of the binary-encoded private data is:

[0080]

[0081] The above formula means that the private data is kept at its original value with probability p, and is perturbed to output any of the other k-1 types with probability (1-p) / , where k is the number of different private data.

[0082] In the specific real-time process, Figure 1 As shown in the figure: Assume that the local user has m personalized privacy protection levels, namely ∈1,∈2,…,∈ m , can be divided into m subgroups G1, G2, ..., G m , the number of users in the subgroups is n1, n2, ..., n m The total sample size is Where m ≥ 2 and is an integer.

[0083] Without loss of generality, assume that each user has only one kind of private data x∈X={x1,x2,…,x k}, where k is the number of different private data, x i is the i-th type of private data. Let private data x i The true probability is P i , after the privacy protection operation, the estimated probability is The mean square error (MSE) is defined to measure the accuracy of statistical estimation:

[0084]

[0085] in, is the estimated prior distribution of private data, P X =[P1,P2,…,P k ] is the true distribution of private data, E represents the expected operation, is the two-norm operation.

[0086] In one embodiment, two classic localized differential privacy techniques include the basic RAPPOR technique or the k-RR technique. Based on the minimum mean square error criterion, the adaptive boundaries of the two classic localized differential privacy techniques are determined, including:

[0087] The first estimation error of the privacy distribution when using the basic RAPPOR technique is calculated based on the maximum likelihood estimation criterion:

[0088]

[0089] The second estimation error of the privacy distribution when using the k-RR technique is calculated based on the maximum likelihood estimation criterion:

[0090]

[0091] where n represents the data amount or the number of users, ∈ is the privacy protection level, also known as the privacy budget, x i is the ith privacy data, the real probability of the privacy data x i is P i , k is the number of different privacy data, is the first estimation error, is the second estimation error;

[0092] The adaptive boundary of two classical localized differential privacy technologies is determined according to the first estimation error and the second estimation error.

[0093] In an embodiment, the adaptive boundary of two classical localized differential privacy technologies is determined according to the first estimation error and the second estimation error, comprising:

[0094] The function is constructed, and the value at the zero point of the ΔMSE is:

[0095]

[0096] where the expressions of u and v are:

[0097]

[0098]

[0099] The ∈ * is taken as the optimal adaptive boundary of the basic RAPPOR technology and the k-RR technology under the minimum MSE criterion.

[0100] Specifically, after the optimal adaptive boundary is obtained, the selection or determination mode of the adaptive perturbation mode is as follows:

[0101] When ∈≥∈ * , there is It is indicated that under the same conditions, the estimation accuracy of the privacy data obtained by the localized differential privacy technology using the k-RR technology is higher than that using the basic RAPPOR technology, and the k-RR technology is adaptively selected as the perturbation mode;

[0102] When ∈<∈ * , there is It is indicated that under the same conditions, the estimation accuracy of the privacy data obtained by the localized differential privacy technology using the basic RAPPOR technology is higher than that using the k-RR technology, and the basic RAPPOR technology is adaptively selected as the perturbation mode.

[0103] For example, the private data X is the patient's condition, with a total of k = 8 different values. The private data value set is X = {lung cancer, liver cancer, heart disease, coronary heart disease, cold, AIDS, indigestion, lung nodules}. At this time, the optimal adaptive boundary ∈ * =7809. The following two cases are explained:

[0104] (1) Assume that the privacy protection requirement of a participant u is ∈=0.5, then ∈<∈ * , based on the minimum MSE criterion, the basic RAPPOR technology is adaptively selected as the perturbation method, and the optimal adaptive perturbation probability The set of private data values ​​is X = {lung cancer, liver cancer, heart disease, coronary heart disease, common cold, AIDS, indigestion, lung nodules}. The private data needs to be encoded, and the encoding result is X = {lung cancer, liver cancer, heart disease, coronary heart disease, common cold, AIDS, indigestion, lung nodules} = {10000000, 01000000, 0010000, 0001000, 00001000, 00000100, 00000010, 00000001}. Assuming that participant u's true private data is x = "liver cancer," then its private encoded data is "01000000." In this case, a one-to-one perturbation process is performed on the private encoded data. For each bit of data, a coin is tossed, with a probability of 0.5622 for heads and 0.4378 for tails. If the face is facing up, the bit remains the same; if the face is facing up, the bit is flipped.

[0105] (2) Assume that the privacy protection requirement of a participant u is ∈=1.0, then ∈>∈ * , based on the minimum MSE criterion, the k-RR technology is adaptively selected as the perturbation method, and the optimal adaptive perturbation probability Suppose participant u's true private data is x = "liver cancer." A coin is tossed, and the probability of heads is 0.2797. If heads, the perturbation output is {"liver cancer"}. If tails, a random value is selected from the set {lung cancer, heart disease, coronary heart disease, cold, AIDS, indigestion, lung nodules} (excluding "liver cancer") as the perturbation output.

[0106] In order to balance privacy and usability, and to obtain higher data utility while satisfying personalized privacy protection, the local end adopts the idea of ​​adaptive privacy protection algorithm to perform personalized data perturbation. The privacy level set is defined as {∈1,∈2,…,∈ m}, where m is the number of different privacy levels. Adaptive data collection framework such as Figure 1 As shown, simple steps:

[0107] Step (1): The local user determines the personalized privacy protection level and sends it to the data collector or aggregator.

[0108] Step (2): The data collector or aggregator groups local users according to the privacy protection level, and users with the same privacy protection level are divided into a subgroup.

[0109] Step (3): The data aggregator determines the adaptive boundary and the optimal perturbation probability based on the privacy protection level in step (1) and the minimum mean square error criterion, so that each subgroup of users adaptively selects the optimal perturbation method (basicRAPPOR technology or k-RR technology) and the optimal perturbation probability, and then sends the adaptive selection result to the local end user.

[0110] Step (4): Users in each subgroup perceive their private data using the best perturbation method, perform privacy protection operations on them using the best perturbation probability, and send the perturbed data to the data aggregator.

[0111] Step (5): The data aggregator aggregates the perturbed data from subgroups at different privacy protection levels, constructs appropriate weighting factors based on the minimum mean square error, and then uses weighted aggregation to obtain better statistical estimates.

[0112] In one embodiment, a weighting factor is constructed based on the minimum mean square error, and the perturbed data sent from each subgroup at different privacy protection levels is aggregated to obtain a statistical estimate of the local user's private data, including:

[0113] Construct weighting factors based on the expectation of minimum mean square error:

[0114]

[0115] w τ is the weighting factor of the τth subgroup and satisfies l is a counting symbol, ranging from 1 to m, MSE τ is the τth subgroup at the privacy protection level ∈ τ The mean square error value of the estimated distribution under;

[0116] Based on the constructed weighting factors, the perturbed data in the m subgroups are weighted and aggregated to obtain a statistical estimate of the local user's private data:

[0117]

[0118] in, is the estimated distribution of private data for the τth subgroup, m is the total number of subgroups, It is a statistical estimate of local user privacy data.

[0119] Specifically, considering the contribution of different privacy groups to the accuracy of statistical estimation, a weighted aggregation method is designed based on MSE to improve the accuracy of statistical estimation.

[0120] For m subgroups, the private data x i The estimated distribution of : Use Figure 1 The weighted aggregation shown is used to obtain the final private data x i Statistical estimates of

[0121] In a specific example, we also use the above example as an example: Assume that there are k = 8 different values, and the set of private data values ​​is X = {lung cancer, liver cancer, heart disease, coronary heart disease, cold, AIDS, indigestion, lung nodules}. At this time, the optimal adaptive boundary ∈ * =7809. Assume that there are m=4 privacy protection levels, which are ∈1=0.1, ∈2=0.5, ∈3=1.0, ∈4=1.5, that is, 4 subgroups are divided, such as Figure 2 The following two cases are described:

[0122] (1) Based on the minimum MSE criterion, the subgroups of privacy levels ∈1 and ∈2 are adaptively treated using the basic RAPPOR technique, and the adaptive perturbation probabilities are

[0123]

[0124]

[0125] (2) Based on the minimum MSE criterion, the subgroups with privacy levels ∈3 and ∈4 are adaptively treated using the k-RR technique, and the adaptive perturbation probabilities are

[0126]

[0127]

[0128] Then, each subgroup performs statistical estimation of private data, and finally performs weighted aggregation to obtain better estimation accuracy.

[0129] In one embodiment, the method further includes: expanding the data in the sub-population using a multi-perturbation data expansion strategy.

[0130] In a specific implementation of the present invention, the multi-perturbation data expansion strategy is to equivalently increase the sample size while leaking additional privacy, further improving the accuracy of statistical estimation.

[0131] Specifically, considering that data providers with low privacy requirements are also willing to provide perturbed data with high privacy protection levels, as long as no additional privacy leakage occurs, i.e., does not exceed the respective maximum privacy budget. According to the composition property of local differential privacy, multiple independent perturbations are performed on the original privacy data, and the privacy budget has an additive property, thereby exceeding the maximum privacy budget, i.e., a cooperative hierarchical gain is generated, causing additional privacy leakage. Based on this, the present application introduces correlation between different perturbed versions with different privacy levels to eliminate the hierarchical gain generated by cooperation between different perturbed versions.

[0132] Based on the concatenation property of symmetric channels in information theory and coding, a multi-level correlation perturbation strategy is designed. The design principle is: without exceeding the maximum privacy budget of the data provider with low privacy requirements, the number of samples of the group with high privacy requirements is increased to further improve the accuracy of statistical estimation, effectively balancing privacy and usability. For ease of description, assume that m , and satisfy: f * f+1 … <∈ m , that is, the first f subgroups use the basic RAPPOR technology, and the last s subgroups use the k-RR technology, where f and s are integers, and f+s=m.

[0133] For the first f subgroups, the privacy requirement τ * The perturbation probability of the privacy data is

[0134]

[0135] At this time, the perturbation processing can generate n τ binary bit strings τ with privacy protection level

[0136] On the other hand, the subgroups with privacy requirement τ can also provide perturbed data with high privacy protection, provided that no additional privacy is leaked. According to the composition property of differential privacy, independent perturbation processing cannot be directly performed on the original privacy data of the subgroups with privacy requirement τ , otherwise the maximum privacy budget will be exceeded due to the additive property. Therefore, correlation can be introduced between different perturbed versions, and perturbation is performed on the perturbed data set , which destroys the composition property of differential privacy, and the correlation perturbation probability p′ τ-1 satisfies:

[0137] ​​​

[0138] where p′ τ-1 Indicates the perturbation version The privacy protection level is obtained by re-perturbing τ-1 The perturbation data At this time, after perturbation processing, n τ The privacy protection level is ∈ τ-1 The perturbation bit string This is equivalent to increasing the privacy protection level to ∈ τ-1 The sample size of the subgroup is determined, and the original dataset is ensured Above is p τ-1 The statistical characteristics of the probability perturbation are equivalent to the perturbed data set p′ τ-1 The statistical characteristics of the probability perturbation are obtained, which improves ∈ τ-1 The accuracy of sub-group privacy data statistics under different privacy protection levels.

[0139] For the last s subgroups, privacy requirements ∈ τ >∈ * Privacy data The perturbation probability is

[0140]

[0141] At this time, after perturbation processing, n τ The privacy protection level is ∈ τ The perturbation data

[0142] Similarly, privacy requirements∈ τ The subgroup of can also provide perturbation data with high privacy protection, provided that their additional privacy is not disclosed. According to the combination characteristics of differential privacy, it is not possible to τ Independent perturbation processing can be performed directly on the original privacy data of , otherwise it will exceed its maximum privacy budget due to accumulation. Perturb the data to destroy the combination characteristics of differential privacy, and the perturbation probability satisfies:

[0143]

[0144] At this time, after perturbation processing, n τ The privacy protection level is ∈ τ-1 The perturbation data This is equivalent to increasing the privacy protection level to ∈ τ-1 The sample size of the subgroup is determined, and the original dataset is ensured Above is p τ-1 The statistical characteristics of the probability perturbation are equivalent to the perturbed data set p′ τ-1 The statistical characteristics of the probability perturbation are obtained, which improves ∈ τ-1 The accuracy of sub-group privacy data statistics under different privacy protection levels.

[0145] Based on steps (4) and (5) of the aforementioned method, the above-mentioned multi-perturbation data expansion strategy is adopted to equivalently increase the sample size while leaking additional privacy, which can further improve the accuracy of statistical estimation.

[0146] See Figure 3 , is a framework diagram of an adaptive privacy protection method (∈1=0.1,∈2=0.5,∈3=1.0,∈4=1.5) that divides data into four subgroups and combines multiple perturbations in an embodiment of the present invention;

[0147] Let’s take the above example as an example: Assume there are k = 8 different values, and the set of private data values ​​is X = {lung cancer, liver cancer, heart disease, coronary heart disease, cold, AIDS, indigestion, lung nodules}. Then the optimal adaptive boundary ∈ * =7809. Assume that there are m = 4 privacy protection levels, ∈1 = 0.1, ∈2 = 0.5, ∈3 = 1.0, and ∈4 = 1.5, which means that there are 4 subgroups, and the amount of data in each subgroup is n1 = 5000, n2 = 3000, n3 = 2000, and n4 = 1000. The following two cases are explained:

[0148] (1) Adaptive basic RAPPOR technology for subgroups 1 and 2: The subgroup with privacy protection level ∈1 is perturbed to obtain perturbed data with privacy protection level ∈1, and a data expansion strategy of multiple perturbations is introduced on the perturbed data of the subgroup with privacy protection level ∈2. The perturbation probability is

[0149]

[0150] That is, by perturbing the perturbed data with a privacy protection level of ∈2 with a probability of p′1, the perturbed data with a privacy protection level of ∈1 can be obtained. At this time, the amount of perturbed data with a privacy protection level of ∈1 is n1+n2=5000+3000=8000. Protecting the perturbed data of the subgroup with a privacy protection level of ∈1 and the perturbed data of the subgroup with a privacy protection level of ∈2 is equivalent to increasing the amount of data of the subgroup with a privacy protection level of ∈1, and will not cause additional privacy leakage.

[0151] (2) For subgroups 3 and 4 of the adaptive k-RR technology, the subgroup with a privacy protection level of ∈3 is perturbed to obtain perturbed data with a privacy protection level of ∈3, and a data expansion strategy of multiple perturbations is introduced on the perturbed data of the subgroup with a privacy protection level of ∈4. The perturbation probability is

[0152]

[0153] That is, on the perturbation data with privacy protection level ∈4, ′ By perturbing the probability of n3=0.2674, we can obtain perturbed data with a privacy protection level of ∈3. At this time, the amount of perturbed data with a privacy protection level of ∈3 is n3+4=2000+1000=3000. The perturbed data of the subgroup with a privacy protection level of ∈3 and the perturbed data of the subgroup with a privacy protection level of ∈4 are equivalent to increasing the amount of data of the subgroup with a privacy protection level of ∈3 without causing additional privacy leakage.

[0154] Then, the sub-population with data expansion is statistically analyzed according to the perturbed data after expansion, and the sub-population without expansion is statistically analyzed according to the original situation, and finally weighted aggregation is performed.

[0155] Since the re-perturbation data expansion strategy is equivalent to increasing the sample size, it can further improve the accuracy of statistical estimation.

[0156] In general, the advantages and beneficial technical effects of the present invention include:

[0157] (1) It fully considers the personalized privacy needs of local users. While achieving personalized privacy protection, it also improves the enthusiasm and initiative of local users in participating in data collection and sharing to a certain extent.

[0158] (2) Based on the minimum mean square error criterion, two adaptive boundaries of classic localized differential privacy are theoretically derived and given. Participants adaptively choose appropriate perturbation methods according to the adaptive boundaries.

[0159] (3) Weighted aggregation based on minimum mean square error further improves the accuracy of privacy distribution estimation compared to direct aggregation;

[0160] (4) The data expansion strategy of multiple perturbations is to increase the sample size equivalently without leaking additional privacy, which can further improve the accuracy of statistical estimation;

[0161] (5) The adaptive privacy protection method proposed in this invention not only realizes personalized privacy protection, but also introduces a data expansion strategy with multiple disturbances to obtain statistical accuracy with higher availability. It is a practical method for statistics and analysis and has strong practical significance.

[0162] Example 2

[0163] Based on the same inventive concept, this embodiment provides an adaptive privacy protection device based on the minimum mean square error criterion, see Figure 4 , the device comprises:

[0164] The privacy protection level receiving module 201, the data aggregator receives the privacy protection level sent by the local end user;

[0165] A group division module 202 is used to group local users according to their privacy protection levels, and divide users with the same privacy protection level into the same subgroup;

[0166] Adaptive result generation module 203 is used to determine the optimal perturbation probability based on the localized differential privacy technology and the privacy protection level, determine the adaptive boundaries of two classic localized differential privacy technologies based on the minimum mean square error criterion, select the optimal data perturbation method based on the adaptive boundaries, and send the adaptive results to users in the corresponding subgroups, so that users in each subgroup use the corresponding optimal data perturbation method to perturb their private data and perform privacy protection operations using the optimal perturbation probability. The perturbed data is obtained and sent to the data aggregator. The adaptive results include the optimal data perturbation method and the optimal perturbation probability.

[0167] The weighted aggregation module 204 is used to construct a weighting factor based on the minimum mean square error, aggregate the perturbed data sent from each subgroup at different privacy protection levels, and obtain a statistical estimate of the local user's private data.

[0168] Since the apparatus described in Example 2 of the present invention is used to implement the adaptive privacy protection method based on the minimum mean square error criterion described in Example 1 of the present invention, the specific structure and variations of the apparatus are readily apparent to those skilled in the art based on the method described in Example 1 of the present invention, and thus will not be further described here. All apparatuses used in the method described in Example 1 of the present invention fall within the scope of protection of the present invention.

[0169] Example 3

[0170] Based on the same inventive concept, the present invention provides an adaptive privacy protection system based on the minimum mean square error criterion, including the adaptive privacy protection device based on the minimum mean square error criterion described in Example 2 and a local user terminal, wherein the local user terminal is used to send a privacy protection level to a data aggregator, select the best perturbation method to perturb the private data according to the adaptive result sent by the data aggregator, and use the best perturbation probability to perform a privacy protection operation to obtain the perturbed data, and send it to the data aggregator.

[0171] Since the system described in Example 3 of the present invention is used to implement the adaptive privacy protection method based on the minimum mean square error criterion in Example 1 of the present invention, the specific structure and variations of this system are readily apparent to those skilled in the art based on the method described in Example 1 of the present invention, and thus will not be further described here. All systems used in the method described in Example 1 of the present invention fall within the scope of protection of the present invention.

[0172] Example 4

[0173] Based on the same inventive concept, see Figure 5 The present invention further provides a computer-readable storage medium 300 on which a computer program 311 is stored. When the program is executed, the method described in the first embodiment is implemented.

[0174] Since the computer-readable storage medium described in Example 4 of the present invention is used to implement the adaptive privacy protection method using the minimum mean square error criterion in Example 1 of the present invention, the specific structure and variations of the computer-readable storage medium are readily apparent to those skilled in the art based on the method described in Example 1 of the present invention, and thus will not be further described here. All computer-readable storage media used in the method of Example 1 of the present invention fall within the scope of protection of the present invention.

[0175] Example 5

[0176] Based on the same inventive concept, the present application also provides a computer device, such as Figure 6 As shown, it includes a memory 401, a processor 402 and a computer program 403 stored in the memory and executable on the processor. When the processor executes the above program, the method in the first embodiment is implemented.

[0177] Since the computer device described in Example 5 of the present invention is used to implement the adaptive privacy protection method using the minimum mean square error criterion in Example 1 of the present invention, the specific structure and variations of the computer device are readily apparent to those skilled in the art based on the method described in Example 1 of the present invention, and thus will not be further described here. All computer devices used in the method described in Example 1 of the present invention fall within the scope of protection of the present invention.

[0178] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0179] The present application is described in reference to the drawings using specific terminology to describe embodiments thereof. But, the description herein is presented for purposes of illustration and description and is not intended to limit the scope of the application in that the scope of the application is limited only by the appended claims. Figure 1 one or more functions specified in the flow or flows and / or blocks Figure 1 one or more functions specified in the flow or flows and / or blocks

[0180] While the preferred embodiments of the application have been described, additional variations and modifications can be made to the preferred embodiments by those of skill in the art once they have the benefit of the present disclosure without departing from the spirit and scope of the application. Accordingly, it is intended that the appended claims be interpreted as including all such additional variations and modifications as fall within the scope of the present application.

[0181] It is apparent that those skilled in the art can make various changes and modifications to the preferred embodiments of the application without departing from the spirit and scope of the application. It is therefore intended that the present application cover all such changes and modifications that are within the scope of the present application as defined by the appended claims and their equivalents.

Claims

1. An adaptive privacy protection method based on the minimum mean square error criterion, characterized by: include: The data aggregator receives the privacy protection level sent by the local end user; Local users are grouped according to their privacy protection levels, and users with the same privacy protection level are grouped into the same subgroup. According to the localized differential privacy technology and the privacy protection level, the optimal perturbation probability is determined. Based on the minimum mean square error criterion, the adaptive boundaries of two classic localized differential privacy technologies are determined. The optimal data perturbation method is selected according to the adaptive boundary, and the adaptive results are sent to the users in the corresponding subgroup, so that the users in each subgroup use the corresponding optimal data perturbation method to perturb their private data, and use the optimal perturbation probability to perform privacy protection operations, obtain the perturbed data, and send it to the data aggregator. Among them, the adaptive results include the optimal data perturbation method and the optimal perturbation probability. The two classic localized differential privacy technologies include basic RAPPOR technology or k -RR technology, which determines the optimal perturbation probability based on localized differential privacy technology and privacy protection level, including: When the localized differential privacy technology used is basic RAPPOR technology, Under the privacy protection level, the optimal perturbation probability for each bit of binary-encoded private data is: in, Privacy protection level; When the localized differential privacy technology used is k -RR technology, Under the privacy protection level, the optimal perturbation probability for each bit of binary-encoded private data is: The above formula represents the privacy data The probability of keeping the original value is The probability of perturbation output other -1 of any kind, is the number of different private data; Based on the minimum mean square error criterion, the adaptive boundaries of two classic localized differential privacy technologies are determined, including: The first estimation error of the privacy distribution when using the basic RAPPOR technique is calculated based on the maximum likelihood estimation criterion: Based on the maximum likelihood estimation criterion, the k -RR technology uses the second estimation error of privacy distribution: in, Indicates the amount of data or the number of users, is the privacy protection level, also known as the privacy budget, For the i private data, private data The true probability is , is the number of different privacy data, is the first estimation error, is the second estimation error; The adaptive boundaries of two classic localized differential privacy techniques are determined based on the first estimation error and the second estimation error; A weighting factor is constructed based on the minimum mean square error, and the perturbed data sent from each subgroup under different privacy protection levels are aggregated to obtain a statistical estimate of the local user's private data.

2. The adaptive privacy protection method based on the minimum mean square error criterion according to claim 1, characterized in that: Based on the first estimation error and the second estimation error, the adaptive boundaries of two classic localized differential privacy techniques are determined, including: Build Function ,but The value at zero point is: in, and The expression is: Will As the basic RAPPOR technique and k -Optimal adaptive boundary of RR technique.

3. The adaptive privacy protection method based on the minimum mean square error criterion according to claim 1, characterized in that: A weighting factor is constructed based on the minimum mean square error. The perturbed data sent by each subgroup at different privacy protection levels is aggregated to obtain a statistical estimate of the local user's private data, including: Construct weighting factors based on the expectation of minimum mean square error: For the The weighting factors of the subgroups satisfy , is a counting symbol, with values ​​ranging from 1 to m , It is Subgroups in the privacy protection level The mean square error value of the estimated distribution under; Based on the constructed weighting factor m The perturbed data in each subgroup is weighted and aggregated to obtain a statistical estimate of the local user's private data: in, For the first The estimated distribution of private data of subgroups, m is the total number of subgroups, It is a statistical estimate of local user privacy data.

4. The adaptive privacy protection method based on the minimum mean square error criterion according to claim 1, characterized in that: The method further includes: expanding the data in the subgroup using a multi-perturbation data expansion strategy to equivalently increase the number of privacy subgroups with high privacy requirements.

5. An adaptive privacy protection device based on the minimum mean square error criterion, characterized in that: include: Privacy protection level receiving module: the data aggregator receives the privacy protection level sent by the local user; The group segmentation module is used to group local users according to their privacy protection level, and divide users with the same privacy protection level into the same subgroup; The adaptive result generation module is used to determine the optimal perturbation probability based on the localized differential privacy technology and the privacy protection level. Based on the minimum mean square error criterion, it determines the adaptive boundaries of the two classic localized differential privacy technologies, selects the best data perturbation method according to the adaptive boundary, and sends the adaptive results to the users in the corresponding subgroup, so that the users in each subgroup use the corresponding optimal data perturbation method to perturb their private data, and use the optimal perturbation probability to perform privacy protection operations, obtain the perturbed data, and send it to the data aggregator. The adaptive results include the optimal data perturbation method and the optimal perturbation probability. The two classic localized differential privacy technologies include basic RAPPOR technology or k -RR technology, which determines the optimal perturbation probability based on localized differential privacy technology and privacy protection level, including: When the localized differential privacy technology used is basic RAPPOR technology, Under the privacy protection level, the optimal perturbation probability for each bit of binary-encoded private data is: in, Privacy protection level; When the localized differential privacy technology used is k -RR technology, Under the privacy protection level, the optimal perturbation probability for each bit of binary-encoded private data is: The above formula represents the privacy data The probability of keeping the original value is The probability of perturbation output other -1 of any kind, is the number of different private data; Based on the minimum mean square error criterion, the adaptive boundaries of two classic localized differential privacy technologies are determined, including: The first estimation error of the privacy distribution when using the basic RAPPOR technique is calculated based on the maximum likelihood estimation criterion: Based on the maximum likelihood estimation criterion, the k -RR technology uses the second estimation error of privacy distribution: in, Indicates the amount of data or the number of users, is the privacy protection level, also known as the privacy budget, For the i private data, private data The true probability is , is the number of different privacy data, is the first estimation error, is the second estimation error; The adaptive boundaries of two classic localized differential privacy techniques are determined based on the first estimation error and the second estimation error; The weighted aggregation module is used to construct a weighting factor based on the minimum mean square error, aggregate the perturbed data sent from each subgroup at different privacy protection levels, and obtain a statistical estimate of the local user's private data.

6. Adaptive privacy protection system based on minimum mean square error criterion, characterized by: The method comprises the adaptive privacy protection device based on the minimum mean square error criterion as described in claim 5 and a local user terminal, wherein the local user terminal is used to send a privacy protection level to a data aggregator, select an optimal perturbation method to perturb the private data according to the adaptive result sent by the data aggregator, and use the optimal perturbation probability to perform a privacy protection operation to obtain the perturbed data, and send the perturbed data to the data aggregator.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed, the method according to any one of claims 1 to 4 is implemented.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Anti-collusion attack data sharing method and device and electronic equipment

    CN114185860A

  • Smart power grid data aggregation method based on local differential privacy

    CN115168423A