Self-adaptive privacy protection mean value estimation method, device and system based on local personalized differential privacy

By integrating IM and NM mechanisms in local differential privacy technology, and adaptively selecting disturbance strategies based on user privacy protection levels, the problem of imbalance between privacy and data availability in traditional technologies is solved, and efficient and accurate mean estimation and data sharing are achieved.

CN120337288APending Publication Date: 2025-07-18HUBEI UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510458812.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional local differential privacy technology has imbalanced privacy and data availability due to the unified privacy protection parameter setting in the mean estimate, which cannot meet the needs of personalized privacy protection, affecting the accuracy of statistical analysis and user data sharing enthusiasm.

Method used

Adaptive privacy protection mean estimation method based on local personalized differential privacy is adopted. By integrating interval-based mean estimation mechanism (IM) and nearest neighbor-based mean estimation mechanism (NM), the optimal perturbation mechanism and strategy are adaptively selected according to the user's privacy protection level, and a weighted merging method is adopted to optimize the data aggregation process.

Benefits of technology

It achieves the improvement of the accuracy and data availability of mean estimates while ensuring data privacy and security, meet personalized privacy protection needs, and enhances users' enthusiasm for participating in data collection and sharing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337288A_ABST
    Figure CN120337288A_ABST
Patent Text Reader

Abstract

The invention discloses an adaptive privacy protection mean value estimation method, device and system based on local personalized differential privacy, deeply integrates the advantages of an interval-based mean value estimation mechanism (IM) and a neighbor-based mean value estimation mechanism (NM), and aims to search and realize the optimal balance between data privacy and availability. By fully considering personalized privacy protection requirements of users, optimal adaptive boundaries of an IM mechanism and an NM mechanism are determined from the perspective of minimizing the maximum mean square error of the worst case, so that a data aggregator can select an optimal disturbance mechanism for the users at a local end in combination with actual situations and privacy protection levels reported by the users; therefore, the optimal perturbation strategy is adopted, personalized privacy protection is realized, and meanwhile, relatively high data effectiveness is ensured. Meanwhile, a weighting factor is constructed based on the maximum mean square error under the worst condition, and the accuracy of mean value estimation is effectively improved by adopting a weighting combination method. In short, the method is a practical method oriented to statistics and analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of privacy data protection, and more specifically, to an adaptive privacy protection mean estimation method, device and system based on local personalized differential privacy. Background Art

[0002] Mean estimation, as a basic statistical tool, has very wide applications and is an indispensable part when analyzing and inferring a data set. In mean estimation, the problem of privacy protection is an increasingly concerned topic, especially when dealing with sensitive data. For example, in fields such as medical care, finance, and social networks, protecting personal privacy is of crucial importance. Traditional mean estimation methods may expose individual information, thus bringing the risk of privacy leakage. Therefore, when performing mean estimation, it becomes particularly important to take appropriate privacy protection measures.

[0003] Among various means of privacy protection mean estimation, Local Differential Privacy (LDP) has unique advantages. Different from traditional differential privacy, local differential privacy does not rely on a central server to access all data, so it has stronger decentralization characteristics and is suitable for scenarios that require distributed data processing, such as mobile devices, Internet of Things devices, etc. The advantage of LDP is that it can effectively protect user privacy while reducing the potential risks brought by centralized storage and processing. However, LDP also faces the challenge that the introduction of noise may lead to a decrease in estimation accuracy, especially in the case of a small amount of data or a small privacy budget. However, the biggest highlight of LDP is that it provides a flexible and efficient solution for protecting personal data privacy while still supporting effective statistical analysis, such as mean estimation.

[0004] When traditional LDP methods are used to estimate the mean of numerical data, they often perform data perturbation at the local end with uniformly set privacy protection parameters. For example, in some large-scale user data collection projects, whether it is medical data, financial data or user behavior data, the same privacy parameters are used to process the data. Although this approach simplifies the operation process to a certain extent, it ignores a key issue, that is, the data distribution characteristics of different users vary greatly, and their privacy preferences are also very different. In an actual scenario, taking medical data as an example, some patients with rare diseases may require a higher level of privacy protection due to the particularity of their data, because once these data are leaked, it may have a serious impact on their lives; while for the general examination data of some common diseases, patients may have relatively lower requirements for privacy. Similarly, in the field of financial data, the financial data of high-net-worth users is usually more sensitive and requires stronger privacy protection, while the privacy protection requirements for some basic financial transaction data of ordinary users are relatively weak.

[0005] Some existing privacy protection technologies, such as the Local Differential Privacy (LDP) protection technology based on Randomized Response (RR), have achieved certain results to some extent. However, most studies have not fully considered the differences in privacy protection needs among different individuals. When the same privacy protection strategy is adopted for all user data, the problem of unbalanced privacy protection will occur: users with high privacy needs may face the risk of privacy leakage due to insufficient protection, while users with low privacy needs may experience a decrease in data availability due to overprotection, affecting the accuracy of statistical analysis. This situation will not only trigger users' concerns about data openness and sharing, hinder the development process of big data, but also greatly reduce the accuracy of statistical estimation and fail to meet the requirements of practical applications. Summary of the Invention

[0006] The purpose of the present invention is to solve the problem of the imbalance between privacy and data availability caused by the unified setting of privacy protection parameters in the mean estimation technology in traditional local differential privacy research when dealing with numerical data. By fully considering the personalized privacy protection of local users, an adaptive privacy protection mean estimation method based on local personalized differential privacy is designed to deeply integrate the advantages of the IM and NM mechanisms and effectively balance data privacy and availability.

[0007] To achieve the above purpose, the first aspect of the present invention provides an adaptive privacy protection mean estimation method based on local personalized differential privacy, including:

[0008] Receiving the privacy protection level of the terminal user, where the privacy protection level of the terminal user is obtained by each terminal user based on the analysis of its own privacy data and privacy needs, and the own privacy data is numerical data;

[0009] According to the privacy protection level of the terminal user, terminal users with the same privacy protection level are divided into the same subgroup;

[0010] Based on minimizing the maximum mean square error of the worst case, the optimal adaptive boundaries of two mechanisms, namely the interval-based mean estimation mechanism and the neighbor-based mean estimation mechanism, are determined. According to the privacy protection level of the local user and the optimal adaptive boundary, the best perturbation mechanism and the corresponding best perturbation strategy are recommended for each subgroup of users; so that the users in each subgroup use the corresponding best perturbation mechanism and the corresponding best perturbation strategy to perturb their privacy data to obtain perturbed data and send it to the data aggregator;

[0011] Aggregating the perturbed data to obtain a mean estimation result.

[0012] In one implementation, the optimal adaptive boundary between the interval-based mean estimation mechanism and the neighbor-based mean estimation mechanism is determined based on minimizing the maximum mean squared error in the worst case, including:

[0013] Obtain the maximum mean squared error in the worst case of the interval-based mean estimation mechanism;

[0014] Obtain the maximum mean squared error in the worst case of the neighbor-based mean estimation mechanism;

[0015] By minimizing the maximum mean squared error in the worst case under the two mechanisms, determine the optimal adaptive boundary between the interval-based mean estimation mechanism and the neighbor-based mean estimation mechanism.

[0016] In one implementation, obtaining the maximum mean squared error in the worst case of the interval-based mean estimation mechanism includes:

[0017] Calculate the variance of the perturbed data under the interval-based mean estimation mechanism;

[0018] Calculate the expectation of the perturbed data under the interval-based mean estimation mechanism;

[0019] According to the variance and expectation of the perturbed data under the interval-based mean estimation mechanism, calculate the mean squared error under the interval-based mean estimation mechanism;

[0020] Based on the mean squared error under the interval-based mean estimation mechanism, obtain the maximum mean squared error in the worst case of the interval-based mean estimation mechanism.

[0021] In one implementation, obtaining the maximum mean squared error in the worst case of the neighbor-based mean estimation mechanism includes:

[0022] Calculate the variance of the perturbed data under the neighbor-based mean estimation mechanism;

[0023] Calculate the expectation of the perturbed data under the neighbor-based mean estimation mechanism;

[0024] According to the variance and expectation of the perturbed data under the neighbor-based mean estimation mechanism, calculate the mean squared error under the neighbor-based mean estimation mechanism;

[0025] Based on the mean squared error under the neighbor-based mean estimation mechanism, obtain the maximum mean squared error in the worst case of the neighbor-based mean estimation mechanism.

[0026] In one implementation, by minimizing the maximum mean squared error in the worst case under the two mechanisms, determining the optimal adaptive boundary between the interval-based mean estimation mechanism and the neighbor-based mean estimation mechanism includes:

[0027] Construct a constructor for the difference between the maximum mean square errors in the worst cases under two mechanisms;

[0028] Use the bisection method to find the zero point to determine the adaptive boundary point, and obtain the optimal adaptive boundary according to the adaptive boundary point.

[0029] In one implementation, using the bisection method to find the zero point to determine the adaptive boundary point, and obtaining the optimal adaptive boundary according to the adaptive boundary point includes:

[0030] Initialize the sub-interval according to the constructor;

[0031] Judge the position of the zero point according to the sign of the value obtained by substituting the initial boundary point into the constructor, and update the sub-interval;

[0032] Judge whether the preset iteration condition is satisfied. If it is satisfied, stop. If it is not satisfied, continue to update the sub-interval. Among them, when the iteration process ends, the midpoint of the latest sub-interval containing the zero point is used as the value of the optimal adaptive boundary.

[0033] In one implementation, aggregating the perturbed data to obtain a mean estimation result includes:

[0034] Each sub-group of users separately performs data aggregation and analysis based on the data perturbed by the users in the group, and estimates the mean value under each privacy protection level as the aggregation result of each sub-group;

[0035] Weightedly combine the aggregation results from sub-groups under different privacy protection levels, where the weighted combination factor is constructed based on minimizing the maximum mean square error in the worst case.

[0036] Based on the same inventive concept, the second aspect of the present invention provides an adaptive privacy protection mean estimation device based on local personalized differential privacy. The device is a data aggregator, including:

[0037] A privacy protection level receiving module, configured to receive the privacy protection level of the terminal user. Among them, the privacy protection level of the terminal user is obtained by each terminal user based on the analysis of its own privacy data and privacy requirements, and the own privacy data is numerical data;

[0038] A sub-group division module, configured to divide the terminal users with the same privacy protection level into the same sub-group according to the privacy protection level of the terminal user;

[0039] The optimal perturbation mechanism and optimal perturbation strategy determination module is used to determine the optimal adaptive boundaries of the interval-based mean estimation mechanism and the neighbor-based mean estimation mechanism based on minimizing the maximum mean square error of the worst case. According to the privacy protection level of the local users and the optimal adaptive boundaries, it recommends the optimal perturbation mechanism and the corresponding optimal perturbation strategy for each subgroup of users; so that the users in each subgroup use the corresponding optimal perturbation mechanism and the corresponding optimal perturbation strategy to perturb their privacy data, obtain the perturbed data, and send it to the data aggregator;

[0040] The aggregation module is used to aggregate the perturbed data to obtain the mean estimation result.

[0041] Based on the same inventive concept, the third aspect of the present invention provides an adaptive privacy protection mean estimation system based on local personalized differential privacy, including the adaptive privacy protection mean estimation device described in the second aspect and the end user. Among them, the end user is used to analyze the privacy protection level based on its own privacy data and privacy needs and send it to the data aggregator, and perturb its privacy data according to the optimal perturbation mechanism and the corresponding optimal perturbation strategy recommended by the data aggregator to obtain the perturbed data, and send it to the data aggregator.

[0042] Based on the same inventive concept, the fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the adaptive privacy protection mean estimation method described in the first aspect.

[0043] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows:

[0044] The present invention provides an adaptive privacy-preserving mean estimation method based on local personalized differential privacy, which deeply integrates the advantages of the interval-based mechanism for mean estimation (IM) and the neighbor-based mechanism for mean estimation (NM), aiming to find and achieve the best balance between data privacy and availability. By fully considering the personalized privacy protection needs of users, from the perspective of minimizing the maximum mean square error in the worst case, the optimal adaptive boundaries of the IM and NM mechanisms are determined, enabling the data aggregator to select the best perturbation mechanism for the users at the local end according to the actual situation and the privacy protection level reported by the users, and then adopting the best perturbation strategy to achieve personalized privacy protection while ensuring high data utility. In addition, a weighting factor is constructed based on the maximum mean square error in the worst case, and a weighted merging method is used to effectively improve the accuracy of mean estimation. In short, this method is a practical method for statistics and analysis, providing a highly innovative and practical solution for the secure processing and efficient application of data, and having strong practical significance. Description of the Drawings

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0046] Figure 1 It is a flowchart of the adaptive privacy-preserving mean estimation method based on local personalized differential privacy in the embodiments of the present invention;

[0047] Figure 2 It is a data collection framework diagram of the adaptive privacy-preserving mean estimation based on local personalized differential privacy in the embodiments of the present invention;

[0048] Figure 3 It is a data collection framework diagram of the adaptive personalized privacy protection divided into three sub-groups in the embodiments of the present invention;

[0049] Figure 4 It is a structural diagram of the adaptive privacy-preserving mean estimation device based on local personalized differential privacy in the embodiments of the present invention. Detailed Embodiments

[0050] Through a large number of studies and practices, the inventors of this application have found that privacy-preserving mean estimation can ensure that individuals' sensitive data is not leaked or misused during statistical analysis. In the traditional local differential privacy technology for mean estimation of numerical data, the common practice is to set a unified privacy protection parameter and apply it to the data perturbation link of all participants. However, in reality, there are significant differences in the privacy preferences of different users. This "one-size-fits-all" approach is difficult to achieve ideal results in terms of protecting privacy and ensuring data availability. Moreover, the interval-based mechanism for mean estimation (IM) has good statistical estimation accuracy for mean estimation in the low privacy range, but poor statistical estimation accuracy in the high privacy range; while the neighbor-based mechanism for mean estimation (NM) has good statistical estimation accuracy for mean estimation in the high privacy range, but poor statistical estimation accuracy in the low privacy range. The present invention hopes to break through the limitations of traditional methods and propose a privacy-preserving mean estimation method based on local personalized differential privacy. This method is applicable to the mean estimation of numerical data, deeply integrates the advantages of the interval-based mechanism for mean estimation (IM) and the neighbor-based mechanism for mean estimation (NM), and aims to find and achieve the best balance between data privacy and availability. The core innovation of the present invention lies in taking into account the personalized privacy protection needs of users, deeply integrating the advantages of the interval-based mechanism for mean estimation (IM) and the neighbor-based mechanism for mean estimation (NM), and realizing the precise and efficient protection and utilization of data. The present invention adaptively selects the IM or NM mechanism according to the personalized privacy protection needs of users at the local end. This adaption includes adaptively selecting the best perturbation mechanism (IM or NM) for data perturbation and adaptively selecting the best perturbation strategy to output perturbed data. This method not only realizes personalized privacy protection, but also can obtain higher data utility through weighted combination. Among them, the maximum mean square error in the worst case of the two mechanisms is theoretically deduced and analyzed. On this basis, with the goal of minimizing the maximum mean square error (MSE) in the worst case, the optimal adaptive boundary of the IM mechanism and the NM mechanism is determined.

[0051] In practical application scenarios, the data aggregator can adaptively select the mechanism that best suits the current privacy protection requirements from the IM and NM mechanisms as the optimal data perturbation mechanism according to the above-defined adaptive boundaries. Furthermore, according to the unique privacy preferences of the data providers, the best perturbation strategy under the corresponding mechanism is adaptively determined, so as to achieve efficient perturbation output of the data, while ensuring data privacy and security and maximizing the accuracy of mean estimation. In addition, to ensure the high availability of the aggregated data, the present invention uses a weighted merging method to effectively merge the data from different privacy groups, and constructs the weighting coefficient using the maximum mean square error in the worst case. In summary, this method is a practical method for statistics and analysis, with strong practical significance.

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0053] Embodiment 1

[0054] This embodiment discloses an adaptive privacy protection mean estimation method based on local personalized differential privacy. Please refer to Figure 1 , including:

[0055] S1: Receive the privacy protection level of the terminal user. Among them, the privacy protection level of the terminal user is obtained by each terminal user based on the analysis of its own privacy data and privacy requirements, and the own privacy data is numerical data.

[0056] Specifically, in this embodiment, the executing entity is the data aggregator or collector. The terminal user first deeply analyzes its own privacy data, combines the actual privacy requirements, clarifies the personalized privacy protection level, and sends this information accurately to the data collector or aggregator.

[0057] S2: According to the privacy protection level of the terminal user, divide the terminal users with the same privacy protection level into the same sub-group.

[0058] In this step, after receiving the user privacy level information, the data collector or aggregator classifies the users with the same privacy protection level into the same sub-group.

[0059] S3: Determine the optimal adaptive boundaries of the interval-based mean estimation mechanism and the neighbor-based mean estimation mechanism based on minimizing the maximum mean squared error in the worst case. According to the privacy protection level of local users and the optimal adaptive boundaries, recommend the optimal perturbation mechanism and the corresponding optimal perturbation strategy for each subgroup of users; so that the users in each subgroup use the corresponding optimal perturbation mechanism and the corresponding optimal perturbation strategy to perturb their privacy data, obtain the perturbed data, and send it to the data aggregator.

[0060] Specifically, S3 includes the data aggregator determining the optimal perturbation mechanism and the optimal perturbation strategy and the users executing the optimal perturbation mechanism and the optimal perturbation strategy. First, based on the privacy protection levels reported by the users, the data aggregator derives the optimal adaptive boundaries based on minimizing the maximum mean squared error in the worst case, accurately recommends the optimal perturbation mechanism (IM or NM) and the corresponding optimal perturbation strategy for each subgroup of users, and sends the perturbation strategy parameters under the corresponding perturbation mechanism to the end users.

[0061] The users in each subgroup of the terminal strictly follow the optimal perturbation mechanism and the optimal perturbation strategy provided by the data aggregator to perturb the local data and send the perturbed data to the data aggregator.

[0062] S4: Aggregate the perturbed data to obtain the mean estimation result.

[0063] Specifically, the aggregation of the perturbed data includes the aggregation of the data of each group of users within each subgroup and the aggregation of the aggregation results of each subgroup among the subgroups.

[0064] In the localized data collection process of the present invention, the participants adaptively select appropriate privacy protection mechanisms and perturbation strategies according to their personalized privacy needs. Among them, the privacy protection needs are measured by the privacy budget parameter: the smaller the privacy budget, the higher the privacy protection intensity, and vice versa. According to the privacy protection requirement parameters, the end users are divided into different subgroups, and the privacy protection levels of the users within the same subgroup are the same. For the privacy protection needs of each subgroup, by means of rigorous theoretical analysis and precise mathematical derivation, the maximum mean squared error in the worst case is given, and the optimal adaptive boundaries of the IM and NM mechanisms are determined based on the criterion of minimizing the maximum mean squared error in the worst case. On this basis, the local participants can adaptively select the most suitable perturbation mechanism from the IM and NM according to their own privacy preferences, and adaptively determine the corresponding optimal perturbation strategy, so as to perform accurate perturbation operations on the privacy data and submit the processed data to the data aggregator for aggregation and weighted combination.

[0065] In the specific implementation process of the present invention, the optimal adaptive boundary in the adaptive perturbation mechanism is derived based on minimizing the maximum mean square error in the worst case, which can make full use of the advantages of the IM and NM mechanisms and effectively improve the accuracy of mean estimation.

[0066] In one implementation, aggregating the perturbed data to obtain a mean estimation result includes:

[0067] Each sub-group of users separately performs data aggregation and analysis based on the data perturbed by the users in the group, estimates the mean value at each privacy protection level, and takes it as the aggregation result of each sub-group;

[0068] The aggregation results from sub-groups under different privacy protection levels are weighted and combined, where the weighted combination factor is constructed based on minimizing the maximum mean square error in the worst case.

[0069] Specifically, the weighted combination method in this implementation is constructed based on minimizing the maximum mean square error in the worst case, fully considering the contribution of each sub-group to the accuracy of the mean estimation of the entire data domain, and further improving the accuracy of mean estimation. Compared with the traditional direct combination method, this weighted combination can deeply explore the data value and significantly improve the accuracy of mean estimation.

[0070] The existing patent application CN115879152A discloses an adaptive privacy protection method based on the minimum mean square error criterion (hereinafter referred to as the comparative document). Compared with the comparative document, the present application mainly has the following differences: (1) The data objects processed are different. The comparative document 1 is a privacy protection method for frequency estimation, aiming at the privacy protection frequency estimation of categorical data, while the present invention is a privacy protection method for mean estimation, aiming at the privacy protection mean estimation of numerical data; (2) The data processing mechanisms and the acquisition methods of the adaptive boundaries adopted are different. The comparative document 1 is for two frequency estimation algorithms, k-RR and Basic RAPPOR, that satisfy ∈-local differential privacy, and solves the adaptive boundary according to the minimum mean square error criterion, while the present invention is for two mean estimation algorithms, the interval-based mean estimation mechanism (IM) and the neighbor-based mean estimation mechanism (NM), that satisfy (∈,δ)-local differential privacy, and solves the adaptive boundary according to the minimization of the maximum mean square error in the worst case; (3) The solution methods of the adaptive boundaries are different. The adaptive boundary of the comparative document 1 can obtain an exact expression through theoretical derivation, while the adaptive boundary of the present invention cannot obtain an exact expression based on theory, so the present invention uses the bisection method to determine the adaptive boundary; (4) The data aggregation methods are different, mainly in the construction method of the weighting factor. The weighting factor in the weighted merging method of the comparative document 1 is constructed based on the minimum mean square error, while the weighted merging method of the present invention is constructed based on the maximum mean square error in the worst case.

[0071] In summary, the present invention has the following remarkable advantages:

[0072] 1. It deeply meets the personalized privacy needs of local users. While providing precise personalized privacy protection for users, it effectively stimulates users' enthusiasm for actively participating in data collection and sharing, and significantly improves users' participation and initiative.

[0073] 2. Relying on strong theoretical support, the optimal adaptive boundaries of the IM and NM mechanisms derived based on the minimization of the maximum mean square error in the worst case provide a scientific and reliable decision-making basis for participants, enabling them to flexibly choose the best perturbation mechanism, ensuring the balance between privacy protection and data availability, and improving the accuracy of mean estimation.

[0074] 3. The weighted merging method constructed based on the minimization of the maximum mean square error in the worst case cleverly utilizes the weight information of the maximum mean square error in the worst case of each sub-group, fully considers the contribution degree of each sub-group to the accuracy of mean estimation for weighting, and gives full play to its advantages in data merging. Compared with the traditional direct merging method, it can further improve the accuracy and stability of mean estimation.

[0075] 4. The proposed adaptive privacy-preserving mean estimation method of the present invention can not only achieve personalized privacy protection, but also significantly improve the data availability and statistical accuracy by introducing a weighted merging method. It is a statistical and analytical method with great practical value and has broad application prospects and important practical significance in actual application scenarios.

[0076] The following is an analysis of the method of the present invention in conjunction with the accompanying drawings.

[0077] Given a data collector and N users {u1, u2, ···, u N}, assuming that the users are independent of each other, and each user has a numerical data x i ∈[-1, 1]. The data collector wants to obtain the mean of the data with relatively high accuracy (Note: In practice, the value range of numerical data may not be the interval [-1, 1]. As long as it is mapped to the interval [-1, 1] by a standardization method, the method of the present invention is applicable to the mean estimation of any value range). However, directly sending the data to the collector will cause user privacy leakage. The present invention considers that u i perturbs the real data using the IM or NM mechanism on the local client, and the perturbed data is denoted as y i . The present invention uses the Mean Square Error (MSE) to evaluate the estimation error of each perturbation mechanism. The smaller the MSE, the smaller the estimation error and the higher the estimation accuracy. The calculation formula of MSE is as follows:

[0078] MSE[y i = Var[y i + (E[y i - x i ) 2

[0079] where, Var[y i represents the variance of the perturbed data, E[y i represents the expected value of the perturbed data, and E[y i - x i represents the difference between the mean of the perturbed data and the unperturbed data.

[0080] Theoretical derivation of the optimal adaptive boundary based on the criterion of minimizing the worst-case MSE

[0081] (1) Mean estimation error under the IM mechanism

[0082] When the local side uses the interval-based mean estimation mechanism (IM) to design the submission of perturbed data that satisfies (∈, δ)-local differential privacy protection, where δ is the relaxation factor and the privacy protection level is measured by the differential privacy budget ∈. Assume the input data of the user is x i ∈[-1, 1], and the perturbed output is y i , and y i ∈[-C, C], where [-C, C] is the output domain and C is a function of the privacy protection parameters (∈, δ). Once the privacy protection parameters (∈, δ) are given, C is a specific constant.

[0083] The user divides the output domain into three intervals: the left interval, the middle interval, and the right interval. The probability that the data x i is perturbed into the three intervals is different. The two boundaries in the output domain depend on the value of the real data x i . The middle interval is closer to the real data. The left interval is defined as [-C, ax i +b), the middle interval is defined as [ax i +b, ax i -b], and the right interval is [ax i -b, C], where a and b are functions of the privacy protection parameters (∈, δ). Once the privacy protection parameters (∈, δ) are given, a and b are specific constants. The specific functional expressions of C, a, and b with respect to ∈ and δ under the optimal perturbation strategy will be given below.

[0084] To obtain a more accurate unbiased estimation result, the perturbed output data is designed to fall in the middle interval with a high probability, that is, the probability that the perturbed data falls in the middle interval [ax i +b, ax i -b] is -2bp, that is, the probability that the perturbed output data falls on each point in the middle interval [ax i +b, ax i -b] is p; similar to the above perturbation, the data is perturbed into the left interval and the right interval with probability q - to any value in [-C, ax i +b) ∪ (ax i -b, C].

[0085] The privacy protection level is measured by the differential privacy budget ∈. Design a perturbation mechanism that satisfies (∈, δ)-local differential privacy, where δ is the relaxation factor. When the privacy protection level is ∈ and the relaxation factor is δ, for the numerical data x i ∈[-1, 1], the best perturbation strategy is:

[0086]

[0087] where, pdf[] represents the probability density function, t is a value in a specific output domain, p and q are perturbation probability values, l(x i ) is the left boundary value in the output domain, which is a function of the real data x i ; r(x i ) is the right boundary value in the output domain, which is also a function of the real data x i . The settings of each parameter under the optimal strategy are as follows:

[0088]

[0089] p = e ∈ q+δ

[0090]

[0091] b = a - C

[0092] l(x i ) = ax i +b

[0093] r(x i ) = ax i -b

[0094] where, q, p, a, C, b, l(x i ), r(x i ) are the perturbation strategy parameters under the IM mechanism.

[0095] Each end user executes the above perturbation strategy and sends the perturbed data to the data collector for aggregation. Its estimated mean is

[0096] According to the definition formula of variance, as well as the corresponding perturbation interval and probability, the variance of the perturbed data y i is described as:

[0097]

[0098] According to the definition of expectation, the expectation of the perturbed data y i is described as:

[0099]

[0100] It can be seen that the mean estimated by the interval-based mean estimation mechanism satisfies the unbiased property. Therefore, according to the definition formula of MSE, the MSE under the IM mechanism is as follows:

[0101]

[0102] The MSE of its estimated mean is described as follows:

[0103]

[0104] It can be seen that the estimation error of the IM mechanism is related to the original true data, and the maximum mean square error in the worst case is when x i = ±1, and the maximum mean square error in the worst case is described as

[0105]

[0106] (2) Mean Estimation Error under the NM Mechanism

[0107] When the local end adopts the neighbor-based mean estimation mechanism (NM) to design the submission of perturbed data that satisfies (∈,δ)-local differential privacy protection, where δ is the relaxation factor, and the privacy protection level is measured by the differential privacy budget ∈. Assume that the input data of the user is x i ∈[-1,1], and the perturbed output is y i , and y i ∈[-h,h + 1], where [-h,h + 1] is the output domain, and h is a function of the privacy protection parameters (∈,δ). Once the privacy protection parameters (∈,δ) are given, h is a specific constant. The user first needs to map the true data x i to x i ′ ∈[0,1], and then perform perturbation. The purpose of this mapping is to better solve for the optimal neighborhood length h.

[0108] For the perturbation process, the user divides the output domain into 3 intervals: the left interval, the middle interval, and the right interval. The true data x i is perturbed to the three intervals with different probabilities. Since each x i is different, for each x i , its left and right intervals and the middle interval are all unique. And for the numerical data x i at the local end, in order to obtain a more accurate estimation result, the perturbed output data is designed to fall in the middle interval with a greater probability, and the probability of the perturbed data falling at each point in the middle interval is set to p; similar to the above perturbation, the original data is perturbed to any point value in the left and right intervals with probability q.

[0109] The privacy protection level is measured by the differential privacy budget ∈, and a perturbation mechanism that satisfies (∈,δ)-local differential privacy is designed, where δ is the relaxation factor. When the privacy protection level is ∈ and the relaxation factor is δ, for the numerical data x i ∈[-1,1], the data x i is formatted to x i ′ =(xi +1) / 2(x i ′ ∈[0,1]), and then design the optimal perturbation strategy under the NM mechanism as follows:

[0110]

[0111] Among them, p' and q' are perturbation probabilities, and are set as follows:

[0112]

[0113] Among them, q', p', and h are perturbation strategy parameters under the NM mechanism.

[0114] According to the definition formula of the mean, as well as the corresponding perturbation interval and probability, the expectation of the perturbed data y i is:

[0115] E[y i = h(p' - q')(x i +1) + q'(2h + 1) / 2

[0116] According to the definition formula of the variance, the variance of the perturbed data y i is described as

[0117]

[0118] According to the definition formula of the MSE, the MSE of the perturbed data y i is:

[0119]

[0120] The MSE of its estimated mean is described as follows:

[0121]

[0122] It can be seen that the NM algorithm can achieve the maximum mean square error of the estimated mean in the worst case when x i = -1:

[0123]

[0124] (3) Optimal adaptive boundary derivation

[0125] Currently, the IM mechanism has good statistical estimation accuracy for mean estimation in the low privacy range, but poor statistical estimation accuracy in the high privacy range; the NM mechanism has good statistical estimation accuracy for mean estimation in the high privacy range, but poor statistical estimation accuracy in the low privacy range. In view of this, the present invention proposes an adaptive privacy protection mean estimation method based on local personalized differential privacy. The core innovation lies in taking into account the personalized privacy protection needs of users, deeply integrating the advantages of the IM mechanism and the NM mechanism, and realizing precise and efficient protection and utilization of data. The following derives the optimal adaptive boundary based on the MSE criterion of minimizing the worst case.

[0126] Now construct a function:

[0127]

[0128] Where, is the maximum mean square error in the worst case under the IM mechanism, is the maximum mean square error in the worst case under the NM mechanism, and ΔMSE(∈) is a function of the privacy protection parameter ∈.

[0129] Since the expressions of the maximum mean square error in the worst case of the estimated means under the IM and NM mechanisms are very complex, it is extremely difficult to directly find the zero point (i.e., the optimal adaptive boundary point) of the constructed function ΔMSE(∈). This embodiment decides to use the bisection method to obtain the optimal adaptive boundary value. Among them, the bisection method is a numerical method for solving the roots (zero points) of equations. Its basic idea is: for a function that is continuous and has different signs on an interval, by continuously dividing the interval into two halves, and based on the signs of the function values at the interval endpoints, gradually narrow the interval where the root is located. Repeat this process until the length of the interval is less than the pre-set accuracy requirement. At this time, the midpoint of the interval can be used as an approximate value of the root of the equation. This midpoint of the interval should be more precise than the actually given value (generally on the order of 10 -4 order of magnitude). Therefore, in the present invention, an interval obtained by the bisection method (this interval is less than 10 -4 order of magnitude) can be applied in practice.

[0130] At the same time, in the existing situation, the order of magnitude of the relaxation factor δ is generally very small, generally below 10 -6 . Therefore, its impact on the optimal adaptive boundary of the privacy budget of the IM and NM perturbation mechanisms is very small. And since both sides of the expressions of the maximum mean square error in the worst case of the estimated means of the IM and NM mechanisms have a factor of 1 / N, when solving the intersection point of the two, 1 / N can be eliminated. Assume that the optimal adaptive boundary point solved according to the bisection method is ∈ * , then the principle for designing the optimal perturbation mechanism is:

[0131] When ∈≥∈ * there is indicating that the estimated accuracy of the privacy data obtained by the local differential privacy technology using the IM technology is higher than that using the NM technology under the same conditions, and the IM technology is adaptively selected as the perturbation method;

[0132] When ∈<∈ * there is indicating that the estimated accuracy of the privacy data obtained by the local differential privacy technology using the NM technology is higher than that using the IM technology under the same conditions, and the NM technology is adaptively selected as the perturbation method.

[0133] The basic solution process of the optimal adaptive boundary is as follows:

[0134] (1) Initialize the sub-interval: By constructing the function ΔMSE(∈), simulate the curve graph of ΔMSE(∈) with respect to the privacy budget ∈, and observe that the boundary point ∈ * roughly belongs to the interval [∈ L , ∈ R . Calculate the corresponding ΔMSE(∈) values under the privacy parameters ∈ L and ∈ R respectively: ΔMSE(∈ L ), ΔMSE(∈ R ). If ΔMSE(∈ L ) > 0 and ΔMSE(∈ R ) < 0, then the zero point is in the interval [∈ L , ∈ R . Take the midpoint ∈ L = (∈ R + ∈ * ) / 2 of the interval [∈ L , ∈ R .

[0135] (2) Update the sub-interval. Determine the position of the zero point according to the sign of ΔMSE(∈ * ), and update the sub-interval: If ΔMSE(∈ * ) > 0, then the zero point is in [∈ * , ∈ R ; if ΔMSE(∈ * ) < 0, then the zero point is in [∈ L , ∈ * . Take the midpoint of the interval containing the zero point and denote it as the new sub-interval [∈ L , ∈ R .

[0136] (3) Iteration termination. Define Δ∈ = ∈ L - ∈ R , when Δ∈ ≤ 10 -4, the iteration stops; otherwise, the iteration continues.

[0137] (4) Iterative update: If the value of Δ∈ does not satisfy Δ∈ ≤ 10 -4 , continue to execute Step 2 and Step (3), and take the midpoint of the latest sub-interval containing the zero point after the iteration process ends as the value of the optimal adaptive boundary.

[0138] A special case is given below to illustrate the solution of the adaptive boundary, which is called Embodiment 1.

[0139] Set N = 10000 and δ = 10 -6 , and observe the boundary point ∈ through the simulation curve graph * is roughly in the interval [1.875, 2]. Then the process of solving the optimal adaptive boundary point by the bisection method is as follows:

[0140] (a) When ∈ = 2, ΔMSE(∈) < 0; when ∈ = 1.875, ΔMSE(∈) > 0. Therefore, it can be determined that the zero point of ΔMSE(∈) is between [1.875, 2], and the midpoint of the interval 1.9375 is taken. When ∈ = 1.9375, ΔMSE(∈) < 0. Therefore, it can be determined that the zero point of ΔMSE(∈) is between [1.875, 1.9375], and the iterative cutoff condition of Δ∈ ≤ 10 -4 is not satisfied, and the midpoint of the interval 1.90625 is taken.

[0141] (b) When ∈ = 1.90625, ΔMSE(∈) > 0. Therefore, it can be determined that the zero point of ΔMSE(∈) is between [1.90625, 1.9375], and the iterative cutoff condition of Δ∈ ≤ 10 -4 is not satisfied, and the midpoint of the interval 1.921875 is taken.

[0142] (c) When ∈ = 1.921875, ΔMSE(∈) < 0. Therefore, it can be determined that the zero point of ΔMSE(∈) is between [1.90625, 1.921875], and the iterative cutoff condition of Δ∈ ≤ 10 -4 is not satisfied, and the midpoint of the interval 1.9140625 is taken.

[0143] (d) When ∈ = 1.9140625, then ΔMSE(∈) < 0. Therefore, it can be determined that the zero point of ΔMSE(∈) is between [1.90625, 1.9140625], and the iterative cutoff condition of Δ∈ ≤ 10 -4 is not satisfied, and the midpoint of the interval 1.91015625 is taken.

[0144] (e) When ∈ = 1.91015625, ΔMSE(∈) > 0. Therefore, it can be determined that the zero point of ΔMSE(∈) is between [1.91015625, 1.9140625], and the condition of Δ∈ ≤ 10 is not satisfied. -4 The iterative cutoff condition is not met, and the midpoint of the interval 1.912109375 is taken.

[0145] (f) When ∈ = 1.912109375, ΔMSE(∈) < 0. Therefore, it can be determined that the zero point of ΔMSE(∈) is between [1.91015625, 1.912109375], and the condition of Δ∈ ≤ 10 is not satisfied. -4 The iterative cutoff condition is not met, and the midpoint of the interval 1.9111328125 is taken.

[0146] (g) When ∈ = 1.9111328125, ΔMSE(∈) < 0. Therefore, it can be determined that the zero point of ΔMSE(∈) is between [1.91015625, 1.9111328125]. At this time, Δ∈ ≤ 10 -4 , meeting the condition for stopping iteration. Therefore, the iteration is stopped, and the midpoint of this interval is used as the final optimal adaptive boundary ∈ * = 1.91064453125.

[0147] Under this implementation, without affecting the selection of the actual solution, the optimal adaptive boundary ∈ of the privacy budget of the IM and NM perturbation mechanisms * can be set to 1.9106. At this time, the adaptive selection of the perturbation mechanism is as follows:

[0148] · When ∈ ≥ 1.9106, adaptively select the IM technology as the perturbation mechanism;

[0149] · When ∈ < 1.9106, adaptively select the NM technology as the perturbation mechanism.

[0150] The following describes the adaptive personalized privacy protection data collection protocol and data aggregation statistical analysis of the present invention in conjunction with the accompanying drawings.

[0151] Design of Adaptive Personalized Privacy Protection Data Collection Protocol

[0152] To balance privacy and usability and obtain high data utility while meeting personalized privacy protection, the local end uses the idea of an adaptive privacy protection algorithm for personalized data perturbation. As Figure 2 shown: Assume that the local end user has m levels of personalized privacy protection, which are ∈1, ∈2, ···, ∈ m , that is, it can be divided into m sub - groups G1, G2, ···, G m , and let the number of users in the sub - groups be n1, n2, ···, nm , the total sample size is where m ≥ 2 and is an integer.

[0153] The optimal perturbation strategy that satisfies (∈ τ , δ)-local differential privacy under the IM mechanism. When ∈ τ ≥ ∈ * , design the optimal perturbation strategy that satisfies (∈ τ , δ)-local differential privacy, where δ is the relaxation factor. When the privacy protection level is ∈ τ , and the relaxation factor is δ, for the numerical data x i ∈[-1, 1], the optimal IM perturbation strategy is:

[0154]

[0155] where, pdf[] represents the probability density function, t is the value in the specific output domain, p and q are the perturbation probability values, l(x i ) is the left boundary value in the output domain, which is a function of the real data x i , r(x i ) is the right boundary value in the output domain, which is also a function of the real data x i . The settings of other parameters under the optimal strategy are as follows:

[0156]

[0157] q τ = e ∈ p τ + δ

[0158]

[0159] b = a - C

[0160] l(x i ) = ax i + b

[0161] r(x i ) = ax i - b

[0162] The optimal perturbation strategy that satisfies (∈ τ , δ)-local differential privacy under the NM mechanism. When ∈ τ <∈ * , design the optimal perturbation strategy that satisfies (∈ τ , δ)-local differential privacy, where δ is the relaxation factor. When the privacy protection level is ∈ τ , and the relaxation factor is δ, for the numerical data x i ∈[-1, 1], the data xi Mapped to x i ′ = (x i + 1) / 2(x i ′ ∈ [0, 1]), and then design the optimal perturbation strategy under the NM mechanism as follows:

[0163]

[0164] where p and q are perturbation probabilities, set as follows:

[0165]

[0166] The implementation steps of the adaptive personalized privacy protection data collection protocol are as follows:

[0167] (1) Determine the privacy protection level. The local user first deeply analyzes its own data, combines the actual privacy requirements, clarifies the personalized privacy protection level, and sends this information accurately to the data collector or aggregator.

[0168] (2) Divide sub - groups. After receiving the user privacy level information, the data collector or aggregator classifies users with the same privacy protection level into the same sub - group.

[0169] (3) Determine the optimal perturbation mechanism and optimal perturbation strategy. Based on the privacy protection level reported by users in step 1, the data aggregator, according to the optimal adaptive boundary derived from minimizing the maximum mean square error in the worst - case scenario, accurately recommends the optimal perturbation mechanism (IM or NM) and the corresponding optimal perturbation strategy for each sub - group of users, and sends the perturbation strategy parameters under the corresponding perturbation mechanism to the end - user.

[0170] (4) Execute the optimal perturbation mechanism and optimal perturbation strategy. Users in each sub - group of the terminal strictly follow the optimal perturbation mechanism and optimal perturbation strategy provided by the data aggregator to perturb their local data, and send the perturbed data to the data aggregator.

[0171] (5) Sub - group data aggregation and statistical analysis. Each sub - group conducts data aggregation and analysis separately to estimate the mean value under each privacy protection level.

[0172] (6) Effective weighted merging. The data aggregator performs weighted merging on the aggregation results from sub - groups under different privacy protection levels, where the weighted merging factor is constructed based on minimizing the maximum mean square error in the worst - case scenario to obtain a high - precision mean value estimation result.

[0173] High - availability data aggregation statistical analysis

[0174] Considering the contribution degree of different privacy groups to the accuracy of mean estimation, a weighted aggregation method is designed based on MSE to improve the accuracy of mean estimation.

[0175] For the mean estimation of the privacy data x by m sub-groups i : Adopt weighted aggregation as shown in Figure 2 to obtain the final mean estimation result of the privacy data:

[0176]

[0177] where, w τ is the weighting factor for described as

[0178]

[0179] where, w τ takes values in the range of 0 to 1 and satisfies MSE τ is the maximum mean square error value of the worst-case mean estimation of the τ-th sub-group under the privacy protection level ∈ τ .

[0180] Please refer to Figure 3 , which is the adaptive personalized privacy protection data collection framework diagram divided into three sub-groups in the embodiment of the present invention, where ∈1 = 1.0, ∈2 = 1.5, ∈3 = 2.0, and N = 10000, δ = 10 -6 . It can be seen from Case 1 that the optimal adaptive boundary ∈ * = 1.9106, and the privacy protection levels of Sub-group 1 and Sub-group 2 satisfy: ∈1, ∈2 < ∈ * , so the NM technology is adaptively selected as the perturbation mechanism, and the (∈1, δ)-local differential privacy and (∈1, δ)-local differential privacy under the NM mechanism are respectively adopted as the optimal perturbation strategies. While the privacy protection level of Sub-group 3 satisfies: ∈3 > ∈ * , so the optimal perturbation strategy that satisfies (∈3, δ)-local differential privacy under the IM mechanism is adopted.

[0181] Embodiment 2

[0182] Based on the same inventive concept, this embodiment discloses an adaptive privacy protection mean estimation method and device based on local personalized differential privacy. Please refer to Figure 4 , the device is a data aggregator, including:

[0183] The privacy protection level receiving module 401 is configured to receive the privacy protection level of the end user. The privacy protection level of the end user is obtained by each end user based on the analysis of their own privacy data and privacy requirements, and the own privacy data is numerical data;

[0184] The sub-group division module 402 is configured to divide the end users with the same privacy protection level into the same sub-group according to the privacy protection level of the end user;

[0185] The optimal perturbation mechanism and optimal perturbation strategy determination module 403 is configured to determine the optimal adaptive boundaries of the interval-based mean estimation mechanism and the nearest neighbor-based mean estimation mechanism based on minimizing the maximum mean square error of the worst case. According to the privacy protection level of the local end user and the optimal adaptive boundary, recommend the optimal perturbation mechanism and the corresponding optimal perturbation strategy for each sub-group of users; so that the users in each sub-group use the corresponding optimal perturbation mechanism and the corresponding optimal perturbation strategy to perturb their privacy data to obtain perturbed data and send it to the data aggregator;

[0186] The aggregation module 404 is configured to aggregate the perturbed data to obtain a mean estimation result.

[0187] Since the device introduced in the second embodiment of the present invention is the device used to implement the adaptive privacy protection mean estimation method based on local personalized differential privacy in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and deformation of the device, so it will not be elaborated here. Any device used in the method of the first embodiment of the present invention belongs to the scope protected by the present invention.

[0188] Embodiment Three

[0189] Based on the same inventive concept, the present invention also provides an adaptive privacy protection mean estimation system based on local personalized differential privacy, including the adaptive privacy protection mean estimation device based on local personalized differential privacy in Embodiment Two and end users. The end users are configured to obtain the privacy protection level based on the analysis of their own privacy data and privacy requirements and send it to the data aggregator, and perform perturbation processing on their privacy data according to the optimal perturbation mechanism and the corresponding optimal perturbation strategy recommended by the data aggregator to obtain perturbed data and send it to the data aggregator.

[0190] Since the adaptive privacy - protected mean - estimation system based on local personalized differential privacy introduced in Embodiment 3 of the present invention is the system adopted by the method for adaptive privacy - protected mean - estimation based on local personalized differential privacy in Embodiment 1 of the present invention, based on the method introduced in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of this system, so it will not be elaborated here. Any system adopted by the method in Embodiment 1 of the present invention falls within the scope of protection of the present invention.

[0191] Embodiment 4

[0192] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in Embodiment 1 is implemented.

[0193] Since the computer device introduced in Embodiment 4 of the present invention is the computer device adopted by the method for adaptive privacy - protected mean - estimation based on local personalized differential privacy in Embodiment 1 of the present invention, based on the method introduced in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of this computer device, so it will not be elaborated here. Any computer device adopted by the method in Embodiment 1 of the present invention falls within the scope of protection of the present invention.

[0194] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer - usable storage media (including but not limited to disk memory, CD - ROM, optical memory, etc.) containing computer - usable program code.

[0195] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general - purpose computer, a special - purpose computer, an embedded processor, or other programmable data - processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data - processing devices generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0196] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. An adaptive privacy-preserving mean estimation method based on local personalized differential privacy, characterized in that Including: Receiving the privacy protection level of the end user, where the privacy protection level of the end user is obtained by each end user based on the analysis of their own privacy data and privacy requirements, and the own privacy data is numerical data; According to the privacy protection level of the end user, end users with the same privacy protection level are divided into the same sub-group; Based on minimizing the maximum mean square error of the worst case, determining the optimal adaptive boundaries of the interval-based mean estimation mechanism and the neighbor-based mean estimation mechanism, and according to the privacy protection level of the local end user and the optimal adaptive boundary, recommending the optimal perturbation mechanism and the corresponding optimal perturbation strategy for each sub-group of users; so that the users in each sub-group use the corresponding optimal perturbation mechanism and the corresponding optimal perturbation strategy to perturb their privacy data to obtain perturbed data and send it to the data aggregator; Aggregating the perturbed data to obtain a mean estimation result.

2. The adaptive privacy protection mean estimation method based on local personalized differential privacy according to claim 1, characterized in that Based on minimizing the maximum mean square error of the worst case, determining the optimal adaptive boundaries of the interval-based mean estimation mechanism and the neighbor-based mean estimation mechanism, including: Obtaining the maximum mean square error in the worst case of the interval-based mean estimation mechanism; Obtaining the maximum mean square error in the worst case of the neighbor-based mean estimation mechanism; By minimizing the maximum mean square error in the worst case of the two mechanisms, determining the optimal adaptive boundaries of the interval-based mean estimation mechanism and the neighbor-based mean estimation mechanism.

3. The adaptive privacy protection mean estimation method based on local personalized differential privacy according to claim 2, characterized in that Obtaining the maximum mean square error in the worst case of the interval-based mean estimation mechanism, including: Calculating the variance of the perturbed data under the interval-based mean estimation mechanism; Calculating the expectation of the perturbed data under the interval-based mean estimation mechanism; According to the variance and expectation of the perturbed data under the interval-based mean estimation mechanism, calculating the mean square error under the interval-based mean estimation mechanism; According to the mean square error under the interval-based mean estimation mechanism, obtaining the maximum mean square error in the worst case of the interval-based mean estimation mechanism.

4. The adaptive privacy-preserving mean estimation method based on local personalized differential privacy according to claim 2, wherein Obtaining the maximum mean square error in the worst case of the neighbor-based mean estimation mechanism, including: Calculating the variance of the perturbed data under the neighbor-based mean estimation mechanism; Calculating the expectation of the perturbed data under the neighbor-based mean estimation mechanism; According to the variance and expectation of the perturbed data under the neighbor-based mean estimation mechanism, calculating the mean square error under the neighbor-based mean estimation mechanism; According to the mean square error under the neighbor-based mean estimation mechanism, obtaining the maximum mean square error in the worst case of the neighbor-based mean estimation mechanism.

5. The adaptive privacy protection mean estimation method based on local personalized differential privacy according to claim 2, characterized in that, By minimizing the maximum mean square error in the worst case of the two mechanisms, determining the optimal adaptive boundaries of the interval-based mean estimation mechanism and the neighbor-based mean estimation mechanism, including: Constructing a construction function for the difference between the maximum mean square errors in the worst case of the two mechanisms; Using the method of dichotomy to find the zero point to determine the adaptive boundary point, and obtaining the optimal adaptive boundary according to the adaptive boundary point.

6. The adaptive privacy protection mean estimation method based on local personalized differential privacy according to claim 5, characterized in that Using the method of dichotomy to find the zero point to determine the adaptive boundary point, and obtaining the optimal adaptive boundary according to the adaptive boundary point, including: Initializing the sub-interval according to the construction function; Based on the sign of the value obtained by substituting the initial boundary point into the constructor, determine the position of the zero point and update the sub-interval. Determine whether the preset iteration condition is satisfied. If it is satisfied, stop; if not, continue to update the sub-interval. After the iteration process ends, the midpoint of the latest sub-interval containing the zero point is used as the value of the optimal adaptive boundary.

7. The adaptive privacy protection mean estimation method based on local personalized differential privacy according to claim 1, characterized in that, Aggregate the perturbed data to obtain the mean estimation result, including: Each sub-group of users separately aggregates and analyzes the data perturbed by the users in the group, and estimates the mean value at each privacy protection level as the aggregation result of each sub-group. Weightedly combine the aggregation results from sub-groups under different privacy protection levels, where the weighted combination factor is constructed based on minimizing the maximum mean square error in the worst case.

8. An adaptive privacy-preserving mean estimation device based on local personalized differential privacy, characterized in that, The device is a data aggregator, including: A privacy protection level receiving module for receiving the privacy protection level of the terminal user, where the privacy protection level of the terminal user is obtained by each terminal user based on the analysis of its own privacy data and privacy requirements, and the own privacy data is numerical data. A sub-group division module for dividing terminal users with the same privacy protection level into the same sub-group according to the privacy protection level of the terminal user. An optimal perturbation mechanism and optimal perturbation strategy determination module for determining the optimal adaptive boundary of the interval-based mean estimation mechanism and the nearest neighbor-based mean estimation mechanism based on minimizing the maximum mean square error in the worst case, and recommending the optimal perturbation mechanism and the corresponding optimal perturbation strategy for each sub-group of users according to the privacy protection level of the local users and the optimal adaptive boundary; so that the users in each sub-group use the corresponding optimal perturbation mechanism and the corresponding optimal perturbation strategy to perturb their privacy data to obtain perturbed data and send it to the data aggregator. An aggregation module for aggregating the perturbed data to obtain the mean estimation result.

9. An adaptive privacy-preserving mean estimation system based on local personalized differential privacy, characterized in that It includes the adaptive privacy protection mean estimation device based on local personalized differential privacy and the terminal user as described in claim 8, where the terminal user is used to obtain the privacy protection level based on the analysis of its own privacy data and privacy requirements and send it to the data aggregator, and perturb its privacy data according to the optimal perturbation mechanism and the corresponding optimal perturbation strategy recommended by the data aggregator to obtain perturbed data and send it to the data aggregator.

10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the adaptive privacy protection mean estimation method based on local personalized differential privacy as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Self-adaptive privacy protection method, device and system based on minimum mean square error criterion

    CN115879152A