Multi-dimensional data mean value estimation method, device and system for dual personalized differential privacy protection

Through the dual personalized differential privacy protection mechanism, combined with the minimization of estimation variance and weighted merging method, the problems of low privacy protection intensity and personalized requirements in multidimensional data are solved, and mean estimation with high accuracy and privacy protection is achieved.

CN120805172APending Publication Date: 2025-10-17HUBEI UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510832012.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-17

Smart Images

  • Figure CN120805172A_ABST
    Figure CN120805172A_ABST
Patent Text Reader

Abstract

Aiming at privacy protection mean value estimation of multi-dimensional numeric data, a personalized privacy protection mechanism meeting-localization differential privacy is designed, and personalized privacy protection of a user level and a data level is provided. According to the invention, each user can select one privacy protection level from a plurality of preset privacy protection levels according to own privacy protection requirements, so that personalized privacy protection of the user level is realized, which is the first personalized privacy protection. And based on the expectation of the minimum estimation variance, determining an optimal dimension extraction parameter, and randomly extracting part of dimension data from all dimensions to carry out disturbance submission. And a user scoring strategy is adopted to determine a distribution strategy of privacy protection parameters, so that personalized privacy protection on a user data level is realized, which is second personalized privacy protection. And finally, a weighting factor is constructed based on the expectation of the estimation variance, and a weighted combination mode is adopted for mean value estimation under multiple privacy levels, so that the accuracy of the overall mean value estimation is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of privacy data protection, and more particularly to a multi-dimensional data mean estimation method, device and system for double personalized differential privacy protection. BACKGROUND

[0002] Mean estimation is a commonly used statistical method, widely used in data analysis and inference, especially when understanding the trend of data in a set. However, when dealing with sensitive data, privacy protection becomes increasingly important. For example, in the fields of medicine, finance, social networks, etc., a large amount of personal privacy information is involved, so it is crucial to protect this information from being leaked. Traditional mean estimation methods directly use raw data for calculation, which may lead to individual information leakage, thereby increasing privacy risks. Therefore, when performing mean estimation, it is essential to use appropriate privacy protection techniques, which can ensure user privacy security without affecting the overall data analysis results. In summary, privacy protection is an important consideration when performing mean estimation, especially in sensitive information fields. Through effective privacy protection methods, not only can the effectiveness of data analysis be ensured, but also the security and legality of data can be ensured.

[0003] Currently, most research on privacy-protected mean estimation focuses on single-dimensional data, and the privacy protection is relatively simple. However, real-world data is usually multi-dimensional. As the dimensionality of data increases, privacy protection becomes more complex, because multi-dimensional data contains multiple features, each of which may involve different types of sensitive information. In this case, how to balance privacy protection and data analysis accuracy becomes a major challenge. Too much noise may lead to a decrease in data analysis accuracy, while too little noise cannot effectively protect privacy. Local differential privacy technology provides an effective means to solve this problem. It adds noise to data at the source of each user's data to ensure that privacy leaks are effectively controlled. This technology is particularly important in some applications that require data collection from a large number of users, as it can prevent privacy leaks at the data collection stage.

[0004] Therefore, when dealing with multi-dimensional data, using local differential privacy technology can not only ensure the privacy and security of data, but also effectively solve the privacy protection problem without significantly affecting the accuracy of analysis.

[0005] Compared with - Local Differential Privacy (LDP) mechanism, The scheme under the local differential privacy model has a smaller error boundary and higher data utility. Existing schemes based on The multi-dimensional data privacy protection mean value estimation method of LDP outputs the privacy data as one of two determined values by probabilistic perturbation, and there is a problem of low privacy protection strength and large error. SUMMARY

[0006] In view of the technical problems of the prior art, the present application proposes a privacy protection mean value estimation for multi-dimensional numerical data, and designs a privacy protection mechanism that meets The personalized privacy protection mechanism of local differential privacy provides personalized privacy protection at the user level and personalized privacy protection at the data level. Each user can select a privacy protection level from a plurality of preset privacy protection levels according to the user's privacy protection requirements, to achieve personalized privacy protection at the user level, which is the first level of personalized privacy protection. Based on the expectation of the minimum estimation variance, the optimal dimension extraction parameter is determined, and part of the dimension data is randomly extracted from all dimensions for perturbation and submission, which effectively improves the accuracy of the mean value estimation. Since the user has personalized privacy protection requirements for the data in different dimensions, a user scoring strategy is used to determine the allocation strategy of the privacy protection parameter, to achieve personalized privacy protection at the data level, which is the second level of personalized privacy protection. Based on the expectation of the estimation variance, a weighting factor is constructed, and the weighted combination is used for the mean value estimation under multiple privacy levels, to further improve the accuracy of the overall mean value estimation. In summary, the present application provides a new solution for personalized privacy protection mean value estimation of multi-dimensional data.

[0007] To achieve the above-mentioned purpose, the first aspect of the present application provides a dual personalized differential privacy protection method for multi-dimensional data mean value estimation, comprising: determining the personalized privacy protection level at the user level; based on the personalized privacy protection level at the user level, users with the same privacy protection level at the user level are classified into the same sub-group; determining the optimal dimension extraction parameter by minimizing the expectation of the estimation variance according to the personalized privacy protection level; based on the personalized perturbation data, data aggregation and analysis are performed for each sub-group to estimate the mean value under each privacy protection level, wherein the personalized perturbation data is obtained after the user performs personalized privacy protection according to the personalized privacy protection level at the data level, and the personalized privacy protection level at the data level is obtained by the user of the corresponding sub-group based on the determined optimal dimension extraction parameter using the user scoring strategy; the aggregation results from the sub-groups under different privacy protection levels are weighted and combined to obtain the mean value estimation result.

[0008] In one embodiment, the personalized privacy protection level at the user level is determined, comprising: a plurality of preset privacy protection levels; receiving a privacy protection level selected by the user from the plurality of preset privacy protection levels as a personalized privacy protection level at the user level.

[0009] In an embodiment, according to the personalized privacy protection level, the optimal dimension extraction parameter is determined by minimizing the expectation of the estimation variance as follows:

[0010] wherein, is the optimal dimension extraction parameter, is the dimension of the user data, is a privacy protection budget in the personalized privacy protection level.

[0011] In an embodiment, the aggregated results from the subgroups under different privacy protection levels are combined by weighting to obtain a mean estimation result, including: constructing a weighting combination factor based on the expectation of the estimation variance of the subgroups; combining the aggregated results from the subgroups under different privacy protection levels according to the weighting combination factor to obtain the mean estimation result.

[0012] Based on the same inventive concept, the second aspect of the present application provides a method for estimating a mean of multi-dimensional data with double personalized differential privacy protection, including: sending the privacy protection level selected from the plurality of preset privacy protection levels to a data aggregator, so that the data aggregator obtains a personalized privacy protection level at the user level, and groups users with the same privacy protection level at the user level into the same subgroup based on the personalized privacy protection level at the user level; receiving the optimal dimension extraction parameter sent by the data aggregator, wherein the optimal dimension extraction parameter is determined by the data aggregator according to the privacy protection level parameter by minimizing the expectation of the estimation variance; based on the optimal dimension extraction parameter, performing privacy protection parameter allocation using a user scoring strategy to determine a personalized privacy protection level at the user data level; according to the personalized privacy protection level at the user data level, performing personalized privacy protection at the data level to obtain personalized perturbed data, and sending the personalized perturbed data to the data aggregator, so that the data aggregator performs data aggregation and analysis on each subgroup based on the personalized perturbed data to estimate the mean under each privacy protection level, and then combines the aggregated results from the subgroups under different privacy protection levels by weighting to obtain a mean estimation result.

[0013] In an embodiment, based on the optimal dimension extraction parameter, a user scoring strategy is adopted for privacy protection parameter allocation, and a personalized privacy protection level at the user data level is determined, including: Based on the optimal dimension extraction parameter, randomly select d dimensions from the user's dimensional data, wherein the optimal dimension extraction parameter is determined; Based on data sensitivity, the data of the extracted dimensions are scored for sensitivity, wherein the higher the sensitivity, the higher the score; The sensitivity scores of the selected dimensions are inversely normalized to construct an inverse normalized weight factor; According to the inverse normalized weight factor, the privacy protection parameters are personalized and allocated to obtain the personalized privacy protection level at the user data level.

[0014] In an embodiment, the sensitivity scores of the selected dimensions are inversely normalized to construct an inverse normalized weight factor, specifically:

[0015] wherein, is the inverse normalized weight factor of the jth dimension of the selected dimensions, is the sensitivity score of the jth dimension of the selected dimensions, i is the ith dimension of the selected dimensions.

[0016] Based on the same inventive concept, the third aspect of the present application provides a dual personalized differential privacy protection multi-dimensional data mean value estimation device, which is a data aggregator, including: determining a personalized privacy protection level at the user level; Based on the personalized privacy protection level at the user level, users with the same privacy protection level at the user level are grouped into the same sub-population; According to the personalized privacy protection level, the optimal dimension extraction parameter is determined by minimizing the expected variance of the estimation; Based on the personalized perturbed data, data aggregation and analysis are performed for each sub-population to estimate the mean value under each privacy protection level, wherein the personalized perturbed data is obtained after the user performs personalized privacy protection according to the personalized privacy protection level at the data level, and the personalized privacy protection level at the data level is obtained by the user of the corresponding sub-population based on the determined optimal dimension extraction parameter using the user scoring strategy; The aggregation results from the subgroups under different privacy protection levels are combined by weighting to obtain the mean estimation result.

[0017] Based on the same inventive concept, the fourth aspect of the present application provides a dual personalized differential privacy protection multi-dimensional data mean estimation device, which is a user end and comprises: The privacy protection level selected from the preset plurality of privacy protection levels is sent to the data aggregator, so that the data aggregator obtains the personalized privacy protection level at the user level, and the users with the same privacy protection level at the user level are classified into the same subgroup based on the personalized privacy protection level at the user level; The optimal dimension extraction parameter sent by the data aggregator is received, wherein the optimal dimension extraction parameter is determined by the data aggregator according to the privacy protection level parameter by minimizing the expected variance of the estimation; Based on the optimal dimension extraction parameter, the user scoring strategy is used for privacy protection parameter allocation to determine the personalized privacy protection level at the user data level; According to the personalized privacy protection level at the user data level, the personalized privacy protection at the data level is performed to obtain the personalized perturbed data, which is sent to the data aggregator, so that the data aggregator performs data aggregation and analysis on each subgroup based on the personalized perturbed data, estimates the mean under each privacy protection level, and then combines the aggregation results from the subgroups under different privacy protection levels by weighting to obtain the mean estimation result.

[0018] Based on the same inventive concept, the fifth aspect of the present application provides a dual personalized differential privacy protection multi-dimensional data mean estimation system, which comprises the dual personalized differential privacy protection multi-dimensional data mean estimation device of the third aspect and the dual personalized differential privacy protection multi-dimensional data mean estimation device of the fourth aspect.

[0019] Compared with the prior art, the present application has the following innovative points and beneficial technical effects: 1. For the problem of multi-dimensional data privacy protection mean estimation, the personalized privacy protection demand of the local end user is fully considered, including the personalized privacy protection at the user level and the personalized privacy protection at the data level, the dual privacy protection is realized, the user is provided with accurate personalized privacy protection, the enthusiasm of the user for actively participating in data collection and sharing is effectively stimulated, and the participation and initiative of the user are significantly improved.

[0020] 2. In the present application, the data collector only knows the user-level personalized privacy protection level, and does not know the user data-level personalized privacy protection parameters, i.e. the randomly selected part of the dimension-specific privacy budget and the relaxation factor allocation result of each user is unknown to the data collector. In fact, the privacy protection parameters at the user data level are also the user's privacy information, because the sensitive value weight of the user's data will also indirectly reveal the user's privacy information. Therefore, the present application further reduces the risk of user privacy leakage while ensuring personalized privacy protection at the data level.

[0021] 3. The present application is based on . - the personalized privacy protection of multi-dimensional data is designed based on the local differential privacy protection model, which extends - the multi-dimensional data mean estimation under the local differential privacy protection model, which takes into account double privacy protection, is a more practical privacy protection model. That is: when , the present application is to meet - the local differential privacy protection model, and the personalized privacy protection is considered, which shows that the present application is a more practical privacy protection model.

[0022] 4. The optimal dimension parameter of the present application is derived based on the minimization of the expected estimation variance, which has a strict theoretical basis, so that each sub-population can select the optimal perturbation dimension parameter for data perturbation, effectively improving the accuracy of multi-dimensional data mean estimation and balancing privacy protection and data availability.

[0023] 5. The weighted merging method constructed based on the minimization of the expected estimation variance, which cleverly uses the expected estimation error of each sub-population to construct the weighting factor, fully considers the contribution of each sub-population to the accuracy of the mean estimation, and gives full play to its advantages in data merging. Compared with the traditional method of directly accumulating and then averaging, it can further improve the accuracy of mean estimation.

[0024] 6. The multi-dimensional data personalized privacy protection mean estimation method proposed in the present application, in order to measure the sensitivity of the selected dimension data, adopts a user scoring strategy. In principle, the more sensitive the data, the higher the score, and then the privacy budget and the relaxation factor are divided according to the inverse normalized weight factor, providing a theoretical basis for personalized privacy protection at the data level.

[0025] 7. The present application realizes double privacy protection of user-level personalized privacy protection and data-level personalized privacy protection while ensuring good statistical estimation accuracy, and is a statistical and analytical method with great practical value, which has broad application prospects and important practical significance in actual application scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0027] Figure 1 Flow chart of the method for estimating the mean value of multi-dimensional data with double personalized differential privacy protection on the data aggregator side in the embodiment of the present application; Figure 2 Flow chart of the method for estimating the mean value of multi-dimensional data with double personalized differential privacy protection on the user side in the embodiment of the present application; Figure 3 Overall framework diagram of the method for estimating the mean value of multi-dimensional data with double personalized differential privacy protection in the embodiment of the present application. DETAILED DESCRIPTION

[0028] Embodiment one The embodiment provides a method for estimating the mean value of multi-dimensional data with double personalized differential privacy protection, please refer to Figure 1 , which comprises: S101: determining the personalized privacy protection level on the user level; S102: based on the personalized privacy protection level on the user level, grouping the users with the same privacy protection level on the user level into the same sub-group; S103: according to the personalized privacy protection level, determining the optimal dimension extraction parameter by minimizing the expectation of the estimation variance; S104: based on the personalized perturbation data, separately performing data aggregation and analysis on each sub-group to estimate the mean value under each privacy protection level, wherein the personalized perturbation data is obtained after the user performs personalized privacy protection according to the personalized privacy protection level on the data level, and the personalized privacy protection level on the data level is obtained by the user of the corresponding sub-group based on the determined optimal dimension extraction parameter using the user scoring strategy; S105: weighting and merging the aggregation results from the sub-groups under different privacy protection levels to obtain the mean value estimation result.

[0029] Specifically, the execution subject of the present embodiment is a data aggregator or collector, S101 determines the personalized privacy protection level on the user level, which is specifically realized by the following way: presetting a plurality of privacy protection levels; receiving the privacy protection level selected by the user from the preset plurality of privacy protection levels as the personalized privacy protection level on the user level.

[0030] In the implementation process, the data collector or aggregator presets multiple privacy protection levels, and the local user selects a privacy protection level (i.e. m -LDP, -LDP, -LDP, -LDP) from the preset

[0031] S102 is sub-group division. After receiving the user-level privacy protection level, the data collector or aggregator classifies users with the same user-level privacy protection level into the same sub-group, and sequentially .

[0032] S103 is the determination of the best dimension extraction parameter, which can be achieved by the following method:

[0033] wherein, is the best dimension extraction parameter, is the dimension of the user data, is the privacy protection budget in the personalized privacy protection level.

[0034] In the implementation process, the data aggregator determines the best dimension parameter based on the user-reported user-level personalized privacy protection level parameter in step S101 , and determines the best value in the perturbation of randomly selecting dimensions of data from d-dimensional data for each sub-group user based on the expectation of the minimum estimation variance.

[0035] S104 is sub-group data aggregation and statistical analysis. Each sub-group performs data aggregation and analysis separately to estimate the mean value under each privacy protection level.

[0036] S105 is to execute the weighted merging method to obtain effective mean value estimation results, which can be achieved by the following method: Construct a weighted merging factor based on the expectation of the estimation variance of the sub-group; According to the weighted merging factor, the aggregation results from the sub-groups under different privacy protection levels are merged to obtain the mean value estimation result.

[0037] Specifically, since the weighted merging factor is constructed based on the expectation of the estimation variance of the sub-group to obtain a high-precision mean value estimation result.

[0038] In the specific implementation process of the present invention, the personalized privacy protection at the user level is measured based on factors such as the user's attitude towards privacy protection and privacy preferences, because each user has different attitudes and preferences for privacy protection. Some users may be very sensitive to privacy protection and require strict privacy protection measures; while other users may have lower privacy protection needs and are willing to share more information for convenience. Based on their privacy protection needs, users can select a preset m Privacy protection levels: -LDP, -LDP... -LDP to achieve personalized privacy protection at the user level. For the h The privacy budget parameters under the privacy protection level of the user level in the subgroup, It is h The relaxation factor parameters under the privacy protection level of the user level in the subgroup are used together to control the privacy protection level of the private data in the group.

[0039] In the specific implementation process of the present invention, for subgroups For users in , the personalized privacy protection level at the user level is , when performing data perturbation, it is not targeted at d All dimensions of data are disturbed and submitted, but randomly selected In other words, when the privacy protection level at the user level is -LDP, when performing data perturbation, the privacy budget parameter and relaxed privacy parameters Assign to randomly selected dimensional data, rather than being evenly distributed d Dimensions of data. Among them, is the dimension of the data. Example 2 This embodiment provides a multi-dimensional data mean estimation method with dual personalized differential privacy protection. Figure 2 ,include: S201: Sending a privacy protection level selected from a plurality of preset privacy protection levels to a data aggregator, so that the data aggregator obtains a personalized privacy protection level at the user level, and groups users with the same user-level privacy protection level into the same subgroup based on the personalized privacy protection level at the user level; S202: receiving the optimal dimension extraction parameter sent by the data aggregator, wherein the optimal dimension extraction parameter is determined by the data aggregator according to the privacy protection level parameter by minimizing the expectation of the estimation variance; S203: based on the optimal dimension extraction parameter, adopting the user scoring strategy to perform privacy protection parameter allocation, and determining the individualized privacy protection level at the user data level; S204: according to the individualized privacy protection level at the user data level, performing individualized privacy protection at the data level to obtain individualized perturbed data, and sending the individualized perturbed data to the data aggregator, so that the data aggregator performs data aggregation and analysis on each sub-population based on the individualized perturbed data, estimates the mean value under each privacy protection level, and then combines the aggregation results from the sub-populations under different privacy protection levels to obtain the mean value estimation result.

[0040] Specifically, the execution subject of the embodiment is the user end, S201 selects the privacy protection level from the preset multiple privacy protection levels, and sends the privacy protection level to the data aggregator to determine the individualized privacy protection level at the user level.

[0041] S202 receives the optimal dimension extraction parameter sent by the data aggregator.

[0042] S203: adopting the user scoring strategy to perform privacy protection parameter allocation, and determining the individualized privacy protection level at the data level. Specifically, the following methods are used to achieve the individualized privacy protection level at the data level: based on the optimal dimension extraction parameter, randomly selecting d dimensions from the user's dimension data, wherein is the optimal dimension extraction parameter; based on the data sensitivity, scoring the data of the extracted dimensions, wherein the higher the sensitivity, the higher the score; performing reverse normalization processing on the sensitivity scores of the selected dimensions to construct a reverse normalized weight factor; according to the reverse normalized weight factor, performing individualized allocation of the privacy protection parameter to obtain the individualized privacy protection level at the user data level. Specifically, for the users in the sub-population , for the data of the d dimensions of the users, randomly selecting dimensions, satisfying: , and based on the data sensitivity, scoring the data of the extracted dimensions, the principle of the scoring being that the higher the sensitivity of the data, the higher the score, and performing reverse normalization processing on the scores of the The scores of the dimensions are subjected to reverse normalization processing, and based on this, a personalized privacy protection parameter at a user level is allocated, and then a personalized privacy protection level at a data level is determined.

[0043] The selected The scores of the sensitivity of the dimensions are subjected to reverse normalization processing, and a reverse normalized weight factor is constructed, and specifically:

[0044] The selected The reverse normalized weight factor of the jth dimension in the selected The score of the sensitivity of the jth dimension in the selected The score of the sensitivity of the jth dimension in the selected The ith dimension in the selected The ith dimension in the selected

[0045] S204: According to the personalized privacy protection level at the data level of the user, the personalized privacy protection at the data level is performed, personalized perturbed data is obtained, and is sent to the data aggregator.

[0046] The present application will be analyzed below in combination with the drawings and specific examples. Please refer to Figure 3 , which is the overall framework diagram of the multi-dimensional data mean estimation method of the double personalized differential privacy protection in the present application.

[0047] Given a data collector or aggregator and N users , it is assumed that the users are independent of each other, and each user has a numerical data containing d dimensions, and the d-dimensional data of the user is defined as , , , without loss of generality, it is assumed that the numerical data of each dimension , the purpose of the data collector is to obtain the mean estimation of each dimension data with high accuracy.

[0048] It should be noted that: in practice, the value range of numerical data may not be the interval [-1, 1], as long as it is mapped to the interval [-1, 1] by linear normalization or other methods, and then the mean estimation result is restored, so the method in the present application is applicable to mean estimation of any value range.

[0049] As known, directly sending user data to the collector will lead to user privacy leakage, so privacy protection operation is needed before data submission, and the present application is based on - The local differential privacy protection model designs personalized privacy protection of multi-dimensional data, realizes personalized privacy protection at a user level and privacy protection at a data level. The present application considers each According to the personalized privacy requirements of its user level from the preset m privacy protection levels: -LDP, -LDP… -LDP, select a suitable privacy protection level, and send this user-level personalized privacy protection level to the data collector or aggregator. Users with the same privacy protection level are divided into a sub-group.

[0050] First, from the single-dimensional data perturbation mechanism: (1) meet -Local differential privacy interval perturbation mechanism for single-dimensional data Assume that the input privacy data of the user is , the perturbed output is , and , where is the output domain, is a function of the privacy protection parameter , and once the privacy protection parameter is given, C is a certain constant. The user divides the output domain into three intervals: the left interval, the middle interval, and the right interval, and the privacy data is perturbed to the three intervals with different probabilities. The privacy protection level is measured by the differential privacy budget and the relaxation factor , and the perturbation mechanism that meets (- ) local differential privacy is designed for numerical data , and the best perturbation strategy is: . Where denotes the probability density function, is the left boundary of the middle interval, is the right boundary of the middle interval, and are functions of the privacy data , , , t is a numerical intermediate variable, p and q are functions of , , and is the left interval, is the middle interval, is the right interval. The specific parameter settings are as follows:

[0051]

[0052]

[0053]

[0054]

[0055]

[0056]

[0057] where, and b are functions of privacy protection parameter , once given privacy protection parameter , and b are specific constants. From the above perturbation mechanism, the probability that the perturbed data falls within the middle interval is , that is, the probability that the perturbed output data falls at each point in the middle interval is ; similarly to the above perturbation, the data is perturbed to the value of any point in the left and right intervals with probability .

[0058] Each end user performs the above perturbation strategy and sends the perturbed data to the data collector for aggregation, and its estimated mean value is:

[0059] where N is the data volume.

[0060] According to the definition of expectation, the expectation of the perturbed data is

[0061] According to the definition of variance, the variance of the perturbed data is described as:

[0062]

[0063]

[0064] Then the variance of the estimated mean value is:

[0065] (2) satisfies - Interval perturbation mechanism of multi-dimensional data of local personalized differential privacy a. Privacy parameter segmentation based on user scoring strategy The existing non-personalized scheme, for the selected part When the data of dimensions are disturbed, the disturbance parameters are evenly divided, such as 、 , such a setting does not take into account the different sensitivities of data in different dimensions, that is, data in different dimensions have personalized privacy protection requirements. Based on this, in order to measure the selected The sensitivity of the data in each dimension is determined by user scoring strategy. In principle, the more sensitive the data is, the higher the privacy protection requirement is, and the higher the score is. Score the sensitive data of each dimension, and set the score range from 0 to 100. Assume that the user selects The data sensitivity scores of the dimensions are as follows: 、 … , then the reverse normalized weight factor for:

[0066] From the above formula, we can get Inverse normalized weighting factors: 、 … ,satisfy: At this point, the selected The privacy protection parameters of the data in each dimension are: 、 … , that is, constructing weight factors based on sensitivity, and then determining the privacy budget and the segmentation strategy of relaxed privacy to achieve personalized privacy protection at the data level and meet -LDP, so the i-th dimension data satisfies -LDP. In principle, the higher the sensitivity of a user's data in a certain dimension, the higher the privacy protection requirement, and the smaller the privacy budget and relaxation factor should be allocated. The above weight factor construction method just meets this principle.

[0067] It is worth noting that in the present invention, the data collector only knows the personalized privacy protection level at the user level, and does not know the personalized privacy protection parameters at the user data level, that is, for subgroups , each user randomly selects The specific privacy protection parameter allocation results of each dimension are unknown to the data collector, which further protects the user's privacy information.

[0068] b. Personalized privacy protection at the data level and data aggregation of sub-groups According to the above privacy budget and relaxed privacy segmentation strategy, the user Selected j-th dimension data Adopting -LDP to perturb. That is, replacing the above interval perturbation mechanism satisfying -localized differential privacy with , , , , and setting the output as , the unselected dimensions are replaced with 0 in the output.

[0069] Since -LDP, the multi-dimensional data mean estimation mechanism is that the user randomly selects dimensions from d dimensions to perturb the output, so the jth dimension of the N users will not be selected, so the above perturbed output needs to be corrected, so the dimension extraction strategy is adopted, and the perturbed output of the jth dimension is described as:

[0070] The unselected dimensions are directly replaced with =0 when submitting data, then The estimated mean of the jth dimension data under -LDP is described as:

[0071] where is the personalized privacy protection level of the user level -LDP is

[0072] c. Optimal dimension parameter determination According to the definition of expectation, the expectation of perturbed data is

[0073] According to the definition of variance, the variance of perturbed data is described as:

[0074] wherein, each parameter is described as follows:

[0075]

[0076]

[0077]

[0078]

[0079] Since the privacy parameter allocation strategy of each user is unknown to the data collector, that is, the data collector only knows the personalized privacy protection level at the user level, but does not know the personalized privacy protection at the data level of each user, that is, the personalized privacy budget and the relaxation factor adopted by the data collector when performing data perturbation on each selected dimension of each user are unknown, therefore, when performing error evaluation, the theoretical analysis cannot be directly performed according to the specific allocated privacy budget and relaxation factor, and the present application adopts the expectation of variance to perform error evaluation: Wherein,

[0080]

[0081] Let = , = , the known parameter , and the parameter , is known, then

[0082]

[0083]

[0084] Since the specific distribution of is unknown, without loss of generality, it is assumed that is uniformly distributed in the interval of -1 to 1, then

[0085] Then the expectation of the variance of the above mean estimation is described as:

[0086] Under -LDP, randomly select dimensions of data for perturbation. In order to minimize the above error estimation, then determine the optimal dimension parameter based on the minimum estimation variance expectation:

[0087] Wherein, .

[0088] d. Weighted combination based on the expectation of the estimated variance ​​Considering the contribution of different privacy groups to the accuracy of mean value estimation, an expectation-based design of merging method is used to improve the accuracy of mean value estimation based on the variance of mean value estimation.

[0089] m The mean value estimation of the jth dimension privacy data of the hth sub-group is in turn: The final mean value estimation result is obtained by using weighted merging as shown in Figure 3

[0090] wherein, is a weighted factor for , and is the weighted factor of the hth sub-group, which is described as

[0091] wherein, is a value in the interval of 0 to 1, and satisfies , is the group Under the privacy protection level -LDP, the average value of the variance of the mean value estimation of the jth dimension.

[0092] Embodiment three Based on the same inventive concept, the embodiment discloses a dual personalized differential privacy protection multi-dimensional data mean value estimation device, which is a data aggregator, comprising: determining the user level personalized privacy protection level; based on the user level personalized privacy protection level, users with the same user level privacy protection level are classified into the same sub-group; According to the personalized privacy protection level, the optimal dimension extraction parameter is determined by minimizing the expectation of the estimation variance; Based on the personalized perturbation data, the data aggregation and analysis of each sub-group are carried out separately to estimate the mean value under each privacy protection level, wherein the personalized perturbation data is obtained after the user performs personalized privacy protection according to the data level personalized privacy protection level, and the data level personalized privacy protection level is obtained by the user of the corresponding sub-group based on the determined optimal dimension extraction parameter using the user scoring strategy; The aggregation results from the sub-groups under different privacy protection levels are weighted and merged to obtain the mean value estimation result.

[0093] ​Since the device introduced in the embodiment three of the present application is the device used for implementing the method in the embodiment one of the present application, the specific structure and deformation of the device can be understood by the person skilled in the art based on the method introduced in the embodiment three of the present application, and thus will not be described here again. The device used in the method in the embodiment three of the present application belongs to the range of the present application.

[0094] Embodiment four Based on the same inventive concept, the embodiment discloses a device for estimating the mean value of multi-dimensional data with double personalized differential privacy protection, which is a user terminal and comprises: The privacy protection level selected from the plurality of preset privacy protection levels is sent to the data aggregator, so that the data aggregator obtains the personalized privacy protection level at the user level, and the users with the same privacy protection level at the user level are classified into the same sub-population based on the personalized privacy protection level at the user level; The optimal dimension extraction parameter sent by the data aggregator is received, wherein the optimal dimension extraction parameter is determined by the data aggregator according to the privacy protection level parameter by minimizing the expected value of the estimation variance; Based on the optimal dimension extraction parameter, the user scoring strategy is used to allocate the privacy protection parameter, and the personalized privacy protection level at the user data level is determined; According to the personalized privacy protection level at the user data level, the personalized privacy protection at the data level is performed to obtain the personalized perturbed data, and the data aggregator is sent, so that the data aggregator performs data aggregation and analysis on each sub-population based on the personalized perturbed data, estimates the mean value under each privacy protection level, and then combines the aggregation results from the sub-populations under different privacy protection levels to obtain the mean value estimation result.

[0095] Since the device introduced in the embodiment four of the present application is the device used for implementing the method in the embodiment two of the present application, the specific structure and deformation of the device can be understood by the person skilled in the art based on the method introduced in the embodiment two of the present application, and thus will not be described here again. The device used in the method in the embodiment two of the present application belongs to the range of the present application.

[0096] Embodiment five Based on the same inventive concept, the present application also provides a personalized differential privacy protection system based on a user scoring strategy, comprising the double personalized differential privacy protection multi-dimensional data mean value estimation device in the embodiment three and the double personalized differential privacy protection multi-dimensional data mean value estimation device in the embodiment four.

[0097] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0098] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0099] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, the present invention is intended to include such changes and modifications to the embodiments of the present invention if they fall within the scope of the claims and their equivalents.

Claims

1. A multi-dimensional data mean estimation method with dual personalized differential privacy protection, characterized by: include: Determine personalized privacy protection levels at the user level; Based on the personalized privacy protection level at the user level, users with the same user-level privacy protection level are grouped into the same subgroup; According to the personalized privacy protection level, the optimal dimension extraction parameters are determined by minimizing the expected estimated variance; Based on personalized perturbation data, data is aggregated and analyzed for each subgroup separately to estimate the mean at each privacy protection level. The personalized perturbation data is obtained by performing personalized privacy protection based on the personalized privacy protection level at the data level. The personalized privacy protection level at the data level is obtained by the users of the corresponding subgroup using a user scoring strategy based on the determined optimal dimension extraction parameters. The aggregation results from subgroups at different privacy protection levels are weighted and merged to obtain the mean estimation result.

2. The multidimensional data mean estimation method with dual personalized differential privacy protection according to claim 1, characterized in that: Determine the level of personalized privacy protection at the user level, including: Preset multiple privacy protection levels; A privacy protection level selected by the user from a plurality of preset privacy protection levels is received as a personalized privacy protection level at the user level.

3. The multidimensional data mean estimation method with dual personalized differential privacy protection according to claim 1, characterized in that: According to the personalized privacy protection level, the optimal dimension extraction parameters are determined by minimizing the expected estimated variance: in, is the optimal dimension extraction parameter, is the dimension of user data, The privacy protection budget in the personalized privacy protection level.

4. The multidimensional data mean estimation method with dual personalized differential privacy protection according to claim 1, characterized in that: The aggregation results from subgroups at different privacy protection levels are weighted and merged to obtain the mean estimation results, including: Construct a weighted pooling factor based on the expected estimated variance of the subpopulations; The aggregation results from subgroups at different privacy protection levels are merged according to the weighted merging factor to obtain the mean estimation result.

5. A multi-dimensional data mean estimation method with dual personalized differential privacy protection, characterized by: include: A privacy protection level selected from a plurality of preset privacy protection levels is sent to a data aggregator, so that the data aggregator obtains a personalized privacy protection level at the user level, and groups users with the same user-level privacy protection level into the same subgroup based on the personalized privacy protection level at the user level; Receive the optimal dimension extraction parameters sent by the data aggregator, where the optimal dimension extraction parameters are determined by the data aggregator by minimizing the expected estimated variance according to the privacy protection level parameter; Parameters are extracted based on the optimal dimension, and a user scoring strategy is used to allocate privacy protection parameters to determine the personalized privacy protection level at the user data level; According to the personalized privacy protection level at the user data level, personalized privacy protection is performed at the data level to obtain personalized perturbation data, which is sent to the data aggregator so that the data aggregator can aggregate and analyze data for each subgroup separately based on the personalized perturbation data, estimate the mean at each privacy protection level, and then perform weighted merger of the aggregation results from subgroups at different privacy protection levels to obtain the mean estimation result.

6. The multi-dimensional data mean estimation method with dual personalized differential privacy protection according to claim 5, characterized in that: Parameters are extracted based on the optimal dimension, and a user scoring strategy is used to allocate privacy protection parameters to determine the personalized privacy protection level at the user data level, including: Extract parameters based on the best dimension from the user d Random selection in dimensional data dimensions, among which Extract parameters for the optimal dimension; Extraction based on data sensitivity The data of each dimension is scored for sensitivity, where the higher the sensitivity, the higher the score; The selected The sensitivity scores of each dimension are reverse normalized to construct the reverse normalized weight factor; The privacy protection parameters are personalized according to the inverse normalized weight factors to obtain the personalized privacy protection level at the user data level.

7. The multi-dimensional data mean estimation method with dual personalized differential privacy protection according to claim 6, characterized in that: The selected The sensitivity scores of each dimension are reverse normalized to construct the reverse normalized weight factors, which are as follows: in, For the selected The weight factor for the inverse normalization of the j-th dimension in the dimensions, For the selected The sensitivity score of the jth dimension among the dimensions, i For the selected The i-th dimension among the dimensions.

8. A multi-dimensional data mean estimation device with dual personalized differential privacy protection, characterized by: The device is a data aggregator and includes: Determine personalized privacy protection levels at the user level; Based on the personalized privacy protection level at the user level, users with the same user-level privacy protection level are grouped into the same subgroup; According to the personalized privacy protection level, the optimal dimension extraction parameters are determined by minimizing the expected estimated variance; Based on personalized perturbation data, data is aggregated and analyzed for each subgroup separately to estimate the mean at each privacy protection level. The personalized perturbation data is obtained by performing personalized privacy protection based on the personalized privacy protection level at the data level. The personalized privacy protection level at the data level is obtained by the users of the corresponding subgroup using a user scoring strategy based on the determined optimal dimension extraction parameters. The aggregation results from subgroups at different privacy protection levels are weighted and merged to obtain the mean estimation result.

9. A multi-dimensional data mean estimation device with dual personalized differential privacy protection, characterized in that: The device is a user terminal, including: A privacy protection level selected from a plurality of preset privacy protection levels is sent to a data aggregator, so that the data aggregator obtains a personalized privacy protection level at the user level, and groups users with the same user-level privacy protection level into the same subgroup based on the personalized privacy protection level at the user level; Receive the optimal dimension extraction parameters sent by the data aggregator, where the optimal dimension extraction parameters are determined by the data aggregator by minimizing the expected estimated variance according to the privacy protection level parameter; Parameters are extracted based on the optimal dimension, and a user scoring strategy is used to allocate privacy protection parameters to determine the personalized privacy protection level at the user data level; According to the personalized privacy protection level at the user data level, personalized privacy protection is performed at the data level to obtain personalized perturbation data, which is sent to the data aggregator so that the data aggregator can aggregate and analyze data for each subgroup separately based on the personalized perturbation data, estimate the mean at each privacy protection level, and then perform weighted merger of the aggregation results from subgroups at different privacy protection levels to obtain the mean estimation result.

10. A multidimensional data mean estimation system with dual personalized differential privacy protection, characterized by: It includes the multidimensional data mean estimation device with dual personalized differential privacy protection as claimed in claim 8 and the multidimensional data mean estimation device with dual personalized differential privacy protection as claimed in claim 9.

Citation Information

Cited By

  • Multi-dimensional data privacy calculation method and device, equipment, medium and product

    CN121413034A

  • A method, apparatus, device, medium, and product for multidimensional data privacy computing.

    CN121413034B