Data privacy iterative optimization method and device and medium

By iteratively optimizing the data noise addition process between the client and the server, combining the decision tree model to evaluate the importance of the data dimension, and dynamically adjusting the privacy budget, the problem of being unable to achieve fine noise control and data processing in the existing technology is solved, and efficient data privacy protection and data quality improvement are achieved.

CN120030591APending Publication Date: 2025-05-23CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510099160.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art cannot achieve fine noise control and data processing by analyzing the overall data set, especially the challenges of ensuring data quality while protecting data privacy.

Method used

A data privacy iterative optimization method is proposed. By performing the first noise addition on the client, the server integrates the first noise addition data from multiple clients, and allocates a second privacy budget for each dimension of the data. The client performs the second noise addition based on the new budget. This method combines the decision tree model to evaluate the importance of data dimensions, dynamically adjusts privacy budgets, and achieves fine control of noise levels.

Benefits of technology

It realizes the improvement of data availability and analysis accuracy while protecting data privacy, ensuring fine control of noise levels and high availability of data quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030591A_ABST
    Figure CN120030591A_ABST
Patent Text Reader

Abstract

The invention provides a data privacy iterative optimization method and device and a medium, and relates to the technical field of data processing, and the method comprises the steps: carrying out the first noise addition of the data of each dimension of first local multi-dimensional data according to a first sub-privacy budget allocated for each dimension of the multi-dimensional data, and obtaining a first sub-privacy budget allocated for each dimension of the first local multi-dimensional data; obtaining first disturbance data of the first local multi-dimensional data; sending the first disturbance data to a server, so that the server allocates a second sub-privacy budget for each dimension of the multi-dimensional data according to the first disturbance data from multiple clients, and sending the second sub-privacy budget to the clients; and according to the second sub-privacy budget of each dimension, performing second noise addition on the data of each dimension of the second local multi-dimensional data to obtain second disturbance data of the second local multi-dimensional data, and sending the second disturbance data to the server. According to the invention, through twice optimization of privacy and noise addition, data with good privacy and high quality are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure at least relates to the field of data processing technology, and in particular to a data privacy iterative optimization method, device and medium. Background Art

[0002] With the rapid development and continuous innovation of information technology, all walks of life are constantly generating huge amounts of data every day. By systematically collecting and deeply mining these data, valuable knowledge treasures can be revealed. Therefore, data collection and analysis activities are becoming more and more common, and their application scope is also expanding.

[0003] However, these data come from different companies or organizations, and the data of companies or organizations is private, and they hope to protect privacy while digging deeper into the value of data. However, the existing technology of adding noise to data locally cannot achieve fine noise control and data processing by analyzing the overall data set. Summary of the invention

[0004] The technical problem to be solved by the present disclosure is to provide a data privacy iterative optimization method, device and medium in view of the above-mentioned deficiencies, so as to solve the problem of how to achieve fine noise control and data processing by analyzing the overall data set.

[0005] In a first aspect, the present disclosure provides a data privacy iterative optimization method, which is applied to a client and includes:

[0006] According to the first privacy budget allocated to each dimension of the multidimensional data, respectively, performing a first noise addition on the data of each dimension of the first local multidimensional data to obtain first perturbed data of the first local multidimensional data;

[0007] Sending the first perturbation data to the server, so that the server allocates a second privacy budget for each dimension of the multi-dimensional data according to the first perturbation data from the multiple clients, and sends the second privacy budget to the client;

[0008] According to the second privacy budget of each dimension, the data of each dimension of the second local multidimensional data is denoised for a second time to obtain second perturbed data of the second local multidimensional data, and the second perturbed data is sent to the server.

[0009] Further, wherein:

[0010] The client selects different noise adding mechanisms to perform the first noise adding and the second noise adding for the data of the corresponding dimension according to whether each dimension corresponds to a discrete data attribute or a continuous data attribute.

[0011] The server uses a decision tree model to obtain the importance score of each dimension of the multidimensional data according to the first perturbation data from multiple clients, and allocates a second privacy budget according to the importance score.

[0012] Further, according to the first privacy budget allocated to each dimension of the multidimensional data, the data of each dimension of the first local multidimensional data is firstly denoised to obtain first perturbed data of the first local multidimensional data, specifically including:

[0013] Obtain a first total privacy budget ε initially preset for the multidimensional data;

[0014] Allocate the first privacy budget ε / n to each dimension attribute according to the number of dimensions n of the multidimensional data;

[0015] In response to the data attribute of the first dimension of the multidimensional data being a discrete data attribute, performing a first noise addition on the data of the first dimension of the first local multidimensional data using a k-random response mechanism according to ε / n;

[0016] In response to the data attribute of the second dimension of the multidimensional data being a continuous data attribute, performing a first noise addition on the data of the second dimension of the first local multidimensional data using a Laplace mechanism according to ε / n;

[0017] The results of the first noise addition of the data in each dimension are integrated to obtain the first disturbance data of the first local multi-dimensional data.

[0018] Further, according to the second privacy budget of each dimension, the data of each dimension of the second local multidimensional data is subjected to a second noise addition to obtain second disturbance data of the second local multidimensional data, and the second disturbance data is sent to the server, specifically including:

[0019] Receive the second privacy budget a of n dimensions from the server i *ε',a i (0﹤a i ﹤1,∑ n a i =1) is the importance score of the i-th (i∈[1,n]) dimension of the multidimensional data, and ε' is the second total privacy budget;

[0020] In response to the data attribute of the first dimension of the multidimensional data being a discrete data attribute, according to a i *ε' uses a k-random response mechanism to perform a second noise addition on the data of the first dimension of the second local multidimensional data;

[0021] In response to the data attribute of the second dimension of the multidimensional data being a continuous data attribute, according to a i*ε' uses the Laplace mechanism to perform a second noise addition on the data of the second dimension of the second local multidimensional data;

[0022] The result after the second noise addition of the data in each dimension is synthesized to obtain the second disturbance data of the second local multi-dimensional data;

[0023] The second perturbation data is sent to the server, so that the server publishes the second perturbation data, and uses the second perturbation data to solve the classification problem or the regression problem.

[0024] In a second aspect, the present disclosure provides a data privacy iterative optimization method, which is applied to a server and includes:

[0025] Receiving first perturbed data from multiple clients, where the first perturbed data is obtained by the clients performing a first noise addition on data of each dimension of first local multidimensional data according to a first privacy budget allocated to each dimension of the multidimensional data;

[0026] Allocate a second sub-privacy budget for each dimension of the multidimensional data according to the first perturbation data from the multiple clients, and send the second sub-privacy budget to the client, so that the client performs a second noise addition on the data of each dimension of the second local multidimensional data according to the second sub-privacy budget of each dimension, so as to obtain second perturbation data of the second local multidimensional data;

[0027] Second perturbation data is received from a plurality of clients.

[0028] Further, wherein:

[0029] The client selects different noise adding mechanisms to perform the first noise adding and the second noise adding for the data of the corresponding dimension according to whether each dimension corresponds to a discrete data attribute or a continuous data attribute.

[0030] The server uses a decision tree model to obtain the importance score of each dimension of the multidimensional data according to the first perturbation data from multiple clients, and allocates a second privacy budget according to the importance score.

[0031] Furthermore, according to the first perturbation data from the multiple clients, a second privacy budget is allocated to each dimension of the multi-dimensional data, specifically including:

[0032] Get the preset multidimensional data problem analysis type;

[0033] According to the type of multidimensional data problem analysis, which is to use multidimensional data to solve classification problems or regression problems, the decision tree model is selected as a classification tree model or a regression tree model;

[0034] Input the first perturbed data from the m clients into the classification tree model or the decision tree model to obtain the minimized impurity of the data of each dimension split in the classification tree model or the minimized mean square error of the data split in the decision tree model;

[0035] According to minimizing impurity or minimizing mean square error, and the second total privacy budget ε', obtain the importance score a of the i-th (i∈[1,n]) dimension i (0﹤a i ﹤1,∑ n a i =1), according to the importance score a i Allocate the second privacy budget a to n dimensions of multidimensional data i *ε'.

[0036] Further, receiving second disturbance data from multiple clients specifically includes:

[0037] Second perturbation data is received from the m clients, the second perturbation data is published, and the second perturbation data is used to solve a classification problem or a regression problem.

[0038] In a third aspect, the present disclosure provides a client, which is used for iterative optimization of data privacy and includes:

[0039] A first noise adding module, configured to perform a first noise adding on the data of each dimension of the first local multidimensional data according to the first privacy budget allocated to each dimension of the multidimensional data, so as to obtain first disturbance data of the first local multidimensional data;

[0040] a privacy acquisition module, connected to the first noise adding module, and configured to send the first disturbance data to the server, so that the server allocates a second privacy budget for each dimension of the multidimensional data according to the first disturbance data from the plurality of clients, and sends the second privacy budget to the client;

[0041] The second noise addition module is connected to the privacy acquisition module, and is used to perform a second noise addition on the data of each dimension of the second local multidimensional data according to the second privacy budget of each dimension, so as to obtain second disturbance data of the second local multidimensional data, and send the second disturbance data to the server.

[0042] In a fourth aspect, the present disclosure provides a server, the server being used for iterative optimization of data privacy, and comprising:

[0043] A first data module is used to receive first perturbed data from multiple clients, where the first perturbed data is obtained by the client performing a first noise addition on data of each dimension of first local multidimensional data according to a first privacy budget allocated to each dimension of the multidimensional data;

[0044] a privacy allocation module connected to the first data module, configured to allocate a second privacy budget to each dimension of the multidimensional data according to the first perturbation data from the plurality of clients, and send the second privacy budget to the client, so that the client performs a second noise addition on the data of each dimension of the second local multidimensional data according to the second privacy budget of each dimension, so as to obtain second perturbation data of the second local multidimensional data;

[0045] The second data module is connected to the privacy allocation module and is used to receive second disturbance data from multiple clients.

[0046] In a fifth aspect, the present disclosure provides a data privacy iterative optimization system, the system comprising:

[0047] Multiple clients, used to implement the data privacy iterative optimization method as described in the first aspect above;

[0048] The server is connected to multiple clients and is used to implement the data privacy iterative optimization method as described in the second aspect above.

[0049] In a sixth aspect, the present disclosure provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the data privacy iterative optimization method as described above is implemented.

[0050] The present disclosure provides a data privacy iterative optimization method, device and medium, which performs the first noise addition on the client, and the server integrates the data after the first noise addition from multiple clients to allocate a privacy budget for each dimension of the data. The client performs the second data noise addition using the privacy budget allocated by the server, and obtains the privacy allocation of each data dimension according to the overall data of multiple clients, thereby achieving a reasonable allocation of the privacy budget, and further achieving fine control of the noise level and high availability of data quality, that is, providing data with good privacy and high quality by optimizing privacy noise twice. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a flow chart of a data privacy iterative optimization method according to an embodiment of the present disclosure;

[0052] Figure 2 is a flow chart of another data privacy iterative optimization method according to an embodiment of the present disclosure;

[0053] Figure 3 is a schematic diagram of the structure of a client according to an embodiment of the present disclosure;

[0054] Figure 4 is a structural diagram of a server according to an embodiment of the present disclosure;

[0055] Figure 5 is a structural diagram of a data privacy iterative optimization system according to an embodiment of the present disclosure;

[0056] Figure 6 is a schematic diagram of the effect of conducting an experiment using the Iris dataset according to an embodiment of the present disclosure;

[0057] Figure 7 It is a schematic diagram of the effect of using a Wine dataset to conduct an experiment in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0058] In order to enable those skilled in the art to better understand the technical solution of the present disclosure, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.

[0059] It should be understood that the specific embodiments and drawings described herein are only used to explain the present disclosure rather than to limit the present disclosure.

[0060] It can be understood that, in the absence of conflict, the various embodiments of the present disclosure and the various features in the embodiments can be combined with each other.

[0061] It can be understood that, for the convenience of description, the drawings of the present disclosure only show parts related to the present disclosure, while parts irrelevant to the present disclosure are not shown in the drawings.

[0062] It can be understood that each module or unit involved in the embodiments of the present disclosure may correspond to only one entity structure, or may be composed of multiple entity structures, or multiple modules or units may be integrated into one entity structure.

[0063] It will be understood that, in the absence of conflict, the functions and steps marked in the flowcharts and block diagrams of the present disclosure may occur in an order different from that marked in the drawings.

[0064] It is understood that the flowcharts and block diagrams of the present disclosure illustrate the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to the embodiments of the present disclosure. Each box in the flowchart or block diagram may represent a module, unit, program segment, or code, which contains executable instructions for implementing the specified functions. Moreover, each box or combination of boxes in the block diagram and flowchart may be implemented by a hardware-based device that implements the specified functions, or by a combination of hardware and computer instructions.

[0065] It can be understood that the modules and units involved in the embodiments of the present disclosure can be implemented by software or hardware, for example, the modules and units can be located in a processor.

[0066] Embodiment 1:

[0067] likeFigure 1 As shown, the present disclosure provides a data privacy iterative optimization method, which is applied to a client and includes:

[0068] S11, performing a first noise addition on the data of each dimension of the first local multidimensional data according to the first privacy budget allocated to each dimension of the multidimensional data, so as to obtain first perturbed data of the first local multidimensional data;

[0069] S12, sending the first disturbance data to the server, so that the server allocates a second privacy budget for each dimension of the multidimensional data according to the first disturbance data from the multiple clients, and sends the second privacy budget to the client;

[0070] S13. Perform a second noise addition on the data of each dimension of the second local multidimensional data according to the second privacy budget of each dimension to obtain second disturbance data of the second local multidimensional data, and send the second disturbance data to the server.

[0071] In this embodiment, by performing the first noise addition on the client, the server integrates the data from multiple clients after the first noise addition, allocates a privacy budget for each dimension of the data, and the client uses the privacy budget allocated by the server to perform the second data noise addition, and obtains the privacy allocation of each data dimension according to the overall data of multiple clients, so as to achieve a reasonable allocation of the privacy budget, thereby achieving fine control of the noise level and high availability of data quality. That is, through two optimizations of privacy noise addition, the data of the first noise addition is used to obtain the privacy items used for the second noise addition, and the privacy items used for the second noise addition take into account the overall characteristics of the data of multiple clients, so as to provide data with good privacy and high quality for comprehensive use on the server side. The method of two noise additions is adopted, and the privacy budget allocation of each attribute is dynamically adjusted based on the analysis of the first noise addition data. This strategy aims to further improve the availability of the second noise addition data while protecting privacy through more reasonable privacy budget allocation.

[0072] and Figure 1 The corresponding method applied to the server is as follows Figure 2 As shown, the corresponding client structure is as follows Figure 3 As shown, the server structure is as follows Figure 4 As shown, the corresponding modules execute the steps with corresponding numbers. Figure 2 The corresponding detailed description of the method is not repeated here, and those skilled in the art can easily deduce the peer action based on the corresponding relationship. In addition, the data privacy iteration optimization system of the integrated client and server is as follows Figure 5 shown.

[0073] Specifically, collecting data from different enterprises and organizations, conducting data analysis, breaking information silos, and mining greater data value are also accompanied by a serious problem - privacy leakage. If sensitive user data is unfortunately leaked, it will lead to a series of public security problems, such as fraud and malicious harassment, bringing significant troubles to society. To effectively address this challenge, the concept of differential privacy (DP) has been proposed, aiming to provide solid privacy protection for data processing activities in various fields. However, differential privacy mainly focuses on centralized datasets and assumes that the server is trustworthy. This means that at the client level, data is not fully protected, and the privacy issue remains unresolved.

[0074] In this context, Local Differential Privacy (LDP), as an innovative extension of DP, has emerged. Under the LDP framework, before sending data, users will perturb the data themselves on the client side and then upload the perturbed data to the server. This mechanism ensures that the server only receives the perturbed version of the user data, while the user's real data always remains on the user's own device and is never exposed to the outside world. Therefore, even if the server is not completely trustworthy, the user's privacy security can be effectively guaranteed. With its excellent security performance, LDP has attracted the attention and research of many well-known institutions and has been widely applied in practice, such as Google Chrome, Apple's iOS and macOS systems, and Microsoft's Windows Insiders project, etc.

[0075] However, compared with differential privacy, local differential privacy cannot perform fine-grained noise control and data processing by analyzing the overall dataset because noise is added locally. Under the framework of Local Differential Privacy (LDP), noise addition processing is performed on user data on the client side, which ensures that the privacy of the data is protected before leaving the user device. However, this data processing at the local level also brings a significant challenge: since the server cannot access the complete and original dataset, it is difficult to conduct in-depth analysis of each data attribute, and thus it is impossible to achieve fine-tuning of the noise level for each attribute and optimization of data processing.

[0076] Most existing methods for dealing with local differential privacy often adopt a relatively simple and direct strategy, that is, evenly distributing the privacy budget for each attribute in the dataset. This "one-size-fits-all" approach is easy to implement but ignores the significant differences and importance differences that may exist between different attributes. As a result, for those critical and sensitive attributes, too low a privacy budget leads to unnecessary noise increase and data quality degradation.

[0077] Therefore, how to ensure the privacy of data while improving the accuracy and availability of data as much as possible under the framework of local differential privacy has become an urgent problem to be solved. This requires that when designing a privacy protection strategy, the specific characteristics and importance of each attribute must be fully considered to achieve a reasonable allocation of the privacy budget and fine control of the noise level.

[0078] In view of this, this embodiment introduces an iterative optimization mechanism, and uses information entropy, decision tree algorithm and other calculations to distinguish the importance of different data attributes, thereby allocating privacy budgets in a gradient manner. This innovative method further improves data availability while ensuring privacy security, and opens up new prospects for the application of localized differential privacy.

[0079] This embodiment proposes a method for collecting multidimensional data under local differential privacy, which can be specifically called an iterative optimization multidimensional data collection method based on local differential privacy. Its purpose is to obtain data that guarantees both data quality and data privacy security by iteratively optimizing the privacy budget of different data dimensions of the collected data in the process of collecting data from multiple clients to the server. The method first performs a first noise addition and aggregation on the user data to obtain a first noisy data set. Secondly, the importance of each attribute is distinguished by a decision tree algorithm. Then, according to the importance of each attribute, the corresponding privacy budget is adjusted, the important attributes increase their privacy budget, and the secondary attributes reduce their privacy budget. Finally, the user data is noised for a second time using the adjusted privacy budget. The data set obtained by aggregation after the second noise addition can improve the availability of the collected data while ensuring the same privacy protection strength (the total privacy budget allocated to the user data remains unchanged).

[0080] In one embodiment, wherein:

[0081] The client selects different noise adding mechanisms to perform the first noise adding and the second noise adding for the data of the corresponding dimension according to whether each dimension corresponds to a discrete data attribute or a continuous data attribute.

[0082] The server uses a decision tree model to obtain the importance score of each dimension of the multidimensional data according to the first perturbation data from multiple clients, and allocates a second privacy budget according to the importance score.

[0083] In this embodiment, when constructing a data processing flow based on local differential privacy (LDP), a series of refined steps are taken to ensure the dual improvement of privacy protection and data quality, including: different noise mechanisms can be used on the client side to improve the effect of privacy protection for data with different data attributes, and decision trees are used on the server side to analyze the importance of each dimension of the data, and privacy allocation is determined according to different importance, so as to improve the quality of the final data when it is used. The decision tree model can use information entropy and the like to determine the importance of data dimensions. It can be understood that the decision tree model is not the only method that can be used to determine the importance of data dimensions. Through different noise processing mechanisms, extremely high diversity is shown in the processing of data types. Whether it is discrete, continuous or mixed data (that is, data containing both discrete and continuous components), it can be effectively processed by this method. This wide range of data type compatibility enables the method of this embodiment to be applied to more diversified data collection scenarios. The first noise data is analyzed by information entropy and decision tree algorithm to distinguish the importance of each attribute. On this basis, the allocation of privacy budget is dynamically adjusted, and the privacy budget is increased for important attributes and reduced for secondary attributes. This strategy aims to maximize data availability under the same privacy budget (privacy protection strength).

[0084] In one embodiment, S11, performing a first noise addition on data of each dimension of the first local multidimensional data according to the first privacy budget allocated to each dimension of the multidimensional data to obtain first perturbed data of the first local multidimensional data, specifically includes:

[0085] Obtain a first total privacy budget ε initially preset for the multidimensional data;

[0086] Allocate the first privacy budget ε / n to each dimension attribute according to the number of dimensions n of the multidimensional data;

[0087] In response to the data attribute of the first dimension of the multidimensional data being a discrete data attribute, performing a first noise addition on the data of the first dimension of the first local multidimensional data using a k-random response mechanism according to ε / n;

[0088] In response to the data attribute of the second dimension of the multidimensional data being a continuous data attribute, performing a first noise addition on the data of the second dimension of the first local multidimensional data using a Laplace mechanism according to ε / n;

[0089] The results of the first noise addition of the data in each dimension are integrated to obtain the first disturbance data of the first local multi-dimensional data.

[0090] In this embodiment, the following is the detailed process after optimization and expansion:

[0091] 1) Initial privacy budget allocation: First, a basic strategy is adopted, which is to divide the privacy budget equally among each attribute (each dimension) in the dataset. For example, in the Iris dataset (a commonly used classification experiment dataset), each data contains 4 attributes (4 dimensions), namely, sepal length, sepal width, petal length, and petal width. If the total privacy budget is ε = 1, the privacy budget allocated to each attribute is ε / 4 = 0.25. This step provides a starting point for subsequent privacy protection processing. Although it may not be optimal, it provides a basis for subsequent iterative optimization.

[0092] 2) Client data perturbation: For each client, for discrete data attributes, the k-RR (k-randomized response) mechanism is used to perturb the user's data. For example, suppose the candidate value of the discrete data attribute A is i (i = 1, 2, 3), the attribute value of A of a piece of original data x is 1, and the privacy budget allocated to the attribute is ε / 4. Pick a random number in the range. If the random number is within If the random number is within the range, the A attribute value of x is still 1; If the value of A of x is within the range, the perturbation of A attribute value is 2; the random number is If the value of x is within the range, the perturbation of attribute A of x is 3. For continuous data attributes, the Laplace mechanism is used to perturb the user's data. For example, assuming that the value range of the continuous data attribute B is [0,2], the attribute value of B of a piece of original data y is 1, and the privacy budget assigned to the attribute is ε / 4, the sensitivity Δf can be obtained as 2 through the value range, and the scale parameter of the Laplace noise is calculated as From the Laplace distribution (probability density function is The noise is generated by sampling in the original data, and the generated Laplace noise is added to the attribute value of the original data to obtain the perturbed attribute value. After the above four attributes are perturbed by the above two mechanisms, they all satisfy ε / 4-local differential privacy. According to the sequence composition of local differential privacy, the perturbed data satisfies ε-local differential privacy. Finally, the perturbed data is uploaded to the server, and the server will first aggregate the perturbed data of all users to obtain the first perturbed data set.

[0093] In one embodiment, S13, performing a second noise addition on the data of each dimension of the second local multidimensional data according to the second privacy budget of each dimension to obtain second disturbance data of the second local multidimensional data, and sending the second disturbance data to the server, specifically includes:

[0094] Receive the second privacy budget a of n dimensions from the server i *ε',a i (0﹤a i ﹤1,∑ n a i =1) is the importance score of the i-th (i∈[1,n]) dimension of the multidimensional data, and ε' is the second total privacy budget;

[0095] In response to the data attribute of the first dimension of the multidimensional data being a discrete data attribute, according to a i *ε' uses a k-random response mechanism to perform a second noise addition on the data of the first dimension of the second local multidimensional data;

[0096] In response to the data attribute of the second dimension of the multidimensional data being a continuous data attribute, according to a i *ε' uses the Laplace mechanism to perform a second noise addition on the data of the second dimension of the second local multidimensional data;

[0097] The result after the second noise addition of the data in each dimension is synthesized to obtain the second disturbance data of the second local multi-dimensional data;

[0098] The second perturbation data is sent to the server, so that the server publishes the second perturbation data, and uses the second perturbation data to solve the classification problem or the regression problem.

[0099] In this embodiment, the detailed process after optimization and expansion on the server side includes:

[0100] 3) Decision tree algorithm learning: After receiving the perturbation data from all clients, the server uses the decision tree algorithm to learn the perturbation data set. First, the decision tree algorithm discretizes the continuous attributes. This is usually done by finding the best split point, which divides the value range of the continuous attribute into two or more intervals. Then, the decision tree algorithm selects the best features for splitting, recursive splitting, generating leaf nodes, pruning and other steps to obtain the final model. The decision tree algorithm can be implemented with reference to the prior art. This embodiment uses the decision tree to obtain the basis for privacy allocation.

[0101] 4) Privacy budget reallocation: During the construction process, the decision tree will be split according to the characteristics of the data to minimize impurity (for classification problems) or error (for regression problems). For classification trees, the importance of a feature can be measured by the impurity it reduces when it splits (such as Gini impurity or entropy). If a feature can significantly reduce impurity at multiple split points, then its importance is higher. For regression trees, the importance of a feature can be measured by the mean square error it reduces when it splits. Similarly, if a feature can significantly reduce error at multiple split points, then its importance is higher. Based on the above principles, the decision tree model can be used to obtain the importance score a of each dimension of multidimensional data i , you can design a i The value is between 0 and 1, and there are n a i The sum is 1. Based on the importance score of each attribute (dimension) in the decision tree model, the server will reallocate the privacy budget of each attribute. Specifically, the privacy budget is allocated according to the proportion of the importance score, so that the privacy budget of attributes with higher importance scores increases, and the privacy budget of attributes with lower importance scores decreases. In this way, the noise level of each attribute can be controlled more accurately, so as to preserve the accuracy and usefulness of the data as much as possible while protecting privacy.

[0102] After that, the detailed process after optimization and expansion also includes:

[0103] 5) Client data re-perturbation: After receiving the new privacy budget, each client will use the k-RR mechanism and Laplace mechanism to perturb the user's original data again. This perturbation is performed under the new privacy budget constraint, so it can more accurately balance privacy protection and data quality. The perturbation mechanism is the same as step 2). The size of the noise added by the mechanism is related to the allocated privacy budget, so the perturbation is constrained by the privacy budget. The perturbed data is uploaded to the server again.

[0104] 6) Server-side data aggregation and statistical analysis: Finally, the server will aggregate all the perturbed data uploaded again by the client and perform statistical analysis, such as using the second perturbed data to analyze classification or regression problems. Since the data at this time has been perturbed twice and the privacy budget has been fine-tuned, it can provide more accurate and reliable analysis results while protecting user privacy. For example, in the field of data collection and release, the perturbed data set collected for the first time is not released, and the perturbed data set collected for the second time is released. For the data collected for the first time, even if the server is untrusted, there is no privacy leakage because local differential privacy is satisfied. There is no risk of privacy leakage for the twice perturbed data if the server is trustworthy; if the server is untrustworthy, there is no risk of privacy leakage because the privacy budgets used in the two mechanisms are different.

[0105] It is understandable that the total privacy budget used for the first noise addition and the total privacy budget used for the second noise addition can be the same or different. The second privacy budget can be directly calculated by the server and sent to the client, or the server can calculate the importance score and send each importance score to the client, and the client calculates the second privacy budget by itself. Such changes in details do not deviate from the basic inventive concept of the present disclosure and are within the scope of protection requested by the present disclosure.

[0106] Through the above series of optimization and expansion steps, we can achieve a dual improvement in privacy protection and data quality under the framework of local differential privacy. This method not only improves the efficiency and accuracy of data processing, but also provides new ideas and methodological references for future privacy protection research.

[0107] The following introduces two experiments using the method of this embodiment to perform privacy processing on actual data sets.

[0108] In order to verify the feasibility of the present invention, verification experiments were conducted using the Iris dataset (iris flower dataset) and the Wine dataset (wine identification dataset). The values ​​of the overall privacy budget were set to 10, 20, 30, 40, and 50, respectively, and the K-means algorithm was selected as the classification algorithm. To ensure the reliability of the accuracy results, for each overall privacy budget, the experiment was repeated 1,000 times and the average accuracy was taken as the final result. The experimental results of the Iris dataset are shown in the figure. Figure 6 As shown in Figure 2, the experimental results on the Wine dataset are as follows: Figure 7 shown.

[0109] The solid line in the figure is the accuracy of data classification after conventional privacy processing, and the dotted line is the accuracy of data classification after privacy processing using the method of this embodiment. It can be seen that after privacy processing using the method of this embodiment, the accuracy of data classification has been improved under the same privacy budget, especially for the Wine dataset. When the total privacy budget is 40, the effect is better than that of the conventional method with a total privacy budget of 50, that is, under better privacy protection, better data availability can be obtained. Therefore, the experimental results show that the method of this embodiment can improve the availability of data while keeping the overall privacy budget (privacy protection strength) unchanged.

[0110] In summary, the advantages of this embodiment are mainly reflected in the following two aspects:

[0111] Local differential privacy framework integrated with iterative optimization mechanism: In view of the fact that traditional local differential privacy technology tends to directly add noise to the data, but less consideration is given to the diversity of data characteristics, this embodiment introduces a privacy iterative optimization mechanism, which aims to perform a more detailed analysis of each data attribute before implementing data privacy protection. Through this preprocessing link, noise addition can be more accurately controlled and data can be processed more scientifically, striving to preserve the original information value of the data as much as possible while maintaining user privacy. This mechanism dynamically adjusts the noise strategy through a continuous iterative optimization process, aiming to find a more ideal balance between privacy protection and data quality.

[0112] Attribute importance evaluation and dynamic adjustment of privacy budget based on decision tree: In order to further improve the practicality of data under privacy protection, this embodiment also proposes a method for evaluating attribute importance using a decision tree. This method aims to measure the importance of each attribute of the data set by obtaining the attribute importance score (indicators such as information entropy, information gain, and Gini coefficient can also be calculated) after learning the data set through the decision tree algorithm. Based on this importance assessment, try to allocate more privacy budget to those more important attributes, in order to provide more detailed protection and higher data authenticity for attributes that have a greater impact on data analysis results while maintaining the overall privacy protection strength. This dynamic adjustment strategy of the privacy budget aims to improve the availability and analysis accuracy of data under privacy protection, while promoting more efficient use of privacy resources.

[0113] Embodiment 2:

[0114] like Figure 2 As shown, the present disclosure provides a data privacy iterative optimization method, which is applied to a server and includes:

[0115] S21, receiving first disturbance data from multiple clients, where the first disturbance data is obtained by the client performing a first noise addition on data of each dimension of first local multidimensional data according to a first privacy budget allocated to each dimension of the multidimensional data;

[0116] S22, allocating a second sub-privacy budget for each dimension of the multidimensional data according to the first perturbation data from the multiple clients, and sending the second sub-privacy budget to the client, so that the client performs a second noise addition on the data of each dimension of the second local multidimensional data according to the second sub-privacy budget of each dimension, so as to obtain second perturbation data of the second local multidimensional data;

[0117] S23. Receive second disturbance data from multiple clients.

[0118] In one embodiment, wherein:

[0119] The client selects different noise adding mechanisms to perform the first noise adding and the second noise adding for the data of the corresponding dimension according to whether each dimension corresponds to a discrete data attribute or a continuous data attribute.

[0120] The server uses a decision tree model to obtain the importance score of each dimension of the multidimensional data according to the first perturbation data from multiple clients, and allocates a second privacy budget according to the importance score.

[0121] In one implementation, in S22, according to the first perturbation data from the multiple clients, a second privacy budget is allocated to each dimension of the multi-dimensional data, specifically including:

[0122] Get the preset multidimensional data problem analysis type;

[0123] According to the type of multidimensional data problem analysis, which is to use multidimensional data to solve classification problems or regression problems, the decision tree model is selected as a classification tree model or a regression tree model;

[0124] Input the first perturbed data from the m clients into the classification tree model or the decision tree model to obtain the minimized impurity of the data of each dimension split in the classification tree model or the minimized mean square error of the data split in the decision tree model;

[0125] According to minimizing impurity or minimizing mean square error, and the second total privacy budget ε', obtain the importance score a of the i-th (i∈[1,n]) dimension i (0﹤a i ﹤1,∑ n a i =1), according to the importance score a i Allocate the second privacy budget a to n dimensions of multidimensional data i *ε'.

[0126] In one implementation, S23, receiving second disturbance data from multiple clients, specifically includes:

[0127] Second perturbation data is received from the m clients, the second perturbation data is published, and the second perturbation data is used to solve a classification problem or a regression problem.

[0128] Embodiment 3:

[0129] like Figure 3 As shown, the present disclosure provides a data privacy iterative optimization device, the device is a client, the client is used for data privacy iterative optimization, and includes:

[0130] A first noise adding module 11 is used to perform a first noise adding on the data of each dimension of the first local multidimensional data according to the first privacy budget allocated to each dimension of the multidimensional data, so as to obtain first disturbance data of the first local multidimensional data;

[0131] The privacy acquisition module 12 is connected to the first noise adding module 11, and is used to send the first disturbance data to the server, so that the server allocates a second privacy budget for each dimension of the multidimensional data according to the first disturbance data from multiple clients, and sends the second privacy budget to the client;

[0132] The second noise addition module 13 is connected to the privacy acquisition module 12, and is used to perform a second noise addition on the data of each dimension of the second local multidimensional data according to the second privacy budget of each dimension, so as to obtain second disturbance data of the second local multidimensional data, and send the second disturbance data to the server.

[0133] In one embodiment, wherein:

[0134] The client selects different noise adding mechanisms to perform the first noise adding and the second noise adding for the data of the corresponding dimension according to whether each dimension corresponds to a discrete data attribute or a continuous data attribute.

[0135] The server uses a decision tree model to obtain the importance score of each dimension of the multidimensional data according to the first perturbation data from multiple clients, and allocates a second privacy budget according to the importance score.

[0136] In one implementation, the first noise adding module 11 specifically includes:

[0137] An initial budget unit, used to obtain a first total privacy budget ε initially preset for multi-dimensional data;

[0138] The privacy equalization unit is connected to the initial budget unit and is used to allocate the first privacy budget ε / n to each dimension attribute according to the number of dimensions n of the multidimensional data;

[0139] a discrete noise adding unit, connected to the privacy averaging unit, for performing a first noise adding on the data of the first dimension of the first local multidimensional data using a k-random response mechanism according to ε / n in response to the data attribute of the first dimension of the multidimensional data being a discrete data attribute;

[0140] a continuous noise adding unit, connected to the privacy averaging unit, and configured to perform a first noise adding on the data of the second dimension of the first local multidimensional data using a Laplace mechanism according to ε / n in response to the data attribute of the second dimension of the multidimensional data being a continuous data attribute;

[0141] The comprehensive noise addition unit, connected to the discrete noise addition unit and the continuous noise addition unit, is used to comprehensively process the results of the first noise addition for each dimension of data to obtain the first perturbed data of the first local multi-dimensional data.

[0142] In one embodiment, the second noise addition module 13 specifically includes:

[0143] The privacy receiving unit is used to receive the second partial privacy budget a i *ε’ of n dimensions from the server, a i (0 < a i < 1, ∑ n a i = 1) is the importance score of the i-th (i ∈ [1, n]) dimension of the multi-dimensional data, and ε’ is the second total privacy budget;

[0144] The discrete noise addition unit, connected to the privacy receiving unit, is used to respond that the data attribute of the first dimension of the multi-dimensional data is a discrete data attribute, and according to a i *ε’ use the k-random response mechanism to perform the second noise addition on the data of the first dimension of the second local multi-dimensional data;

[0145] The continuous noise addition unit, connected to the privacy receiving unit, is used to respond that the data attribute of the second dimension of the multi-dimensional data is a continuous data attribute, and according to a i *ε’ use the Laplace mechanism to perform the second noise addition on the data of the second dimension of the second local multi-dimensional data;

[0146] The comprehensive noise addition unit, connected to the discrete noise addition unit and the continuous noise addition unit, is used to comprehensively process the results of the second noise addition for each dimension of data to obtain the second perturbed data of the second local multi-dimensional data;

[0147] The data sending unit, connected to the comprehensive noise addition unit, is used to send the second perturbed data to the server, so that the server publishes the second perturbed data and uses the second perturbed data to solve the classification problem or the regression problem.

[0148] Example 4:

[0149] As Figure 4 shown, the present disclosure provides a data privacy iterative optimization device, and the device is a server, and the server is used for data privacy iterative optimization and includes:

[0150] The first data module 21 is used to receive the first perturbed data from multiple clients, and the first perturbed data is obtained by the clients respectively performing the first noise addition on the data of each dimension of the first local multi-dimensional data according to the first partial privacy budget assigned to each dimension of the multi-dimensional data;

[0151] The privacy allocation module 22 is connected to the first data module 21, and is used to allocate a second privacy budget for each dimension of the multidimensional data according to the first perturbation data from multiple clients, and send the second privacy budget to the client, so that the client performs a second noise addition on the data of each dimension of the second local multidimensional data according to the second privacy budget of each dimension, so as to obtain second perturbation data of the second local multidimensional data;

[0152] The second data module 23 is connected to the privacy allocation module 22 and is used to receive second disturbance data from multiple clients.

[0153] In one embodiment, wherein:

[0154] The client selects different noise adding mechanisms to perform the first noise adding and the second noise adding for the data of the corresponding dimension according to whether each dimension corresponds to a discrete data attribute or a continuous data attribute.

[0155] The server uses a decision tree model to obtain the importance score of each dimension of the multidimensional data according to the first perturbation data from multiple clients, and allocates a second privacy budget according to the importance score.

[0156] In one implementation, the privacy allocation module 22 specifically includes:

[0157] A problem type unit, used to obtain a preset multi-dimensional data problem analysis type;

[0158] A model selection unit is connected to the problem type unit and is used to select the decision tree model as a classification tree model or a regression tree model according to the type of multidimensional data problem analysis, which is to solve a classification problem or a regression problem using multidimensional data;

[0159] A decision tree analysis unit, connected to the model selection unit, for inputting the first disturbance data from the m clients into the classification tree model or the decision tree model to obtain a minimized impurity of the data of each dimension split in the classification tree model or a minimized mean square error of the data split in the decision tree model;

[0160] The privacy allocation unit is connected to the decision tree analysis unit and is used to obtain the importance score a of the i-th dimension (i∈[1,n]) according to minimizing the impurity or minimizing the mean square error and the second total privacy budget ε'. i (0﹤a i ﹤1,∑ n a i =1), according to the importance score a i Allocate the second privacy budget a to n dimensions of multidimensional data i *ε'.

[0161] In one implementation, the second data module 23 specifically includes:

[0162] A data receiving unit, configured to receive second disturbance data from m clients;

[0163] The data publishing and using unit is connected to the data receiving unit and is used to publish the second disturbance data and use the second disturbance data to solve the classification problem or the regression problem.

[0164] Embodiment 5:

[0165] like Figure 5 As shown, the present disclosure provides a data privacy iterative optimization device, the device is a data privacy iterative optimization system, and the system includes:

[0166] The multiple clients as described in Example 3 are used to implement the data privacy iterative optimization method as described in Example 1;

[0167] The server as described in Example 4 is connected to multiple clients to implement the data privacy iterative optimization method as described in Example 2.

[0168] Embodiment 6:

[0169] Embodiment 6 of the present disclosure provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the data privacy iterative optimization method as described in Embodiment 1 or 2 is implemented, or the data privacy iterative optimization device as described in any one of Embodiments 3-5 is implemented.

[0170] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program elements or other data). Computer-readable storage media include, but are not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable read only memory), flash memory or other memory technology, CD-ROM (Compact Disc Read-Only Memory), digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.

[0171] In addition, the present disclosure may also provide a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes the data privacy iterative optimization method as described in Embodiment 1 or 2. The computer device may be the data privacy iterative optimization device as described in any one of Embodiments 3-5.

[0172] The memory is connected to the processor, the memory may be a flash memory or a read-only memory or other memory, and the processor may be a central processing unit or a single-chip microcomputer.

[0173] Embodiments 1-6 of the present disclosure provide a method, device and medium for iterative optimization of data privacy, wherein the first noise addition is performed on the client, and the server integrates the data after the first noise addition from multiple clients, allocates a privacy budget for each dimension of the data, and the client performs a second data noise addition using the privacy budget allocated by the server, and obtains the privacy allocation for each data dimension based on the overall data of multiple clients, thereby achieving a reasonable allocation of the privacy budget, and further achieving fine control of the noise level and high availability of data quality, that is, providing data with good privacy and high quality by optimizing privacy noise twice.

[0174] It is to be understood that the above embodiments are merely exemplary embodiments used to illustrate the principles of the present disclosure, but the present disclosure is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and substance of the present disclosure, and these modifications and improvements are also considered to be within the scope of protection of the present disclosure.

Claims

1. A data privacy iterative optimization method, characterized in that: The method is applied to a client and comprises: According to the first privacy budget allocated to each dimension of the multidimensional data, respectively, performing a first noise addition on the data of each dimension of the first local multidimensional data to obtain first perturbed data of the first local multidimensional data; Sending the first perturbation data to the server, so that the server allocates a second privacy budget for each dimension of the multi-dimensional data according to the first perturbation data from the multiple clients, and sends the second privacy budget to the client; According to the second privacy budget of each dimension, the data of each dimension of the second local multidimensional data is denoised for a second time to obtain second perturbed data of the second local multidimensional data, and the second perturbed data is sent to the server.

2. The method according to claim 1, characterized in that in: The client selects different noise adding mechanisms to perform the first noise adding and the second noise adding for the data of the corresponding dimension according to whether each dimension corresponds to a discrete data attribute or a continuous data attribute. The server uses a decision tree model to obtain the importance score of each dimension of the multidimensional data according to the first perturbation data from multiple clients, and allocates a second privacy budget according to the importance score.

3. The method according to claim 2, characterized in that According to the first privacy budget allocated to each dimension of the multidimensional data, the data of each dimension of the first local multidimensional data is firstly denoised to obtain first perturbed data of the first local multidimensional data, specifically including: Obtain a first total privacy budget ε initially preset for the multidimensional data; Allocate the first privacy budget ε / n to each dimension attribute according to the number of dimensions n of the multidimensional data; In response to the data attribute of the first dimension of the multidimensional data being a discrete data attribute, performing a first noise addition on the data of the first dimension of the first local multidimensional data using a k-random response mechanism according to ε / n; In response to the data attribute of the second dimension of the multidimensional data being a continuous data attribute, performing a first noise addition on the data of the second dimension of the first local multidimensional data using a Laplace mechanism according to ε / n; The results of the first noise addition of the data in each dimension are integrated to obtain the first disturbance data of the first local multi-dimensional data.

4. The method according to claim 2 or 3, characterized in that: According to the second privacy budget of each dimension, the data of each dimension of the second local multidimensional data is subjected to a second noise addition to obtain second perturbation data of the second local multidimensional data, and the second perturbation data is sent to the server, specifically including: Receive the second privacy budget a of n dimensions from the server i *ε',a i (0﹤a i ﹤1,∑ n a i =1) is the importance score of the i-th (i∈[1,n]) dimension of the multidimensional data, and ε' is the second total privacy budget; In response to the data attribute of the first dimension of the multidimensional data being a discrete data attribute, according to a i *ε' uses a k-random response mechanism to perform a second noise addition on the data of the first dimension of the second local multidimensional data; In response to the data attribute of the second dimension of the multidimensional data being a continuous data attribute, according to a i *ε' uses the Laplace mechanism to perform a second noise addition on the data of the second dimension of the second local multidimensional data; The result after the second noise addition of the data in each dimension is synthesized to obtain the second disturbance data of the second local multi-dimensional data; The second perturbation data is sent to the server, so that the server publishes the second perturbation data, and the second perturbation data is used to solve the classification problem or the regression problem.

5. A data privacy iterative optimization method, characterized in that: The method is applied to a server and comprises: Receiving first perturbed data from multiple clients, where the first perturbed data is obtained by the clients performing a first noise addition on data of each dimension of first local multidimensional data according to a first privacy budget allocated to each dimension of the multidimensional data; Allocate a second sub-privacy budget for each dimension of the multidimensional data according to the first perturbation data from the multiple clients, and send the second sub-privacy budget to the client, so that the client performs a second noise addition on the data of each dimension of the second local multidimensional data according to the second sub-privacy budget of each dimension, so as to obtain second perturbation data of the second local multidimensional data; Second perturbation data is received from a plurality of clients.

6. The method according to claim 5, characterized in that in: The client selects different noise adding mechanisms to perform the first noise adding and the second noise adding for the data of the corresponding dimension according to whether each dimension corresponds to a discrete data attribute or a continuous data attribute. The server uses a decision tree model to obtain the importance score of each dimension of the multidimensional data according to the first perturbation data from multiple clients, and allocates a second privacy budget according to the importance score.

7. The method according to claim 6, characterized in that According to the first perturbed data from multiple clients, a second privacy budget is allocated to each dimension of the multi-dimensional data, specifically including: Get the preset multidimensional data problem analysis type; According to the type of multidimensional data problem analysis, which is to use multidimensional data to solve classification problems or regression problems, the decision tree model is selected as a classification tree model or a regression tree model; Input the first perturbed data from the m clients into the classification tree model or the decision tree model to obtain the minimized impurity of the data of each dimension split in the classification tree model or the minimized mean square error of the data split in the decision tree model; According to minimizing impurity or minimizing mean square error, and the second total privacy budget ε', obtain the importance score a of the i-th (i∈[1,n]) dimension i (0﹤a i ﹤1,∑ n a i =1), according to the importance score a i Allocate the second privacy budget a to n dimensions of multidimensional data i *ε'.

8. The method according to claim 7, characterized in that Receiving second disturbance data from multiple clients, specifically comprising: Second perturbation data is received from the m clients, the second perturbation data is published, and the second perturbation data is used to solve a classification problem or a regression problem.

9. A client, characterized in that: The client is used for iterative optimization of data privacy and includes: A first noise adding module, configured to perform a first noise adding on the data of each dimension of the first local multidimensional data according to the first privacy budget allocated to each dimension of the multidimensional data, so as to obtain first disturbance data of the first local multidimensional data; a privacy acquisition module, connected to the first noise adding module, and configured to send the first disturbance data to the server, so that the server allocates a second privacy budget for each dimension of the multidimensional data according to the first disturbance data from the plurality of clients, and sends the second privacy budget to the client; The second noise addition module is connected to the privacy acquisition module, and is used to perform a second noise addition on the data of each dimension of the second local multidimensional data according to the second privacy budget of each dimension, so as to obtain second disturbance data of the second local multidimensional data, and send the second disturbance data to the server.

10. A server, characterized in that: The server is used for iterative optimization of data privacy and includes: A first data module is used to receive first perturbed data from multiple clients, where the first perturbed data is obtained by the client performing a first noise addition on data of each dimension of first local multidimensional data according to a first privacy budget allocated to each dimension of the multidimensional data; a privacy allocation module connected to the first data module, configured to allocate a second privacy budget to each dimension of the multidimensional data according to the first perturbation data from the plurality of clients, and send the second privacy budget to the client, so that the client performs a second noise addition on the data of each dimension of the second local multidimensional data according to the second privacy budget of each dimension, so as to obtain second perturbation data of the second local multidimensional data; The second data module is connected to the privacy allocation module and is used to receive second disturbance data from multiple clients.

11. A data privacy iterative optimization system, characterized in that: The system comprises: Multiple clients, used to implement the data privacy iterative optimization method according to any one of claims 1 to 4; A server, connected to multiple clients, is used to implement the data privacy iterative optimization method as described in any one of claims 5-8.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the data privacy iterative optimization method as described in any one of claims 1-4 or 5-8 is implemented.