A Differential Privacy Data Publishing Method Satisfying Personalized Privacy Budget Allocation
The personalized privacy budget allocation method addresses the challenge of protecting sensitive healthcare data by classifying attributes, clustering with a modified k-prototype algorithm, and applying differential privacy to ensure both privacy and data usability.
Patent Information
- Application Number
- CN202210137381.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-02-15
AI Technical Summary
The existing differential privacy data publishing methods are difficult to balance between protecting user privacy and data availability, especially when facing background knowledge attacks and combination attacks, it is impossible to effectively protect sensitive information. At the same time, there is an imbalance in the allocation of privacy budgets, resulting in low data availability or insufficient privacy protection.
By grading the attributes in the dataset at a sensitive level, grouping them using mutual information and attribute association relationships, and combining the optimal matching theory to build a privacy budget division diagram, matching the corresponding privacy budgets for sensitive attributes at different levels, improving the k-prototype algorithm for clustering, differential privacy protection is carried out for clustering centers, and generating data sets to be published that meet differential privacy.
It realizes the protection and release of patient data in health institutions, which not only ensures user privacy, but also meets the analysis needs of researchers, improves the balance of data availability and privacy protection, and reduces the risk of data leakage.
Smart Images

Figure CN114491644B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a differential privacy data publishing method, in particular to a differential privacy data publishing method that satisfies personalized privacy budget allocation. Background Art
[0002] With the advent of the digital information era, the publishing and utilization of data have become particularly important. The government and relevant departments can use data resources to provide scientific decision-making plans and predict market trends. However, the data to be published usually contains a large amount of sensitive information, and directly publishing such data will inevitably lead to the leakage of user privacy information. How to maximize the data availability while ensuring that the user's sensitive information is not leaked is the key to studying the data publishing problem. In recent years, researchers have proposed some methods for privacy protection in data publishing, mainly including data anonymization publishing methods and data distortion publishing methods.
[0003] k-anonymity is a typical representative of data anonymization publishing methods. It groups and generalizes the quasi-identifier attributes of the data to be published, making each record indistinguishable from at least k-1 other records, thereby protecting data privacy. In response to the defects in the k-anonymity model, subsequent l-diversity and t-closeness methods have been proposed. Although anonymization methods can protect privacy information in data to a certain extent, these methods are effective only on the premise that the attacker does not have any background knowledge and cannot resist background knowledge attacks and combined attacks. Differential privacy, as a data distortion-based publishing method, has been widely studied and applied. Because it makes no assumptions about the background knowledge of the attacker, it adds a certain amount of available noise to the data to be published for perturbation, thereby providing strong privacy guarantees. For example, medical researchers can obtain the general distribution and clinical manifestations of diseases by clustering medical big data, so as to better diagnose and treat diseases and study the causes of diseases. However, the data after clustering often contains a large amount of personal privacy of patients, and if not properly processed, it is easy to be maliciously analyzed by attackers, resulting in the leakage of patients' privacy information. Summary of the Invention
[0004] The present invention designs and develops a differential privacy data publishing method that satisfies personalized privacy budget allocation. By grading the sensitive levels of attributes in the dataset, allocating privacy budgets to different levels of sensitive attributes, and then clustering the original dataset through an improved k-prototype algorithm, differential privacy protection is performed on the clustering centers in each cluster to generate a dataset to be published that satisfies differential privacy and meets various query requirements. This method can protect and publish patient data in a health institution, and can protect user privacy and provide researchers with analysis and use while maximizing data availability.
[0005] The technical solution provided by the present invention is as follows:
[0006] A differential privacy data publishing method that meets personalized privacy budget allocation, including:
[0007] Step 1: Classify sensitive attributes through mutual information and attributes according to the sensitivity levels of attributes in the dataset;
[0008] Step 2: Construct a bipartite graph for privacy budget division, and match corresponding privacy budgets to sensitive attributes at different levels;
[0009] Step 3: Cluster the original dataset, and perform differential privacy protection on the cluster center values through the privacy budgets allocated to sensitive attributes to generate a dataset to be published that meets differential privacy.
[0010] Preferably, in Step 1, according to the degree of association of mutual information between attributes in the dataset, they are divided into: high-sensitivity attribute group, medium-sensitivity attribute group, and low-sensitivity attribute group.
[0011] Preferably, Step 2 includes:
[0012] Set the privacy loss P of each original sensitive attribute i to 0, calculate the privacy consumption C corresponding to each privacy budget ij and the information loss value l of each sensitive attribute S :
[0013] C ij = ε j × S i ;
[0014] l S = C ij - P i ;
[0015] Among them, the privacy protection intensity ratio between the high-sensitivity attribute group, medium-sensitivity attribute group, and low-sensitivity attribute group is 5:3:2;
[0016] Construct a privacy budget division graph through a loss function;
[0017] When there is a privacy budget that optimally matches the sensitive attribute in the privacy budget division graph, output the relationship graph between the sensitive attribute and the privacy budget;
[0018] When there is no privacy budget that optimally matches the sensitive attribute in the privacy budget division graph, increase the privacy loss of a sensitive attribute and continue to construct the graph for matching until an optimal matching result is obtained.
[0019] Preferably, Step 3 includes:
[0020] In dataset O, the number of clusters is k, and the privacy budget is {ε1, ε2......ε m}
[0021] Select the initial centers and calculate the local density ρ of each tuple record O i i
[0022]
[0023] Calculate the distance δ between the records whose local density is higher than that of tuple record O i and the record that is the closest to tuple record O j i :
[0024]
[0025] The ρ i and δ i of the cluster centers need to satisfy the formula:
[0026] Z p =(o i |δ i >μ(δ), ρ i >μ(ρ), 1≤i≤n}
[0027] where μ(δ) and μ(ρ) represent the means of all tuple records ρ i and δ i ;
[0028] Calculate a comprehensive parameter α of ρ p and δ i for each tuple record in z i by formula (13). i ,
[0029] The larger this value (specifically, the value of which parameter) is, the closer the attribute value of this tuple record is to becoming a cluster center;
[0030] Sort all tuple records by α i from largest to smallest, and the tuple record with the largest α i is the cluster center;
[0031] For other preselected center points, calculate the distances of the tuple records before it in order,
[0032] When the distance is greater than 2d l , then this tuple record can be used as an initial cluster center, and select the first k tuple records as the initial cluster centers,
[0033]
[0034] For numerical attributes, add Laplace noise to the mean of each attribute in the cluster:
[0035]
[0036] Generate a differentially private dataset o'.
[0037] Preferably, the third step further includes:
[0038] When the clustering loss function changes, 1 ≤ i ≤ n, calculate d(o i , o j );
[0039] The loss function is:
[0040]
[0041] Assign o i to the cluster with the minimum dissimilarity
[0042] Recalculate the center points within the cluster and calculate the central value of each attribute
[0043] Calculate the value of E i , until the value of E i no longer changes, and obtain the clustering result.
[0044] The beneficial effects of the present invention are as follows:
[0045] 1. A sensitivity grading method is proposed using mutual information and the correlation relationship between attributes for different sensitive levels of attributes in the dataset, enabling users to quantify the degree of attention to sensitive attributes and matching corresponding privacy protection levels for attributes at different levels;
[0046] 2. Combining the optimal matching theory, a bipartite graph for privacy budget division is constructed to achieve the allocation of privacy budgets for sensitive attributes at different levels, which better solves the problems of low data availability and insufficient protection of sensitive attributes caused by the average allocation of privacy budgets.
[0047] 3. By improving the dissimilarity measurement method and the selection of the initial center, a new k-prototype algorithm is proposed to cluster the original dataset and perform differential privacy protection on the clustering centers in each cluster, generating a dataset to be published that satisfies differential privacy and meets various query requirements.
[0048] 4. By using this method, it is possible to specifically count the number of patients suffering from a specific disease in a medical institution and send it to the medical institution, which not only ensures that user data is not fully disclosed but also meets the statistical needs of the medical institution. Brief Description of the Drawings
[0049] Figure 1(a) is a graph showing the variation of the sum of squared errors with different numbers of clustering clusters when ε is 0.01 on the Adult dataset.
[0050] Figure 1(b) is a graph showing the variation of the sum of squared errors with different numbers of clustering clusters when ε is 0.1 on the Adult dataset.
[0051] Figure 1(c) is a graph showing the variation of the sum of squared errors with different numbers of clustering clusters when ε is 1 on the Adult dataset.
[0052] Figure 1(d) is a graph showing the variation of the sum of squared errors with different numbers of clustering clusters when ε is 5 on the Adult dataset.
[0053] Figure 2(a) is a graph showing the variation of the record association value with different numbers of clustering clusters when ε is 0.01 on the Adult dataset
[0054] Figure 2(b) is a graph showing the variation of the record association value with different numbers of clustering clusters when ε is 0.1 on the Adult dataset.
[0055] Figure 2(c) is a graph showing the variation of the record association value with different numbers of clustering clusters when ε is 1 on the Adult dataset.
[0056] Figure 2(d) is a graph showing the variation of the record association value with different numbers of clustering clusters when ε is 5 on the Adult dataset.
[0057] Figure 3(a) is a graph showing the variation of the sum of squared errors with different numbers of clustering clusters when ε is 1 on the CI dataset.
[0058] Figure 3(b) is a graph showing the variation of the record association value with different numbers of clustering clusters when ε is 1 on the CI dataset.
[0059] Figure 4 is the initial privacy budget bipartite graph.
[0060] Figure 5 is the attribute generalization hierarchy graph. Detailed implementation manner
[0061] The following further elaborates on the present invention with reference to the accompanying drawings, so that those skilled in the art can implement it according to the description in the specification.
[0062] As shown in Figures 1-5, the present invention provides a differential privacy data publishing method that meets personalized privacy budget allocation. By grading the sensitivity levels of attributes in the dataset, privacy budgets are allocated to different levels of sensitive attributes. Then, the improved k-prototype algorithm is used to cluster the original dataset, and differential privacy protection is performed on the cluster centers in each cluster to generate a dataset to be published that meets differential privacy, satisfying various query requirements, including:
[0063] Step 1: According to the different sensitivity levels of attributes in the dataset, sensitive attributes are graded through the mutual information and the correlation relationship between attributes;
[0064] Step 2: Combining the optimal matching theory, a bipartite graph for privacy budget division is constructed to match corresponding privacy budgets to sensitive attributes at different levels;
[0065] Step 3: Use the improved k-prototype algorithm to cluster the original dataset, and perform differential privacy protection on the cluster center values through the privacy budgets allocated to sensitive attributes to generate a dataset to be published that meets differential privacy.
[0066] Among them, Step 1 specifically includes:
[0067] The importance of each sensitive attribute in the dataset is different, and the intensity of protection required for it also varies. Therefore, it is necessary to perform a grading operation on sensitive attributes; use the size of the mutual information between attributes to compare their correlation degrees; divide them according to the correlation degrees, and group the attributes with higher correlation degrees in the same group. If there is an overlap in the attribute sensitivity between sensitive attributes, for example: the size relationship of the mutual information between sensitive attribute tuples x1 and x2, x1 and x3, x2 and x3 is I(x1,x2) > I(x1,x3) > I(x2,x3), then they are divided into one group; after grouping, calculate the average attribute sensitivity of each group, and divide the attribute levels according to the size of the sensitivity. Select the first 1 / 3 of the groups with the largest average sensitivity as the high-sensitive attribute group (HAS), the middle 1 / 3 of the groups as the medium-sensitive attribute group (MAS), and the remaining 1 / 3 of the groups as the low-sensitive attributes (LAS). The pseudo-code of its algorithm is as follows:
[0068] Algorithm 1:
[0069] Input: Sensitive attribute r in dataset O i ;
[0070] Output: Graded sensitive attribute groups P1, P2, P3;
[0071] 1.For each r i in O:
[0072] 2. Calculate r iwith r j I(X:Y) between;
[0073] 3. End for
[0074] 4. Compare r i with r j correlation degree
[0075] 5. If the size relationship of the mutual information between sensitive attribute tuples is: I(r i : r j ) > I(r i : r k ) > I(r m : r n )
[0076] 6. Divide the sensitive attributes r i and r j into P1;
[0077] r k into P2;
[0078] r m and r n into P3;
[0079] 7. For P1, P2, P3:
[0080] 8. Calculate the average attribute sensitivity in each group and denote it as
[0081] 9. If and P1 belongs to the first 1 / 3 group, P2 belongs to the middle 1 / 3 group, P3 belongs to the last 1 / 3 group;
[0082] 10. P1 is HAS; P2 is MAS; P3 is LAS;
[0083] 11. End for
[0084] 12. Return P1, P2, P3.
[0085] Step 2 includes: personalized privacy budget allocation scheme
[0086] Combined with the optimal matching theory, allocate privacy budgets for sensitive attributes. For the protection intensity of sensitive attributes, subjective human settings may cause over - or under - setting of the corresponding privacy protection intensity. Optimize it according to the sensitive attribute grading strategy. After automatically grading the sensitive attributes in the dataset, then set their protection intensity, and finally use the optimal matching theory to allocate privacy budgets for sensitive attributes at each level.
[0087] Define the following parameters:
[0088] Privacy consumption: Assume that the privacy protection strength of the sensitive attribute level is S i , and a privacy budget ε is randomly assigned to each sensitive attribute j , then the privacy consumption of this sensitive attribute is calculated as follows: C ij = ε j × S i
[0089] Optimal matching: Construct a bipartite graph (x, y). If there is a maximum number of matches between x and y, and |x| = |y| = the number of matches, then this scheme is called the optimal matching of the bipartite graph;
[0090] Privacy budget partition graph: Under the condition of satisfying differential privacy, link the privacy budget with the largest mutual information and the sensitive attribute level that makes the data to be released and the already released data satisfy differential privacy protection, and the constructed bipartite graph is called the privacy budget partition graph;
[0091] As Figure 4 shown, given a set of privacy budgets ε j , and a set of sensitive attributes. Assume that the ID number and disease are HAS, then a higher protection strength is required for them, a larger amount of noise needs to be added, and the allocated privacy budget is less, so the privacy budget allocated to them is {0.01, 0.03}; the home address and exam scores are MAS, and the privacy budget allocated to them is {0.2, 0.35}; the postal code and education level are LAS, and a small amount of noise can be added to the low-sensitive attributes, so the privacy budget allocated to them is {1.2, 1.35}. Construct an initial privacy budget partition graph, where ls represents the information loss of each sensitive attribute. One ε j corresponds to multiple sensitive attributes. When this ε j matches a certain sensitive attribute, if the calculated ls is the smallest or no longer changes, then this ε j and the corresponding sensitive attribute are an optimal match. When all ε j are connected to the sensitive attributes and the overall ls is the smallest, the bipartite graph formed by the sensitive attributes and the privacy budget at this time is the optimal privacy budget partition graph.
[0092] Specific matching idea: Given a set of sensitive attributes with n sensitive levels and a set containing n random privacy budgets, match the sensitive attribute levels with the privacy budgets that maximize the availability of the published data. The algorithm process is as follows:
[0093] Algorithm 2:
[0094] Input: Sensitive attribute r in dataset O i , privacy protection strength S of each sensitive attribute i , a set of different privacy budgets ε j ;
[0095] Output: Optimal matching graph of sensitive attributes and privacy budget;
[0096] 1. For each r i in O:
[0097] 2. Set the privacy loss P i = 0;
[0098] 3. For each privacy budget ε j :
[0099] 4. Calculate the privacy consumption C ij ;
[0100] 5. ls = C ij - P i ;
[0101] 6. End for;
[0102] 7. End for;
[0103] 8. Construct the initial privacy budget partition graph PM;
[0104] 9. If the optimal matching exists:
[0105] 10. End;
[0106] 11. Then this PM is the optimal matching result graph of the sensitive level and the privacy budget;
[0107] 12. Else:
[0108] 13. P i + 1;
[0109] 14. Return to 8;
[0110] 15. End.
[0111] Steps 1 - 7, set the privacy loss P of each original sensitive attribute i to 0, calculate the privacy consumption corresponding to each privacy budget and the ls value of each sensitive attribute. ls represents the information loss of each sensitive attribute, where the privacy protection strength ratio among HAS, MAS, and LAS is 5:3:2; Steps 8 - 15 use the loss function to construct the privacy budget partition graph and check if there is an optimal matching in the graph. If there is, output the optimal matching relationship graph between the sensitive attributes and the privacy budget. Otherwise, increase the P of this sensitive attribute i by one unit, return to Step 8, continue to construct the graph for matching until the optimal matching result is obtained.
[0112] Step 3 includes: A differential privacy data publishing algorithm based on improved k-prototype clustering
[0113] As a classic clustering algorithm for processing mixed-attribute datasets, the k-prototype algorithm determines the initial clustering centers artificially or randomly and uses the Euclidean distance and the Hemingway formula to simply measure the distance between attribute values, reducing the stability and accuracy of the clustering algorithm. Improve the selection of its initial clustering centers and the formula for calculating the dissimilarity of attribute values to enhance the clustering effect and thereby improve the usability of the data after differential privacy protection.
[0114] Method for measuring the dissimilarity between tuple attribute values
[0115] 1. Measurement of dissimilarity for numerical attributes
[0116] Given a dataset O containing n data records and m-dimensional attributes, where k are numerical attributes and l are categorical attributes. Then the formula for calculating the information entropy of the p-th (1 ≤ p ≤ k) dimensional numerical attribute is as follows:
[0117]
[0118] Among them, C ip represents the number of tuple records o i containing the p-th dimensional numerical attribute, then the weight of this numerical attribute is Then the formula for measuring the dissimilarity of the p-th dimensional numerical attribute between data records o h and o j is: d r (o h , o j ) = w p (o hp - o ip ) 2 .
[0119] 2. Measurement of dissimilarity for categorical attributes
[0120] For the q-th dimensional categorical attribute, the information entropy H q calculation formula is:
[0121]
[0122] Among them, C tq = n tq / n represents the proportion of the t-th categorical attribute value in the q-th dimensional categorical attribute, and n q represents the number of the q-th dimensional categorical attribute.
[0123] Then the weight value y q of this categorical attribute is:
[0124]
[0125] Among them, H q is the information entropy of the q-th dimensional classification attribute.
[0126] The classification attribute is not a specific value. Therefore, a generalization hierarchy tree needs to be constructed for the categorical attribute to measure the dissimilarity between attribute values. For example, for the sensitive attribute 'disease', its generalization hierarchy tree is constructed as Figure 5 shown.
[0127] Then, for the data record o h and o j the dissimilarity measure formula for the q-th dimensional classification attribute between them is;
[0128]
[0129] Among them, |o hq | and |o jq | are respectively the total number of attribute values of the q-th dimensional classification attribute; and are respectively the number of leaf nodes contained in the sub-tree generalized to the same level t h and o j in o q .
[0130] For each cluster, for numerical attributes, the mean value V l r of the attribute values is selected as its clustering center, and for categorical attributes, the value V l c with the highest frequency of occurrence among the attribute values is selected as its clustering center. Then, the formula for calculating the mixed dissimilarity between data records is:
[0131]
[0132] k-prototype clustering needs to use a suitable loss function to measure the distance of numerical and categorical variables to the clustering center. If the distance is appropriate, the clustering result is the optimal and the clustering ends. Assume that the data set O contains n data records, the number of clusters is k, and y il takes values of 0 or 1, indicating whether the tuple i exists in the l-th cluster. If it exists, the value is 1, otherwise it is 0. This loss function is:
[0133]
[0134] Initial clustering center selection
[0135] For the selection of the initial center, attribute values with relatively balanced dissimilarity to each data record should be selected as much as possible. The local density and high-density distance definitions in the data distribution are introduced to select the initial center.
[0136] Define two parameters:
[0137] 1) The local density of each tuple record o i and other tuples o j is calculated using the Gaussian kernel method:
[0138]
[0139] 2) The distance between each tuple record o i and the tuple o j with a higher local density and the closest distance:
[0140] where d(o i , o j ) is the distance between tuple records d i and d j , d l is the truncation distance, that is, the value of the first 1% to 2% after sorting all tuple record distances from smallest to largest.
[0141] To measure the distance between tuples efficiently and accurately, define the distance formula between tuple records as:
[0142] dis(o i , o j ) = dis r (o i , o j ) + dos c (o i , o j );
[0143] Among them, the numerical attribute distance is calculated using the weighted Euclidean distance:
[0144]
[0145] The categorical attribute distance is calculated using the weighted Hamming distance:
[0146]
[0147]
[0148] The local density of the cluster center in its respective cluster should be the largest, and the distance between the cluster centers of different clusters should be as far as possible, ensuring a high dissimilarity between clusters and a high homogeneity within clusters after clustering. The ρ i and δ i of the cluster center need to satisfy the formula:
[0149] Z p = {o i | δ i> μ(δ), ρ i > μ(ρ), 1 ≤ i ≤ n}
[0150] Wherein, Z p is the set of preselected clustering centers, and μ(ρ) and μ(δ) represent the means of all tuple records ρ i and δ i .
[0151] To prevent multiple centers from appearing in the same cluster, the distance between the centers should be specified, and a comprehensive ρ p of each tuple record in Z i and δ i is calculated using the formula, and the parameter α i of ρ i and δ i is obtained. The larger α i , the more likely the attribute value of the tuple record is to become a clustering center; sort all tuple records in descending order of α l . The tuple record with the largest α
[0152]
[0153] Noise addition mechanism
[0154] To prevent the leakage of private information in the dataset after clustering, differential privacy protection is performed on the clustering centers:
[0155] a) For numerical attributes, Laplace noise is added to the mean of each attribute in the cluster:
[0156]
[0157] Wherein, is the average value of the p-th dimensional numerical attribute, and lap(Δf / ε) is the Laplace noise added to satisfy the privacy budget of size ε.
[0158] b) For categorical attributes, the exponential mechanism is used to probabilistically select the best clustering center in combination with the given privacy budget and scoring function to satisfy differential privacy protection.
[0159] Algorithms 1 and 2 allocate different privacy budgets to sensitive attributes, and use the improved k-prototype algorithm to cluster the original dataset. The clustered dataset is divided into k clusters, and the privacy budget allocated to each attribute in the cluster is 1 / k of the overall budget of the attribute. The data records in the cluster are added noise using the allocated privacy budget. The algorithm flow is as follows:
[0160] Algorithm Three:
[0161] Input: Dataset O, number of clusters k, privacy budget {ε1, ε2......ε m};
[0162] Output: Differentially private dataset o′;
[0163] 1. For each o i in O:
[0164] 2. Calculate ρ i and δ i ;
[0165] 3. If ρ i > μ(ρ) and δ i > μ(δ):
[0166] 4. o i is the initial center of the preselected cluster;
[0167] 5. For each o i :
[0168] 6. Calculate the α i value;
[0169] 7. If α i is max(α i ):
[0170] 8. Then o i must be the initial center of the cluster;
[0171] 9. If dis(o j , o i ) ≥ 2d l :
[0172] 10. Then o j is the initial center of the cluster;
[0173] 11. End for;
[0174] 12. End for;
[0175] 13. Initial cluster centers {o1, o2......o k};
[0176] 14. While the clustering loss function changes:
[0177] 15. For 1≤i≤n :
[0178] 16. Calculate d(o i , oj );
[0179] 17. Incorporate o i into the cluster with the minimum dissimilarity;
[0180] 18. Recalculate the center point within the cluster;
[0181] 19. Calculate the central value of each attribute;
[0182] 20. Calculate E i , until the value of E i no longer changes, and output the clustering result;
[0183] 21. End for;
[0184] 22. For 1 ≤ j ≤ k:
[0185] 23. For numerical attributes, calculate the mean of the attribute values as the clustering center of the attribute; for categorical attributes, generate the set of all attribute values;
[0186] 24. End for;
[0187] 25. For the clustering center value of numerical attributes, add Laplace noise to it using the privacy budget obtained from the division of each attribute; for categorical attributes, use the exponential mechanism to select the attribute value with the highest output probability as the center value;
[0188] 26. Generate the differentially private dataset o'.
[0189] Feasibility analysis of the algorithm:
[0190] Prove the usability of the DP-IMKP algorithm.
[0191] Define the query function f, which returns a record o in the dataset O i , and the query sensitivity can be effectively reduced by the k-prototype clustering algorithm.
[0192] Proof: The k-prototype clustering algorithm acts on the dataset O and divides it into k clusters, each of which contains at least n records. When the query function f acts on the dataset O, the difference between two data records is distributed among n records, and the query result returns the centroid of the cluster where a certain record is located, so the sensitivity is at most Δf / n. For the outliers in the clustering operation, since they are not assigned to a certain cluster, the overall mean of a certain attribute is used to replace them, and the query sensitivity will be less than that of the normal points assigned to the cluster. Therefore, when querying the dataset after k-prototype clustering, the sensitivity is less than or equal to Δf / n. Q.E.D.
[0193] The DP-IMKP algorithm satisfies ε-differential privacy.
[0194] Proof: Given two adjacent data sets O1 and O2, the output result sets after DP-IMKP calculation are M(O1) and M(O2) respectively. For the differential privacy data set, the query results of the query function f are f(O1) and f(O2) respectively. H and G are the results of all numerical attribute data and categorical attribute data output by the DP-IMKP algorithm on the adjacent data sets O1 and O2. For query results.
[0195] According to the definition of differential privacy, for numerical attributes, Then:
[0196]
[0197] For categorical attributes Then:
[0198]
[0199] The k-prototype clustering algorithm is used to divide the data set O into k non-overlapping subsets. According to the parallelism of differential privacy, the privacy budget ε allocated to each subset i is the overall privacy budget ε of the DP-IMKP algorithm. Therefore, DP-IMKP satisfies ε-differential privacy. Q.E.D.
[0200] Experimental Analysis
[0201] Experimental Environment and Data Sets
[0202] The experiment uses Python 3.8 as the development environment, Intel(R) Core(TM) i5-1135G7 @ 2.40 GHz, 16 GB of memory, and the operating system is Microsoft Windows 10. To verify the scalability of this algorithm, experiments are carried out on the Adult data set and the Census-Income data set respectively.
[0203] The Adult data set is a collection of US census statistical information data and is also one of the commonly used data sets in data analysis. It includes 14 numerical and categorical attributes. After removing the records containing missing attributes, there are 30,158 available records. The present invention selects 3 numerical attributes and 5 categorical attributes in the Adult data set as experimental attributes, as shown in Table 1:
[0204] Table 1 Adult Data Set
[0205]
[0206]
[0207] The Census-Income (hereinafter referred to as CI) dataset is the current population income survey conducted by the U.S. Census Bureau in 1994 and 1995, including 40 numerical and categorical attributes. After removing duplicate records, there are 299,285 available records in total. Using the MIndUtilmineset-to-mlc partitioning method, the present invention selects 99,762 of them as experimental records and selects 4 numerical attributes and 4 categorical attributes as experimental attributes, as shown in Table 2.
[0208] Table 2 CI dataset
[0209]
[0210] Experimental result measurement criteria
[0211] The experiment measures the data availability and the degree of data leakage risk as the standard for verifying the superiority of the algorithm.
[0212] Regarding the data availability, the sum of squared errors (SSE) is used as the measurement standard. It is used to calculate the sum of squared errors between the original data records and the data records after differential privacy protection. The smaller the SSE value, the less data information is lost. Its calculation formula is as (16):
[0213]
[0214] Among them, d(o i , o′ i ) represents the dissimilarity size between records. For the dissimilarity of numerical attributes in records, the Euclidean distance is used for calculation; for categorical attributes, the method proposed in the text is used for calculation.
[0215] The risk of data leakage is measured by record linkage (RL). This value represents the percentage of the dataset after differential privacy processing that matches the records in the original dataset. Its calculation formula is:
[0216]
[0217] Among them, o′ is the dataset after differential privacy protection, D is the set of original records with the smallest dissimilarity to o′. If the true original record o is in D, then pr(o′) represents the probability of guessing the original record o in D as 1 / |D|, where |D| is the number of records in D; if not, then pr(o′) = 0; the larger the RL, the higher the degree of association between the records in the dataset after differential privacy protection and the records in the original dataset, and the higher the risk of data leakage.
[0218] Analysis of Experimental Results
[0219] To verify the superiority of the DP-IMKP algorithm in improving data availability and reducing the risk of privacy leakage, on different datasets, it is compared with the DP-MDAV algorithm in the literature
[10] , the DP-k-prototype algorithm in the literature
[11] and the DCKPDP algorithm in the literature
[12] . For the mixed-attribute dataset, all three algorithms use different methods to perform clustering operations on the original dataset. Among them, after improving the classical MDAV algorithm, the DP-MDAV algorithm performs micro-aggregation on the dataset; DP-k-prototype uses the classical k-prototype algorithm to cluster the dataset; DCKPDP improves the initial clustering center and the dissimilarity calculation method of the k-prototype algorithm, clusters the mixed-attribute dataset, and finally performs differential privacy protection. By comparing the three algorithms, it can be proved that the DP-IMKP algorithm has the advantage of using the improved k-prototype algorithm and personalized partitioning of the privacy budget.
[0220] Set ε to {0.01, 0.1, 1, 5}. As the value of K changes, the SSE changes of each algorithm on the Adult dataset are shown in Figures 1(a)-(d).
[0221] It can be seen from Figure 1 that after performing differential privacy on the clustered dataset, the SSE value generally shows a downward trend as the number of clustering clusters increases, verifying the authenticity of Theorem 3. When the value of K reaches When (n is the total number of records in the dataset), the SSE value is almost optimal, and as the number of clustering clusters increases, the clustering result gradually approaches the optimization, and the SSE value also tends to be stable. When the value of ε is 0.01, due to the too small privacy budget, a large amount of noise is added to the dataset. Even if the clustering algorithm is used to process the original dataset, the data availability is still very low. When the value of ε is 0.1, the change range of SSE is the most obvious, but at this time the SSE value is relatively high, and the cost of information loss is large. When the values of ε are 1 and 5, the SSE values are relatively low, and as the privacy budget increases, the added noise decreases, and the information loss decreases, so the difference between the two is relatively small. However, no matter what the value of ε is, the SSE value of the proposed DP-IMKP algorithm is generally the lowest, that is, the data availability is the highest. This is because the DP-MDAV algorithm requires the data records to be sorted according to the distance to the boundary records, so the clustering generated by the MDAV algorithm is uneven, resulting in a large amount of information loss. In the process of clustering using the k-prototype algorithm, the DP-k-prototype algorithm artificially or randomly sets the initial clustering center and does not set weights for numerical attributes, resulting in poor clustering effect and large information loss. Although the DCKPDP algorithm improves the selection of clustering centers and the calculation method of dissimilarity, when performing differential privacy protection, it adopts the equal distribution principle for the privacy budget allocation of each attribute, which is likely to cause uneven and wasteful privacy budget allocation. The DP-IMKP algorithm adopts a hierarchical allocation of privacy budget for sensitive attributes, which can make full and reasonable use of the budget, minimize budget waste, and improve data availability. Setting ε as {0.01, 0.1, 1, 5}, as the number of clustering clusters changes, the sum of squared errors of each algorithm on the Adult dataset is shown in Figures 1(a)-(d).
[0222] As can be seen from Figure 1, for the differentially private clustered dataset, the sum of squared errors generally shows a downward trend as the number of clustering clusters increases. When the number of clustering clusters reaches When n (the total number of records in the dataset) is such that the sum of squared errors almost reaches the optimum, and as the number of clustering clusters increases, the clustering result gradually approaches the optimization, and the sum of squared errors also tends to be stable. When the value of ε is 0.01, due to the too small privacy budget, a large amount of noise is added to the dataset. Even if the original dataset is processed using the clustering algorithm, the data availability is still very low. When the value of ε is 0.1, the change range of the sum of squared errors is the most obvious, but at this time the sum of squared errors is relatively large, and the cost of information loss is large. When the values of ε are 1 and 5, the values of the sum of squared errors are low, and as the privacy budget becomes larger, the added noise decreases and the information loss decreases, so the difference between the two is small. However, no matter what value ε takes, the SSE value of the DP-IMKP method of the present invention is generally the lowest, that is, the data availability is the highest. This is because the DP-MDAV algorithm requires the data records to be sorted according to the distance to the boundary records, so the clustering generated by the MDAV algorithm is uneven, resulting in a large amount of information loss. In the process of clustering using the k-prototype algorithm by the DP-k-prototype algorithm, the initial clustering centers are set artificially or randomly, and no weights are set for the numerical attributes, resulting in poor clustering effect and thus large information loss. Although the DCKPDP algorithm improves the selection of clustering centers and the calculation method of dissimilarity, when performing differential privacy protection, the equal division principle is adopted for the privacy budget allocation of each attribute, which is likely to cause uneven and wasteful privacy budget allocation. The DP-IMKP algorithm adopts a hierarchical allocation of privacy budget for sensitive attributes, which can make the budget fully and reasonably utilized, minimize the budget waste, and improve the data availability.
[0223] Set ε to {0.01, 0.1, 1, 5}. As the number of clustering clusters changes, the changes in the record association values of each algorithm on the Adult dataset are as Figures 2(a) - 2(d) shown.
[0224] As can be seen from Figure 2, as ε increases, the noise added to the differentially private protected dataset decreases, and the record association value gradually rises. When the record association value is 0.01, the record association value is the lowest. At this time, although the privacy protection strength is high, the data availability is the worst. When ε takes 0.1 and 1 respectively, the difference in the record association values of the DP-IMKP algorithm is controlled within 0.05, and the differences in the record association values of the three algorithms of DP-MDAV, DP-K-prototype, and DCKPDP are all around 0.04. When ε takes 5, the privacy protection strength is the lowest at this time, and the superiority of using differential privacy cannot be reflected. Since the three algorithms of DP-MDAV, DP-k-prototype, and DCKPDP adopt the equal division principle to allocate privacy budgets for each attribute, therefore, regardless of the value of ε, the RL difference between these three algorithms remains at around 0.025. It can be considered that the protection strengths of these three algorithms are equivalent. However, the DP-IMKP algorithm adopts personalized allocation of privacy budgets, enabling it to be reasonably utilized, so that attributes with higher sensitivity levels are fully protected. Therefore, compared with the three algorithms of DP-MDAV, DP-k-prototype, and DCKPDP, the record association value of the DP-IMKP algorithm of the present invention is reduced by 3%-5%, and the data can be more fully protected.
[0225] According to the experimental analysis on the Adult dataset, when the value of ε is 1, the performance of the DP-IMKP algorithm is optimal. In order to reflect the scalability of the DP-IMKP method, ε is set to 1, and the same experiment is carried out on the dataset CI.
[0226] As can be seen from Figure 3(a) and Figure 3(b), when ε takes 1 and the experiment is carried out on the dataset CI, the DP-IMKP algorithm is also optimal in terms of both the sum of squared errors and the record association value. That is, the DP-IMKP algorithm of the present invention has advantages in improving data availability and reducing the risk of privacy leakage compared with similar algorithms.
[0227] By studying the privacy protection problem in data publishing, the present invention proposes a differentially private data publishing method that satisfies personalized privacy budget allocation. For the studied problem, a sensitive attribute grading strategy is proposed, and the most suitable privacy budget is matched for each level of sensitive attributes. The improved k-prototype clustering method is used to cluster the original dataset to obtain better clustering results. Finally, combined with the privacy budgets allocated to each sensitive attribute, differential privacy protection is carried out on the dataset. Through experimental verification, the DP-IMKP algorithm better solves the problems of low data availability and insufficient privacy protection strength existing in the current methods. However, when conducting experiments in the present invention, only some attributes in the dataset are selected. When dealing with a dataset containing high-dimensional attributes, there may be problems such as low publishing efficiency.
[0228] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated examples described herein.
Claims
1. A differential privacy data publishing method that meets personalized privacy budget allocation, characterized in that, Including: Step 1: Classify sensitive attributes according to the different sensitivity levels of attributes in the dataset through the mutual information and the correlation relationship between attributes; Step 2: Combine the optimal matching theory to construct a bipartite graph for privacy budget division, and match corresponding privacy budgets to sensitive attributes at different levels; Set the privacy loss of each original sensitive attribute to 0, and calculate the privacy consumption corresponding to each privacy budget and the information loss value of each sensitive attribute : ; ; Among them, the privacy protection intensity ratio among the high-sensitivity attribute group, medium-sensitivity attribute group, and low-sensitivity attribute group is 5:3:2; Construct a privacy budget division graph through a loss function; When there is a privacy budget that optimally matches the sensitive attributes in the privacy budget division graph, output the relationship graph between the sensitive attributes and the privacy budget; When there is no privacy budget that optimally matches the sensitive attributes in the privacy budget division graph, add the privacy loss of a sensitive attribute and continue to construct the graph for matching until an optimal matching result is obtained; Step 3: Use the improved k-prototype algorithm to cluster the original dataset, and perform differential privacy protection on the cluster center values through the privacy budgets allocated to the sensitive attributes to generate a to-be-published dataset that satisfies differential privacy; In the dataset the number of clusters is and the privacy budget is Select the initial centers and calculate the local density of each tuple record and respectively , ; Among them, is the tuple record and the distance between, is the truncation distance, that is, the value of the first 1% to 2% after sorting all tuple record distances from small to large; Calculate the distance between the record with a local density higher than the tuple record and the record closest to the tuple record : : ; Of the cluster center and need to satisfy the formula: ; Among them, and represent the mean of all tuple records and ; Calculated by formula A synthesis of each tuple record in And Parameter of : ; The larger it is, the closer the attribute value recorded by the tuple is to becoming a clustering center; for all tuple records Sort them from largest to smallest, The tuple with the largest value is recorded as the clustering center; For other preselected center points, calculate the distances of the tuple records before it in sequence. When the distance is greater than , then this tuple record can be used as an initial clustering center. Select the first tuple records as the initial clustering centers. For numerical attributes, add Laplace noise to the mean value of each attribute in the cluster; ; For categorical attributes, using the exponential mechanism, combined with the given privacy budget and scoring function, select the optimal clustering center that satisfies differential privacy protection to generate a differentially private dataset .
2. The differential privacy data publishing method for meeting personalized privacy budget allocation according to claim 1, characterized in that In the said Step 1, according to the degree of correlation of mutual information between attributes in the dataset, they are divided into: high-sensitivity attribute group, medium-sensitivity attribute group, and low-sensitivity attribute group.
3. The differential privacy data publishing method for meeting personalized privacy budget allocation according to claim 2, characterized in that, The said Step 3 also includes: When the clustering loss function changes, , calculate ; The loss function is as follows: ; Assign it to the cluster with the minimum dissimilarity, recalculate the center point within the cluster, calculate the central value of each attribute, calculate the value until the F value no longer changes, and obtain the clustering result. F
Citation Information
Patent Citations
Differential privacy protection-oriented k-means clustering method adopting
CN108280491A
A data fusion publishing algorithm based on differential privacy
CN109726758A