Differential privacy collection method for two-stage iterative segmentation key value data
Through a two-stage iterative segmentation method for key-value data, the problems of false key-value pair generation and privacy budget segmentation in key-value data collection are solved, higher estimation accuracy and lower computational cost are achieved, and the security and accuracy of the data are ensured.
Patent Information
- Application Number
- CN202510735321.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-12
AI Technical Summary
When collecting key-value data under local differential privacy, existing technologies face the problems of generating false key-value pairs or frequently splitting the privacy budget, resulting in reduced aggregation or estimation accuracy.
A two-stage iterative method for segmenting key-value data is adopted. First, the key-value data is preliminarily divided and disturbed on the user side, then the first frequency and mean estimation is performed on the server side, the value range division interval is updated, and the interval division is optimized through an iterative process to improve the estimation accuracy.
While ensuring the key-value correlation, the accuracy of frequency and mean estimation is improved, the computing, storage and communication costs are reduced, and the accuracy and availability of data are improved.
Smart Images

Figure CN120632933A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data privacy protection, and more particularly, to a differential privacy collection method for two-stage iterative segmentation of key-value data. Background Art
[0002] With the increasing popularity of big data analytics, application service providers are keen to collect user data to improve their services. However, this data collection carries the risk of privacy leakage, impacting both users and application service providers. To address this issue of privacy-preserving data collection, local differential privacy technology has been proposed.
[0003] Early research on local differential privacy focused on basic statistics of simple data, and subsequent research has gradually expanded to more complex data structures and query patterns. However, there is currently little research on mixed or heterogeneous data types. Key-value data, which is widely used in big data analysis, is such an example. The key-value data model is a data model widely used in non-relational databases, and key-value pairs are its basic units of data storage. The key-value storage model maps keys to corresponding values, where the key is the index of the key-value pair storage address and the value is the data stored in the key-value pair. Current mechanisms for collecting key-value data under local differential privacy mainly face problems such as the generation of false key-value pairs or frequent partitioning of the privacy budget, resulting in reduced aggregation or estimation accuracy. Summary of the Invention
[0004] The purpose of this paper is to design and develop a two-stage iterative segmentation differential privacy collection method for key-value data, which improves the accuracy while ensuring the correlation between keys and values and without splitting the privacy budget.
[0005] The technical solution provided by the present invention is:
[0006] A two-stage iterative segmentation differential privacy collection method for key-value data includes the following steps:
[0007] Step 1: The user terminal preliminarily divides the value range of its key-value data, and then perturbs the key-value data to obtain perturbation data;
[0008] Step 2: The user terminal sends the disturbance data to the server terminal, and the server terminal performs an initial frequency estimation and a mean estimation on the disturbance data from the user terminal;
[0009] Wherein, the mean estimate satisfies:
[0010]
[0011] Where, is the mean estimate of key k, is the frequency estimate of key k, is the number of keys k that are 1 in the perturbation data, is the number of key k equal to -1 in the perturbation data, n is the number of users, a is the first perturbation probability of controlling invalid key-value pairs, and p is the perturbation probability of controlling valid key-value pairs;
[0012] Step 3: Update the value range of the key value data;
[0013] Step 4: The server sends the updated value range partition interval of the key-value data to the user terminal, and the user terminal iterates again according to the updated value range partition interval of the key-value data to obtain a secondary frequency estimate and a mean estimate;
[0014] Step 5: The server calculates the rate of change of two adjacent mean estimates. If Then stop the iteration and output the quadratic frequency estimate and mean estimate;
[0015] in, is the rate of change between two adjacent mean estimates, and θ is the error threshold.
[0016] Preferably, the preliminary division is to divide the value range interval of the key-value data into 10 small intervals.
[0017] Preferably, the perturbation of the key-value data specifically includes:
[0018] Step 1: For the i-th user u i The key-value pair set S i Discretize the value data in:
[0019]
[0020] In the formula, v is the value data, v * is the discretized value data;
[0021] Step 2: The key-value pair consisting of the key data and the discretized value data is represented by a bit vector x. The length of the bit vector x is the same as the size of the key range, and the key position in the bit vector x that is the same as the corresponding position in the key range is replaced by v. * , all other positions are set to 0;
[0022] Step 3: Perturb the bit vector to obtain the perturbed vector:
[0023] When x[j]=0, the perturbation is:
[0024]
[0025] When x[j]≠0, it is perturbed as:
[0026]
[0027] Where x[j] is the data of the jth bit in the bit vector, x * [j] is the data after the jth bit in the bit vector is disturbed, and b is the second disturbance probability of controlling invalid key-value pairs.
[0028] Preferably, the disturbance probability of controlling the effective key-value pairs satisfies:
[0029]
[0030] Where ε is the privacy budget and g is the optimal hash domain.
[0031] Preferably, the optimal hash domain satisfies:
[0032]
[0033] Preferably, the frequency estimate of the key k satisfies:
[0034]
[0035] Preferably, the step three specifically includes:
[0036] If frep < δ, then merge the small interval with the adjacent small interval;
[0037] If frep>2δ, the small interval is divided into two equal small intervals;
[0038] Among them, frep is the value range of the key-value data, and δ is the optimization threshold.
[0039] Preferably, the rate of change of the two adjacent mean estimates satisfies:
[0040]
[0041] Where, is the quadratic mean estimate.
[0042] Preferably, the value of the interval optimization threshold satisfies: δ∈[0.005%, 0.02%].
[0043] Preferably, the error threshold is 0.01%.
[0044] The beneficial effects of the present invention are:
[0045] The present invention designs and develops a two-stage iterative segmentation differential privacy collection method for key-value data. Users perform operations such as partitioning and encoding perturbations on key-value data locally, which can ensure that the key and value remain associated. In an iterative manner, user data is counted in the partitioning stage, which improves the accuracy of frequency and mean estimation. The error threshold is used to avoid the computing, storage and communication costs caused by infinite loops and excessive calculations, thereby improving the usability and accuracy of privacy data and reducing errors while ensuring security. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a schematic diagram of the frequency mean square error comparison of the five algorithms described in the present invention based on the TalkingData dataset.
[0047] Figure 2 This is a schematic diagram comparing the mean and square deviation of the five algorithms described in the present invention based on the TalkingData dataset.
[0048] Figure 3 Schematic diagram of the comparison of RE values of the five algorithms described in the present invention based on the TalkingData dataset.
[0049] Figure 4 This is a schematic diagram of the comparison of the frequency mean square error of the five algorithms described in the present invention based on the JData dataset.
[0050] Figure 5 This is a schematic diagram comparing the mean and square error of the five algorithms described in the present invention based on the JData dataset.
[0051] Figure 6 This is a schematic diagram comparing the RE values of the five algorithms described in the present invention based on the JData dataset.
[0052] Figure 7 This is a schematic diagram of the comparison of the frequency mean square error of the five algorithms described in the present invention based on the GUASS dataset.
[0053] Figure 8 This is a schematic diagram comparing the mean and square error of the five algorithms described in the present invention based on the GUASS dataset.
[0054] Figure 9 Schematic diagram of the comparison of RE values of the five algorithms described in the present invention based on the GUASS dataset.
[0055] Figure 10 This is a schematic diagram of the comparison of the frequency mean square error of the five algorithms described in the present invention based on the PLAW dataset.
[0056] Figure 11 This is a schematic diagram comparing the mean and square error of the five algorithms described in the present invention based on the PLAW dataset.
[0057] Figure 12 Schematic diagram of the comparison of RE values of the five algorithms described in the present invention based on the PLAW dataset.
[0058] Figure 13 This is a schematic diagram of the comparison of the mean and square error of the five algorithms described in the present invention based on the TalkingData dataset on the query length of [0,0.4].
[0059] Figure 14 This is a schematic diagram of the comparison of the mean and square error of the five algorithms described in the present invention based on the TalkingData dataset on the query length of [0, 0.8].
[0060] Figure 15 This is a schematic diagram comparing the mean and square error of the five algorithms described in the present invention based on the JData dataset on the query length of [0, 0.4].
[0061] Figure 16 This is a schematic diagram of the comparison of the mean and square error of the five algorithms described in the present invention based on the JData dataset on the query length of [0, 0.8].
[0062] Figure 17 This is a schematic diagram of the comparison of the mean and square error of the five algorithms described in the present invention based on the GUASS dataset on the query length of [0, 0.4].
[0063] Figure 18 This is a schematic diagram of the comparison of the mean and square error of the five algorithms described in the present invention based on the GUASS dataset on the query length of [0, 0.8].
[0064] Figure 19 This is a schematic diagram of the comparison of the mean and square error of the five algorithms described in the present invention based on the PLAW dataset on the query length of [0, 0.4].
[0065] Figure 20 This is a schematic diagram of the comparison of the mean and square error of the five algorithms described in the present invention based on the PLAW dataset on the query length of [0, 0.8].
[0066] Figure 21 This is a schematic diagram comparing the mean square error of the three algorithms described in the present invention based on the GUASS dataset in terms of the number of iterations when the privacy budget is 0.2.
[0067] Figure 22 This is a schematic diagram comparing the mean square error of the three algorithms described in the present invention based on the GUASS dataset in terms of the number of iterations when the privacy budget is 0.8.
[0068] Figure 23 This is a schematic diagram comparing the mean square error of the three algorithms described in the present invention based on the GUASS dataset in terms of the number of iterations when the privacy budget is 3.2.
[0069] Figure 24 This figure shows a comparison of the mean square error of the three algorithms described in the present invention based on the PLAW dataset in terms of the number of iterations when the privacy budget is 0.2.
[0070] Figure 25 This is a schematic diagram comparing the mean square error of the three algorithms described in the present invention based on the PLAW dataset in terms of the number of iterations when the privacy budget is 0.8.
[0071] Figure 26 This is a schematic diagram comparing the mean square error of the three algorithms described in the present invention based on the PLAW dataset in terms of the number of iterations when the privacy budget is 3.2. DETAILED DESCRIPTION
[0072] The present invention is described in further detail below so that those skilled in the art can implement the invention with reference to the description.
[0073] The present invention provides a two-stage iterative segmentation key-value data differential privacy collection method, the system model used includes a data server and a set of users U of size |U|=n, each user has one or more key-value pairs<k,v> , where k∈K, v∈V, the domain size of key k is d, that is, K={1,2,…,d}, the domain of value v is V=[-1,1], and the set of key-value pairs owned by the user is denoted as S.
[0074] Step 1: After the user terminal preliminarily divides the value range of its key-value data, it converts the key-value data into a code and performs optimal local hashing (OLH) perturbation to obtain perturbed data;
[0075] The preliminary division is that the user end divides the value range L of the key-value data into 10 equal small intervals;
[0076] The step of converting the key-value data into a code and then performing the optimal local hashing (OLH) perturbation specifically includes the following steps:
[0077] Step 1: For the i-th user u i The key-value pair set S i The value data v in is discretized, so that the value of the key-value pair is discretized from floating point to binary:
[0078]
[0079] In the formula, v is the value data, v * is the discretized value data;
[0080] At this time, the domain of all key-value pairs becomes {<1,1>, <1,-1>}. To reduce communication costs, it is simplified to three single values {1,-1} corresponding to each;
[0081] Step 2: The key-value pair consisting of the key data and the discretized value data is represented by a bit vector x for encoding. The length of the bit vector x is the same as the value range of the key, and the key position in the bit vector x that is the same as the corresponding position in the value range of the key is replaced by v. * , all other positions are set to 0;
[0082] Step 3: Perturb the bit vector to obtain the perturbed vector:
[0083] When x[j]=0, the perturbation is:
[0084]
[0085] When x[j]≠0, it is perturbed as:
[0086]
[0087] Where x[j] is the data of the jth bit in the bit vector, x * [j] is the data after the jth bit in the bit vector is disturbed, a is the first disturbance probability of controlling invalid key-value pairs, b is the second disturbance probability of controlling invalid key-value pairs, and p is the disturbance probability of controlling valid key-value pairs.
[0088] The perturbation probability of controlling the effective key-value pairs satisfies:
[0089]
[0090] Where ε is the privacy budget and g is the optimal hash domain.
[0091] In this embodiment, when The best effect is achieved during the disturbance process, that is, and
[0092] Step 2: The user terminal sends the disturbance data to the server terminal, and the server terminal performs an initial frequency and mean estimation on the disturbance data from the user terminal;
[0093] Wherein, the frequency estimation satisfies:
[0094]
[0095] Where, is the frequency estimate of key k, is the number of keys k that are 1 in the perturbation data, is the number of keys k with -1 in the perturbation data, and n is the number of users;
[0096] The mean estimate satisfies:
[0097]
[0098] Where, is the mean estimate of key k;
[0099] Step 3: The server updates the value range of the key-value data based on the initial frequency and mean estimation, specifically including:
[0100] The server analyzes the data distribution frequency of each small interval L' initially divided:
[0101] The intervals where the data distribution frequency frep is less than the interval optimization threshold δ are identified as low-frequency or zero-frequency intervals, and the server merges them with adjacent intervals;
[0102] An interval with a data distribution frequency frep greater than twice the interval optimization threshold δ is considered a high-frequency interval. The server divides it into two equal subintervals. The left end of the first subinterval is the left end of the high-frequency interval, and the right end is the median of the high-frequency interval. The left end of the second subinterval is the median of the high-frequency interval, and the right end is the right end of the high-frequency interval.
[0103] Return the optimized interval L';
[0104] Among them, the distribution frequency frep will remain unchanged in the interval between δ and 2δ.
[0105] Step 4: The server sends the updated value range partition interval of the key-value data to the user terminal, and the user terminal iterates again according to the updated value range partition interval of the key-value data to obtain a secondary frequency and mean estimation;
[0106] Step 5: The server calculates the rate of change of two adjacent mean estimates like Then stop the iteration and output the quadratic frequency and mean estimation;
[0107] Among them, the rate of change between two adjacent mean estimates is satisfy:
[0108]
[0109] Where, is the quadratic mean estimate.
[0110] In this embodiment, the value of the interval optimization threshold satisfies: δ∈[0.005%, 0.02%].
[0111] In this embodiment, the error threshold is 0.01%.
[0112] The specific process of the differential privacy collection method for two-stage iterative segmentation of key-value data described in the present invention is as follows:
[0113] Algorithm 1 IAPKV algorithm
[0114] Input: User U = {u1,u2,…,u n} has a key value set S = {S1, S2, ..., S n}, privacy budget ε, perturbation probabilities p and q, range interval list L, error threshold θ, number of iterations c, interval optimization threshold δ;
[0115] Output: Frequency estimate Mean estimation
[0116] 1)
[0117] 2) Phase 1
[0118] 3) User side:
[0119] 4) Initialize the partition range list L = []
[0120] 5) For user u i , i=1~n
[0121] 6)x * [j]←IAPKV-OLH(S i ,ε)
[0122] 7) The user terminal will x * ={x * [j], j=1~d} is sent to the server
[0123] 8) Server side:
[0124] 9)
[0125] 10) L' = IAPKV-Divide(L, δ) / / Update the partition interval
[0126] 11) Phase II
[0127] 12) Notify the user to update the partition interval
[0128] 13) The user performs the first stage operation again
[0129] 14) R′←IAPKV-Analyze(R,θ)
[0130] 15) Stop iteration
[0131] 16) Output the iterated
[0132] As shown in Algorithm 1, the value range of the key-value data is first preliminarily divided. Then, the user uses the IAPKV-OLH algorithm to discretize the values in the key-value data and convert the key-value data into one-hot encoding. Next, the perturbed data is sent to a third-party server. The server then aggregates the data from the user and makes an estimate. It updates the preliminarily divided intervals and merges the new intervals. The new divided intervals are notified to the user end. The user end then executes the first stage steps to send the data to the server, determines whether the estimated value of two consecutive iterations is less than the error threshold, and stops the iteration. Finally, the server completes the estimation based on the collected data.
[0133] Algorithm 2 IAPKV-OLH algorithm
[0134] Input: User's key-value pair set S i , user set U, privacy budget ε, perturbation probability p, a and b
[0135] Output: perturbed vector x *
[0136] 1) For each user u i Key-value pairs<k,v> Do
[0137] 2) The key-value pair<k,v> The value v is discretized into v *
[0138]
[0139] 3) Create a bit vector x of length d and encode x
[0140] 4) Perturb each bit in the bit vector x
[0141] 5) For j=1~d, do
[0142] 6) When x[j]=0
[0143] 7)
[0144] 8) When x[j]≠0
[0145] 9)
[0146] 10) Output x *
[0147] As shown in Algorithm 2, after the value range is divided, the key-value pairs are perturbed by OLH. The IAPKV-OLH algorithm involves three steps: discretization, encoding, and perturbation. First, the value of the key-value pair is discretized, and then a bit vector x of length d is created for encoding. In x, only the positions corresponding to the key are set to the value, and the remaining positions are set to 0. After encoding, all key-value pair data in the key-value pair set become {<1,-1>,<1,1>,<0,0>}. To reduce communication costs, it is simplified to three states {-1,1,0}. The perturbation probability during the perturbation process includes p and a, and it is ensured that the probability after perturbation meets the definition of OLH, so that each bit in the vector is perturbed.
[0148] Algorithm 3 IAPKV-Aggregate Algorithm
[0149] Input: perturbation vector x * , counting statistics
[0150] Output: Frequency estimate Mean estimation
[0151] 1)
[0152] 2) Receive the perturbed vector x from the user *
[0153] 3) Collect two statistical results:
[0154]
[0155] 4) Calculate frequency and mean estimates
[0156]
[0157] 5) Output
[0158] Algorithm 3 is an aggregation mechanism. After the server receives the perturbed data from the user, it uses the IAPKV-Aggregate method to aggregate and estimate, collects the count values of 1 and -1 under the key k, and then calculates the frequency and mean estimation.
[0159] Algorithm 4 IAPKV-Divide algorithm
[0160] Input: range list L, interval optimization threshold δ, data distribution frequency frep
[0161] Output: Optimized interval list L
[0162] 1) The first stage: initial uniform division
[0163] 2) for each user u in U i do
[0164] 3) Divide the range into 10 equal-width intervals
[0165] 4) / / Statistical distribution frequency frep
[0166] 5)end for
[0167] 6) / / Second stage: adaptive optimization and partitioning
[0168] 7) for each small interval in L do
[0169] 8)if frep<δ / / process low frequency range
[0170] 9) Merge the current interval and the next interval
[0171] 10) eilffrep>2δ / / Processing high frequency range
[0172] 11) Divide the high frequency interval into two sub-intervals
[0173] 12) else / / remain unchanged
[0174] 13)end if
[0175] 14)end for
[0176] 15)return L
[0177] As shown in Algorithm 4, the IAPKV-Divide algorithm still operates in two phases. In the first phase, an initial uniform partition is performed, dividing the value data interval into 10 equal parts and calculating the data distribution frequency within the interval. In the second phase, an adaptive optimization partition is performed based on the data distribution. The server analyzes the data distribution frequency of each interval in the initial partition and then performs adaptive partition adjustments based on the analysis results: intervals with a distribution frequency less than the interval optimization threshold δ are identified as low-frequency or zero-frequency intervals and merged with adjacent intervals; intervals with a distribution frequency greater than twice the interval optimization threshold δ are identified as high-frequency intervals and split into two sub-intervals. Intervals with a distribution frequency between δ and 2δ remain unchanged. Adaptive optimization based on the data distribution during the partitioning process can achieve the goal of improving calculation accuracy.
[0178] Algorithm 5IAPKV-Analyze Algorithm
[0179] Input: mean estimation result Error threshold θ, historical analysis results R
[0180] Output: Analysis result R′
[0181] 1)
[0182] 2)
[0183] 3) When
[0184] 4) Determine when to stop the iteration
[0185] 5) Output R'
[0186] As shown in Algorithm 5, the operating rules of IAPKV-Analyze are as follows: the server records the mean estimation results in each round of iteration records and calculates the change rate of the mean estimation results between two adjacent rounds to evaluate the convergence of the overall algorithm. When the change rate of the results between two adjacent rounds is less than the preset error threshold, the iteration is considered complete.
[0187] The method described in this invention is now analyzed from the perspectives of privacy and usability. The privacy of the method described in this invention is explained from the perspective of ε-local differential privacy, and its usability is described from the perspectives of unbiasedness and variance:
[0188] 1. IAPKV-OLH satisfies ε-local differential privacy
[0189] Assumptions<k1,v1> ,<k2,v2> For two different key-value pairs, the perturbed key-value pair is <k * ,v * > indicates that the perturbation vector is x * According to the maximum probability that two key-value pairs have different output vectors, it is proved that IAPKV-OLH satisfies ε-local differential privacy.
[0190]
[0191] or
[0192]
[0193] Using IAPKV-OLH, the perturbed vectors are indistinguishable:
[0194]
[0195] The proof is complete.
[0196] 2. IAPKV satisfies unbiased estimation
[0197] In IAPKV, frequency estimation and mean estimates In the perturbation probability and To prove that the theorem holds, we can prove it from the following two aspects:
[0198] (1)
[0199] To prove Let n k is the true count for key k, and So
[0200]
[0201] therefore,
[0202] (2)
[0203] according to It is known that as long as If established, Establish, prove You can get
[0204] According to the multivariate Taylor expansion of the random variable function, let but The expectation of the quotient of two random variables X and Y can be approximated as:
[0205]
[0206] The variance of the random variable Y is:
[0207] The covariance is:
[0208]
[0209] Var[Y] and Cov X,Y Then calculate them by remembering their exact values respectively.
[0210]
[0211] The proof is complete.
[0212] In order to verify the performance of the IAPKV algorithm described in the present invention, the following experiments were performed:
[0213] The experimental platform is a 6-core Intel i7-9750H CPU (2.6GHz), 16GB memory, and Windows 11. All algorithms are implemented in Python.
[0214] Experiments were conducted on the following datasets: TalkingData, JData, GUASS, and PLAW. The TalkingData dataset is collected from the TalkingData SDK and contains usage data of 306 applications on 60,822 devices. The JData dataset, from the JD.com e-commerce platform, records more than 50.6 million sales records generated by 442 brands in 2016, involving 105,180 users. GUASS and PLAW are synthetic datasets in which the number of keys, values, and key-value pairs follows Gaussian distribution and power-law distribution. The details of the datasets are shown in Table 1.
[0215] Table 1 Description of information of four datasets
[0216]
[0217] In order to ensure the accuracy of frequency and mean estimation, mean square error (MSE) and relative error (RE) are used for evaluation:
[0218] (1) Mean square error:
[0219] The mean squared error is used to measure the accuracy of the estimated frequency (or mean) relative to the actual frequency (or mean) of all keys. For k∈K, let f k and Represent the true frequency and estimated frequency respectively, m k and denote the true mean and the estimated mean, respectively, where d = |K|, which is defined as follows:
[0220]
[0221] Since the variance of MSE is often large and there is a square term, large errors will be severely amplified. Therefore, Log(MSE), that is, logarithmic mean square error, is used. It can compress the numerical range and significantly reduce the overall variance:
[0222]
[0223] Assume the error term z i =log(e i 2 +ω), the variances of MSE and Log-MSE are derived as follows:
[0224]
[0225] When e i Obey the normal distribution N(0,σ 2):The variance of MSE ∝σ 4 , the variance of Log-MSE ∝log 2 (σ 2 ), when σ is large, the variance of Log-MSE is significantly smaller than that of MSE. To ensure numerical stability, use:
[0226]
[0227] Here, ω is a small positive number to prevent the logarithm from being zero.
[0228] (2) Relative error:
[0229] Use the common error measurement method - relative error to measure the difference between the estimated frequency (or mean) and the actual frequency (or mean), let f k and Represent the true frequency and estimated frequency respectively, m k and They represent the true mean and the estimated mean, respectively, and are defined as follows:
[0230]
[0231]
[0232] 1. Comparison of MSE and RE based on real datasets TalkingData and JData
[0233] Figures 1 to 6 Describes IAPKV, KSKV, PCKV, PrivKVM, and PrivKVM respectively * Comparison results of MSE and RE values of five algorithms on TalkingData and JData datasets. Figure 1 、 2 and Figure 4 、 5 It can be found that the Log(MSE) values of the five algorithms decrease with the increase of ε, and the IAPKV algorithm performs better than the other four algorithms and guarantees a lower Log(MSE) value under the same privacy budget. When ε=3.2, the accuracy of the IAPKV method far exceeds that of the other algorithms. Figure 3 and Figure 6 We found that the RE of the five algorithms decreases as ε varies from 0.1 to 3.2. Meanwhile, when ε is fixed, the IAPKV method achieves a lower RE than the other methods, with lower error and higher precision. Experimental results demonstrate that the IAPKV algorithm can provide more accurate frequency and mean estimates while preserving privacy.
[0234] 2. Comparison of MSE and RE of algorithms based on synthetic datasets GUASS and PLAW
[0235] Figure 7 and Figure 12 Describes IAPKV, KSKV, PCKV, PrivKVM, and PrivKVM respectively * Comparison results of MSE and RE values of five algorithms on GUASS and PLAW datasets. Figure 7 、 8 and Figure 10 、 11 It can be found that the Log(MSE) values of the five algorithms decrease with the increase of ε, and the IAPKV algorithm performs better than the other four algorithms and guarantees a lower Log(MSE) value under the same privacy budget. Moreover, when ε=3.2, the accuracy of the IAPKV method is nearly twice that of the PCKV algorithm. Compared with other methods, the accuracy of this method is also better than other methods. Figure 9 and Figure 12 We found that the RE of the five algorithms decreases continuously as ε varies from 0.1 to 3.2. Meanwhile, when ε is fixed, the IAPKV method achieves a lower RE than the other methods, with lower error and higher precision. Experimental results demonstrate that the proposed method achieves higher accuracy under both different and fixed privacy budgets.
[0236] 3. Range query based on real datasets TalkingData, JData and synthetic datasets GUASS, PLAW
[0237] Because the data is distributed in [-1, 1], the mean square error of the mean estimate of the IAPKV algorithm is measured at two query lengths of [0, 0.4] and [0, 0.8] and is still measured using Log(MSE), such as Figures 13 to 20 The query results of five methods are presented. Within the privacy budget range of 0.1-3.2, the IAPKV method consistently outperforms the other four methods in accuracy, and the larger the ε, the higher the accuracy of this method. Experimental results demonstrate that the proposed method has a small deviation from the true value, has good accuracy, can effectively protect the collection of key-value data, and is highly practical.
[0238] 4. Iteration number experiment based on synthetic datasets GUASS and PLAW
[0239] We select three different privacy budgets of ε = 0.2, ε = 0.8, and ε = 3.2 for comparison. We also compare the performance of PrivKVM, PrivKVM*, and IAPLV under the same privacy budget ε. We compare the mean square error of the mean estimation. Figures 21 to 26Comparative results of the three algorithms on the GUASS and PLAW datasets are presented. The experimental data show that the Log(MSE) values of the three algorithms decrease with increasing iterations and gradually converge to a certain value. The IAPKV algorithm is more accurate than the other two algorithms, and when the number of iterations is the same, the IAPKV algorithm has a lower Log(MSE) value. These experimental results demonstrate that the method proposed in this paper is more accurate, can provide more accurate estimates with a limited number of iterations, and has high practicality and effectiveness.
[0240] This paper designs and develops a two-stage iterative segmentation method for differentially private key-value data collection. While ensuring the correlation between keys and values and maintaining the privacy budget, this method performs segmentation, discretization, and perturbation operations locally on the user, protecting their key-value data. Intervals are adaptively split or merged based on the frequency of data distribution within the interval and the interval optimization threshold, improving the precision of statistical results and enhancing the accuracy of private data without sacrificing security.
[0241] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and embodiments shown and described herein.
Claims
1. A two-stage iterative segmentation differential privacy collection method for key-value data, characterized by: include: Step 1: The user terminal preliminarily divides the value range of its key-value data, and then perturbs the key-value data to obtain perturbation data; Step 2: The user terminal sends the disturbance data to the server terminal, and the server terminal performs an initial frequency estimation and a mean estimation on the disturbance data from the user terminal; Wherein, the mean estimate satisfies: Where, is the mean estimate of key k, is the frequency estimate of key k, is the number of keys k that are 1 in the perturbation data, is the number of key k equal to -1 in the perturbation data, n is the number of users, a is the first perturbation probability of controlling invalid key-value pairs, and p is the perturbation probability of controlling valid key-value pairs; Step 3: Update the value range of the key value data; Step 4: The server sends the updated value range partition interval of the key-value data to the user terminal, and the user terminal iterates again according to the updated value range partition interval of the key-value data to obtain a secondary frequency estimate and a mean estimate; Step 5: The server calculates the rate of change of two adjacent mean estimates. If Then stop the iteration and output the quadratic frequency estimate and mean estimate; in, is the rate of change between two adjacent mean estimates, and θ is the error threshold.
2. The differential privacy collection method for key-value data with two-stage iterative segmentation as claimed in claim 1, characterized in that: The preliminary division is to divide the value range of the key-value data into 10 small intervals.
3. The differential privacy collection method for two-stage iterative segmentation of key-value data as described in claim 2 is characterized in that: The perturbation of the key-value data specifically includes: Step 1: For the i-th user u i The key-value pair set S i Discretize the value data in: In the formula, v is the value data, v * is the discretized value data; Step 2: The key-value pair consisting of the key data and the discretized value data is represented by a bit vector x. The length of the bit vector x is the same as the size of the key range, and the key position in the bit vector x that is the same as the corresponding position in the key range is replaced by v. * , all other positions are set to 0; Step 3: Perturb the bit vector to obtain the perturbed vector: When x[j]=0, the perturbation is: When x[j]≠0, it is perturbed as: Where x[j] is the data of the jth bit in the bit vector, x * [j] is the data after the jth bit in the bit vector is disturbed, and b is the second disturbance probability of controlling invalid key-value pairs.
4. The differential privacy collection method for two-stage iterative segmentation of key-value data as described in claim 3 is characterized in that: The disturbance probability of controlling the effective key-value pairs satisfies: Where ε is the privacy budget and g is the optimal hash domain.
5. The differential privacy collection method for two-stage iterative segmentation of key-value data as claimed in claim 4, characterized in that: The optimal hash domain satisfies:
6. The differential privacy collection method for two-stage iterative segmentation of key-value data according to claim 5, characterized in that: The frequency estimate of the key k satisfies:
7. The differential privacy collection method for two-stage iterative segmentation of key-value data according to claim 6, characterized in that: The step three specifically includes: If frep < δ, then merge the small interval with the adjacent small interval; If frep>2δ, the small interval is divided into two equal small intervals; Among them, frep is the value range of the key-value data, and δ is the optimization threshold.
8. The differential privacy collection method for two-stage iterative segmentation of key-value data according to claim 7, characterized in that: The rate of change of the two adjacent mean estimates satisfies: Where, is the quadratic mean estimate.
9. The differential privacy collection method for two-stage iterative segmentation of key-value data as claimed in claim 8, characterized in that: The value of the interval optimization threshold satisfies: δ∈[0.005%,0.02%].
10. The differential privacy collection method for two-stage iterative segmentation key-value data according to claim 9, characterized in that: The error threshold is 0.01%.