Privacy protection method for user position data query

By adaptively dividing sparse and dense areas, and using quad-tree index and frequency estimation methods in dense areas, combined with LRR perturbation algorithm and OUE method, the problems of privacy protection and low efficiency in user location data query are solved, and high-precision and high-efficiency data query are achieved.

CN120046188APending Publication Date: 2025-05-27LIAONING UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510115347.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively resist new attacks when processing user location data, and the spatial division method based on differential privacy has problems such as uneven data distribution and low query efficiency.

Method used

A privacy protection method for user location data query is designed. By grouping random sampling technology, sparse areas and dense areas are adaptively divided, and quad-tree index and frequency estimation are used in dense areas for further decomposition, combining LRR perturbation algorithm and OUE method to achieve data privacy protection.

Benefits of technology

It improves the accuracy and operation efficiency of user location data query, meets the protection requirements of ε-localized differential privacy, reduces noise error and non-uniform error, and optimizes the data acquisition process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046188A_ABST
    Figure CN120046188A_ABST
Patent Text Reader

Abstract

The invention discloses a privacy protection method for user position data query, which comprises the following steps of: 1, acquiring user position data through a sensor, grouping based on a random sampling technology, and enabling the grouped user data to meet the parallel combinatorial property of epsilon-localization differential privacy; 2, decomposing the spatial data based on a hierarchical decomposition method according to the spatial data distribution, and dividing the space into a sparse region and a dense region according to a set density threshold value in the first-layer decomposition; wherein the sparse region is not divided continuously; the dense region is further divided through a quadtree index; and step 3, completing post-disturbance data aggregation and spatial range query based on an RQT algorithm. Privacy protection can be performed on user position data query, and query precision and operation efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a privacy protection method for querying user location data and belongs to the technical field of the Internet of Things. Background Art

[0002] In the era of big data, with the rapid development of information technology and the wide popularization of intelligent devices, the huge value contained in spatial data enables it to be widely applied in many fields such as the Internet of Things, geographic information systems, global positioning systems, crowdsourcing systems, etc., such as recommending surrounding places for users and providing precise navigation, greatly facilitating people's lives.

[0003] However, spatial data contains a large amount of sensitive personal information and has a high risk of privacy leakage. Traditional privacy protection mechanisms such as k-anonymity, l-diversity, t-closeness, etc. need to make assumptions about specific attacks and background knowledge when dealing with spatial data partitioning and are difficult to resist new types of attacks.

[0004] Although differential privacy can resist attacks under any background knowledge, existing spatial partitioning methods based on it, such as UG, AG, etc., have problems such as not considering uneven data distribution and low query efficiency. Methods based on local differential privacy also have limitations, such as large query errors and large computational overheads. Therefore, a method is needed to solve the privacy protection problem of querying user location data. Summary of the Invention

[0005] The present invention designs and develops a privacy protection method for querying user location data, which can protect the privacy of querying user location data and improve query accuracy and operation efficiency.

[0006] The technical solution provided by the present invention is as follows:

[0007] A privacy protection method for querying user location data, comprising:

[0008] Step 1: Obtain user location data through sensors and group them based on random sampling technology so that the grouped user data satisfies the parallel combination property of ε-local differential privacy;

[0009] Step 2: According to the spatial data distribution, based on the hierarchical decomposition method, in the first-level decomposition, divide the space into sparse regions and dense regions according to a set density threshold;

[0010] Among them, for the sparse regions, no further partitioning will be carried out. Each group of users in the sparse regions uses the local random response method to perturb their location data and send the perturbed results to the data collector to complete the privacy protection of the location of users in the sparse regions;

[0011] For dense regions, construct a quadtree, perform post - processing on it, further divide the dense regions through quadtree indexing. Users at each level perform local random perturbations and send the perturbation results to the data collector for frequency construction. When the estimated frequency of a leaf node is less than the set frequency threshold or reaches the set tree height, stop the iteration to complete the privacy protection of the user locations in the dense regions;

[0012] Step 3: Based on the RQT algorithm, complete the aggregation of perturbed data and the spatial range query.

[0013] Preferably, the second step further includes:

[0014] Set the data set as D, the privacy budget as ε, the number of grid cells as m, the tree height as h, the density threshold as 1, and the frequency threshold as θ;

[0015] Randomly divide the data set D into n subsets to form multiple data chunks;

[0016] Perform adaptive division of sparse and dense regions according to the density threshold;

[0017] Construct the root node V 0 , and insert the divided network unit child nodes into the tree, and set the frequency estimate of the root node V 0 as 1;

[0018] Each user group uses the local random response method to locally perturb its own location data and send these perturbed results to the data collector;

[0019] Compare the user density within the region with the density threshold,

[0020] When the user density is greater than the density threshold, it is a dense region;

[0021] When the user density is less than or equal to the density threshold, it is a sparse region;

[0022] Perform quadtree index decomposition on the dense region, recursively until the estimated frequency of the leaf node is less than or equal to the frequency threshold or reaches the set tree height;

[0023] Perform post - processing on the tree, query the result, and return the query frequency.

[0024] Preferably, it further includes:

[0025] In the dense region, by comparing the estimated value with the threshold, each interval is separately divided into four parts. During the process of spatial decomposition, the decomposition of each interval is accompanied by noise error and non - uniform error. Distribute the frequency of the estimated node evenly to each child node to obtain the desired estimation error, and satisfy the formula:

[0026]

[0027] Wherein, X is the added noise, f is the estimated frequency, and Var is the probability of the estimated frequency of the leaf node;

[0028] The non-uniform error is: k ∈ 1...4.

[0029] Preferably, the calculation formula of the frequency threshold is:

[0030]

[0031] Wherein, h is the tree height, N is the value range of the data, and e is the exponent.

[0032] Preferably, it further includes:

[0033] Based on the optimal unary coding protocol, the perturbed vector b' is generated by perturbing the coding vector b, and the perturbation rule is:

[0034]

[0035] Wherein, Pr[] is the probability of the event occurring, p and q are probabilities, and b i is different value cases.

[0036] Preferably, in the second step, the data collector aggregates all the perturbed vectors collected, estimates the frequency of each value, and based on the coding and perturbation waterproofing of OUE, the calculation formula of the estimated variance is:

[0037]

[0038] The beneficial effects of the present invention: The privacy protection method for querying user location data provided by the present invention first uses the adaptive grid division method to divide the sparse area and the dense area, solving the noise error caused by data collection. Then, on this basis, the adaptive hierarchical decomposition method is carried out, and the dense area is further decomposed by using the quadtree index and frequency estimation method. Finally, an efficient post-processing technology is carried out for the frequency of each node, improving the accuracy of data query. Through experiments and comparisons with three algorithms, namely GT-R, PrivAG, and ASDQT, the results show that when facing datasets with uneven data distribution or dense areas, this method has better query accuracy and running efficiency.

[0039] This method satisfies ε-local differential privacy to protect the privacy of data. The LRR perturbation algorithm and the OUE method work together to ensure that user data in data processing cannot be easily obtained by the outside world. Through the user grouping strategy, the privacy budget segmentation is reduced, the noise error is reduced, and the data acquisition process is optimized. When processing dense regions, a frequency threshold is set to limit the quadtree refinement, balancing the noise error and the non-uniform error, ensuring that the overall error is lower than the existing UG algorithm. It not only reasonably divides the space, avoids performance degradation, but also improves the accuracy of spatial range queries, maintaining the availability of data while protecting privacy. Description of the Drawings

[0040] Figure 1 It is a flowchart of the privacy protection method for querying user location data according to the present invention.

[0041] Figure 2a It is a comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the user check-in dataset with different privacy budgets when the query range is [5%, 50%].

[0042] Figure 2b It is a comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the user check-in dataset with different privacy budgets when the query range is [10%, 55%].

[0043] Figure 2c It is a comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the user check-in dataset with different privacy budgets when the query range is [15%, 60%].

[0044] Figure 2d It is a comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Beijing taxi dataset with different privacy budgets when the query range is [5%, 50%].

[0045] Figure 2e It is a comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Beijing taxi dataset with different privacy budgets when the query range is [10%, 55%].

[0046] Figure 2f It is a comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Beijing taxi dataset with different privacy budgets when the query range is [15%, 60%].

[0047] Figure 2gComparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Tokyo dataset with different privacy budgets and a query range of [5%, 50%].

[0048] Figure 2h Comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Tokyo dataset with different privacy budgets and a query range of [10%, 55%].

[0049] Figure 2i Comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Tokyo dataset with different privacy budgets and a query range of [15%, 60%].

[0050] Figure 3a Comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the user check-in dataset with a privacy budget ε of 0.3 and different query ranges.

[0051] Figure 3b Comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the user check-in dataset with a privacy budget ε of 0.6 and different query ranges.

[0052] Figure 3c Comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the user check-in dataset with a privacy budget ε of 0.9 and different query ranges.

[0053] Figure 3d Comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Beijing taxi dataset with a privacy budget ε of 0.3 and different query ranges.

[0054] Figure 3e Comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Beijing taxi dataset with a privacy budget ε of 0.6 and different query ranges.

[0055] Figure 3f Comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Beijing taxi dataset with a privacy budget ε of 0.9 and different query ranges.

[0056] Figure 3gComparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Tokyo dataset under different query ranges with a privacy budget ε of 0.3.

[0057] Figure 3h Comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Tokyo dataset under different query ranges with a privacy budget ε of 0.6.

[0058] Figure 3i Comparison of the mean squared errors of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Tokyo dataset under different query ranges with a privacy budget ε of 0.9.

[0059] Figure 4a Comparison of the running times of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the user check-in dataset under different privacy budgets with a query range of [5%, 50%].

[0060] Figure 4b Comparison of the running times of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the user check-in dataset under different privacy budgets with a query range of [10%, 55%].

[0061] Figure 4c Comparison of the running times of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the user check-in dataset under different privacy budgets with a query range of [15%, 60%].

[0062] Figure 4d Comparison of the running times of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Beijing taxi dataset under different privacy budgets with a query range of [5%, 50%].

[0063] Figure 4e Comparison of the running times of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Beijing taxi dataset under different privacy budgets with a query range of [10%, 55%].

[0064] Figure 4f Comparison of the running times of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Beijing taxi dataset under different privacy budgets with a query range of [15%, 60%].

[0065] Figure 4gComparison of the running times of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Tokyo dataset with different privacy budgets within the query range of [5%, 50%].

[0066] Figure 4h Comparison of the running times of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Tokyo dataset with different privacy budgets within the query range of [10%, 55%].

[0067] Figure 4i Comparison of the running times of four algorithms, namely GT-R, PrivAG, ASDQT, and LDP-AST, for the Tokyo dataset with different privacy budgets within the query range of [15%, 60%]. Detailed implementation manner

[0068] The following further elaborates on the present invention in conjunction with the accompanying drawings so that those skilled in the art can implement it with reference to the text of the specification.

[0069] As Figure 1 shown in -4, the present invention provides a privacy protection method for querying user location data, including:

[0070] Step 1: Obtain user location data through sensors, group it based on random sampling technology, so that the grouped user data satisfies the parallel combination property of ε-local differential privacy;

[0071] Step 2: According to the spatial data distribution, based on the hierarchical decomposition method, in the first-layer decomposition, divide the space into sparse regions and dense regions according to the set density threshold;

[0072] Among them, for the sparse regions, no further division will be carried out. Each group of users in the sparse regions uses the local random response method to perturb their location data and send the perturbed results to the data collector to complete the privacy protection of the user locations in the sparse regions;

[0073] For the dense regions, construct a quadtree and perform post-processing on it. Further divide it through the quadtree index. Each layer of users performs local random perturbation and sends the perturbed results to the data collector for frequency construction. When the estimated frequency of the leaf node is less than the set frequency threshold or reaches the set tree height, the iteration stops to complete the privacy protection of the user locations in the dense regions;

[0074] Step 3: Based on the RQT algorithm, complete the aggregation of the perturbed data and the spatial range query.

[0075] Regarding ε-local differential privacy settings, specifically including: assuming there is a random algorithm M, if and only if two different inputs x, x' (x, x' ∈ Dom), the resulting output y (y ∈ ran) satisfies: Then M satisfies ε-local differential privacy.

[0076] Where Dom(M) is all possible spatial data of users, Ran(M) is all possible outputs, Pr[] represents the probability of an event occurring, and ε is the privacy budget. The smaller the value of ε, the higher the privacy protection level of the algorithm.

[0077] Sequential composition: Given a dataset D and random algorithms {M 1 , M 1 , M 1 ,..., M j}, if the algorithm M satisfies ε-local differential privacy, then the sequential composition {M 1 , M 1 , M 1 ,..., M j} on the dataset D satisfies ε-local differential privacy

[0078] Parallel composition: Assuming several independent and disjoint subsets {D 1 , D 2 , D 3 ,..., D j} on the dataset D, and there are random algorithms {M 1 , M 1 , M 1 ,...,, M j} that satisfy the ε-local differential privacy of the sequential composition on the subset D i , then {M 1 , M 1 , M 1 ,...,, M j} provides ε-local differential privacy on the dataset D.

[0079] Frequency Estimation FO provides a tool for obtaining information about the data distribution and satisfies local differential privacy. Most FO protocols include three steps: encoding, perturbation, and aggregation. Optimal Unary Encoding OUE is a typical FO protocol. While protecting the privacy of user data, it can still effectively perform data analysis. OUE encodes the value v using unary encoding, and the encoded result is a binary vector b of length d, where only the v-th bit is 1 and the other bits are 0: Encode(v) = [0,…,0,1,0,…,0]. OUE generates the perturbed vector b' by perturbing the encoded vector b, and the perturbation rule is as follows:

[0080]

[0081] where ε is the privacy budget, which determines the perturbation strength.

[0082] In the data collection stage, the data collector aggregates all the perturbation vectors collected and estimates the frequency of each value i. Due to the encoding and perturbation method of OUE, its estimation variance is:

[0083]

[0084] The common methods for existing spatial decomposition are divided into two types: data-dependent decomposition and data-independent decomposition. KD-tree and R-tree are two common data-dependent decomposition methods, which rely on the distribution characteristics of the data itself to divide the space. In local differential privacy, since there is no trusted third-party data collector that can access user data, the data-dependent decomposition method cannot be directly used. Data-independent decomposition does not depend on the actual distribution of the data, but uses fixed rules to divide the space. The most common method is the quadtree, which is often used to solve two-dimensional space problems. It recursively divides a space into four subspaces, and the recursion stops when the area of the sub-region is small enough.

[0085] A quadtree is a tree-shaped data structure where each non-leaf node has four child nodes. It is suitable for efficiently dividing and managing two-dimensional space. When constructing a quadtree, taking the center point of the space as a reference, through perpendicular horizontal and vertical lines, the space is recursively and continuously divided into four equal-sized quadrants or regions. The recursion stops until a predetermined termination condition is reached, such as the size of each region or the number of objects it contains.

[0086] The quadtree divides the geographical space into four equal-sized quadrants through a recursive method and can be adaptively subdivided according to the distribution of the data. This recursive division method enables it to handle geographical space data such as GIS and remote sensing images well. In contrast, the KD-tree will add redundant divisions for regular data, and the R-tree is more suitable for managing irregular-shaped data. In addition, the quadtree structure is simple, with each node only divided into four child nodes. Therefore, its memory consumption is lower than the complex node management of the R-tree and is also more resource-saving than the dynamic division structure of the KD-tree. This makes the quadtree perform better in cases where resources are limited, especially in large-scale data processing tasks.

[0087] Regarding the problems of decomposition of dense space and data query, this method first uses random sampling technology to group users so as to satisfy the parallel composition property of ε-local differential privacy; then, according to the spatial data distribution, a hierarchical decomposition method is adopted. In the first-level decomposition, the space is divided into sparse regions and dense regions based on the calculated density threshold; for the sparse regions, no further division is carried out. At this time, each group of users in the sparse region uses the local random response method to locally perturb their location data and send the perturbed results to the data collector; finally, for the dense regions, a quadtree index is used for further division. Each layer of user groups performs local random perturbation and sends the perturbed results to the data collector for frequency construction. When the estimated frequency of the leaf node is less than the given frequency threshold or reaches the specified tree height, the iteration stops.

[0088] Abbreviate the privacy protection method for user location data query as: LDP-ASDT algorithm.

[0089] The LDP-ASDT algorithm, as Algorithm 1, includes:

[0090] Algorithm 1 LDP-ASDT algorithm

[0091] Input data set D, privacy budget ε, number of grid cells m, tree height h, density threshold l, frequency threshold

[0092] θ

[0093] Output query result

[0094]

[0095] The LDP-ASDT algorithm randomly divides the data set D into n subsets by grouping users and dividing spatial regions, forming multiple data chunks. Compared with the traditional UG method, when dividing regions, the LDP-ASDT does not simply divide the space into m×m grid cells, but adaptively divides the sparse regions and dense regions according to the density threshold. This method can more flexibly adapt to different data density distributions, effectively narrow the range of dense regions, and optimize the noise error generated in the data collection process (Steps 1 - Step 2). Construct the root node (V 0 ) and insert the divided grid cells as child nodes into the tree, and set the frequency estimate of V 0 to 1 (Step 3). Each user group uses the local random response (LRR) method to locally perturb its own location data and sends these perturbed results to the data collector. The LRR ensures the privacy of user data, and even during data transmission, it is difficult for the outside world to directly obtain the true location of users.

[0096] Compare the user density within a region with the density threshold. When the user density is greater than the density threshold, it is a dense region; otherwise, it is a sparse region. For a dense region, further determine whether the estimated frequency of the leaf node is less than the frequency threshold, and decompose the region using quadtree indexing recursively until the estimated frequency of the leaf node is less than or equal to the frequency threshold or reaches the specified tree height (Steps 4 - Step 13). Finally, perform post-processing operations on the tree, query the result, and return the query frequency (Steps 14 - Step 16).

[0097] Algorithm 2 LRR Algorithm

[0098] Input the location of the user (1 < i <= n), the index tree, and the privacy budget ε

[0099] Output the perturbed location information of the user

[0100]

[0101] Design the LRR algorithm with reference to the GT - R algorithm. The data collector sequentially distributes the index tree obtained in the previous step to each user group. The user determines the leaf node to which their spatial location belongs in the index tree, independently encodes the belonging leaf node, the length of the encoded result vector is equal to the total number of leaf nodes, and finally uses the optimized random response mechanism to perturb and generate a report.

[0102] Algorithm 3 Post processing

[0103] Input the index tree T

[0104] Output the index tree T

[0105]

[0106] As shown in Algorithm 3, since there is no constraint on the node frequency during the process of constructing the tree, the nodes are processed through post - processing. Use frequency normalization to standardize each node so that it has comparability in the subsequent processing. On this basis, traverse all the nodes of the tree from bottom to top to update the frequency, so that the weighted average frequency value of each node is the sum of the frequency value of the node itself and the weighted average frequency values of all its child nodes. Then perform mean consistency processing from top to bottom, and evenly distribute the total difference between the estimated frequency of the parent node and the sum of the estimated frequencies of the child nodes to the four child nodes.

[0107] Algorithm 4 RQT

[0108] Input the tree T and the spatial range query region Q

[0109] Output the query result

[0110]

[0111] Algorithm 4 shows the process of perturbed data aggregation and spatial range query. First, set the query area to be empty, mark all nodes in the tree as unvisited and set them to 1 (Steps 1 - Step 2). Traverse all nodes from top to bottom, mark the visited nodes as 0. When the node does not intersect with the query Q, it is directly ignored; if the node is completely contained by Q, add the estimated frequency of this point to the estimated frequency of the response result; if the node is not a leaf node but partially intersects with Q, traverse the child nodes of this node again; if the node is a leaf node and partially intersects with Q, calculate the overlapping part between this node and Q, and add the estimated frequency of the overlapping part to the result (Steps 3 - Step 17).

[0112] Perform privacy analysis

[0113] Prove that the LDP - ASDT method satisfies ε - local differential privacy

[0114] Let V be the node encoding of the tree, Z be the perturbed output result, and the user group be G, then we have

[0115]

[0116] In the formula, ω i (V i ) is the encoding value generated under different values.

[0117] The OUE method satisfies ε - local differential privacy. When any input is and the output is b, we have

[0118]

[0119] Therefore, both the LRR perturbation algorithm and the OUE method satisfy ε - local differential privacy, so the LDP - ASDT satisfies ε - local differential privacy.

[0120] The error of the LDP - ASDT method satisfies:[[]]

[0121]

[0122] Prove that the noise error of the LDP - ASDT method comes from the perturbation of OUE. Each time OUE is performed, a perturbation is required. At the same time, for the interval, if it is a dense space, there are more divided intervals, the number of queries used to answer is more, and the noise is higher. The non - uniform error comes from the distribution of the data set. If the data in the interval is uniformly distributed, there is no non - uniform error. Therefore, the finer the interval division, the smaller the non - uniform error, where the calculation of the non - uniform error

[0123] For the dense area, LDP-ASDT divides each interval into four parts separately by comparing the estimated value with the threshold, and uses ,, to represent the true frequency values of the four child nodes; during the spatial decomposition process, the decomposition of each interval is accompanied by noise error and non-uniform error. The frequency of the estimated node is evenly distributed to each child node to obtain the expected estimation error:

[0124]

[0125] where X is the added noise and f is the estimated frequency.

[0126] To balance the noise error and non-uniform error, the frequency threshold is set to

[0127]

[0128] where h is the specified height of the tree. This threshold ensures that the overall error of using this algorithm in the dense area is lower than that of the UG algorithm.

[0129] Example

[0130] Experimental settings

[0131] The experiment was written in the Python language using the PyCharm 2024.1.3 development platform, and the hardware environment was an Intel(R) Core i7-8565U CPU@1.80GHz processor and Windows 10 operating system

[0132] Experimental data: The dataset used three real geographical datasets: the user check-in dataset, the Beijing taxi dataset, and the Tokyo dataset. The specific parameters are shown in Table 1. Among them, the user check-in dataset comes from the Gowalla website and records the check-ins and location information of Gowalla users from February 2009 to October 2010. The Beijing taxi dataset covers the GPS trajectories of 10,357 taxis in Beijing from February 2 to 8, 2008, with a total of about 15 million points and a total distance of 9 million kilometers. The Tokyo dataset contains 573,703 Tokyo check-in information from April 12, 2012 to February 16, 2013.

[0133] Table 1 Dataset information table

[0134]

[0135] To verify the performance of the Local Differential Privacy-based Adaptive Grid Decomposition Algorithm LDP-ASDT, in the present invention, as a preference, the Mean-square error (MSE) is used to evaluate the gaps in the query range, the accuracy of the privacy budget, and the running time among the Grid-based Spatial Range Query (GT-R), the Private Adaptive Network (PrivAG), the Adaptive Network Decomposition based on Quad-tree (ASDQT), and LDP-ASDT. The formula for the mean-square error is:

[0136]

[0137] where Q is the query range, and q is any query within the query interval, is the estimated frequency of q in the query area, and f is the true frequency of q in the query area.

[0138] The impact of the privacy budget ε on MSE

[0139] The present invention evaluates the impact of different privacy budgets on the query error of the algorithm. As shown in Figures 2(a) to (i), the abscissa represents the privacy budget, and the ordinate represents the mean-square error MSE. It can be seen from the figure that when the privacy budget ranges from 0.1 to 0.9 for the three datasets, the mean-square errors of several algorithms change. From the changes in the experimental results, it can be seen that in all cases, as the privacy budget increases, the mean-square error of the present invention gradually decreases. The fundamental reason for this result is that as the privacy budget increases, the degree of privacy protection of the data by the algorithm decreases, resulting in a reduction in the degree of data perturbation and an increase in data accuracy. Therefore, the mean-square error of this method gradually decreases.

[0140] At the same time, as can be seen from Figures 2(a) to (i), the mean-square errors of other algorithms are all higher than that of this method, which indicates that the data availability of this method is relatively high. Specifically, compared with GT-R, the query accuracy of LDP-ASDT has increased by nearly two orders of magnitude, and there are also significant improvements compared with PrivAG and ASDQT. This is attributed to the fact that LDP-ASDT can not only define dense regions by setting appropriate thresholds, but also the algorithm itself can flexibly perform regional division: adopting a coarse-grained division method in sparse regions and a fine-grained division method in dense regions, which effectively reduces the query error. In addition, the performance of the LDP-ASDT algorithm remains stable on different datasets, while the performance of the ASDQT algorithm is greatly affected by the datasets. On datasets with relatively uniform data distribution and low skewness, ASDQT performs relatively well. However, since the LDP-ASDT algorithm can process data more effectively by dividing dense regions, its results are more stable.

[0141] The impact of the query range on MSE

[0142] As shown in Figures 3(a) to 3(i), the precisions of four algorithms under different query ranges are compared. In the experiment, the query ranges cover [5%, 50%], [10%, 55%], and [15%, 60%] of three datasets. To ensure the accuracy of the experiment, the query ranges are randomly queried 300 times, and the average value is taken as the final result.

[0143] As can be seen from Figures 3(a) to 3(i), when the privacy budget is fixed, the query range and the query precision are negatively correlated, that is, the smaller the query range, the relatively higher the query precision; the larger the query range, the lower the query precision may be. However, it can be found from Figures 3(a) to 3(b) that when the query range changes from [5%, 50%] to [10%, 55%], the mean square error decreases with the increase of the query range, that is, the query precision is greater. This is because the data distribution in the user check-in dataset is uneven. As the query range increases, the number of quadtree index nodes within the range becomes smaller, making the precision more accurate.

[0144] Comparison of running times

[0145] As shown in Figures 4(a) to 4(i), for three different datasets, as the privacy budget continuously increases, the running times of four algorithms are compared. The running time of this experiment is the time for space division, tree construction, and range query. It can be seen from the experimental results that the operation efficiencies of the two algorithms LDP-ASDT and ASDQT are basically stable in different datasets, while the running times of the GT-R and PrivAG algorithms are greatly affected by the privacy budget. This is because the grid granularities of the GT-R and PrivAG algorithms are affected by the privacy budget. The running time of LDP-ASDT is shorter than that of the ASDQT algorithm because the LDP-ASDT algorithm can divide out dense regions and further decompose them, and has better running efficiency in datasets with uneven or dense data distributions.

[0146] The privacy protection method for user location data query provided by the present invention aims at the privacy protection problem of user location data query. This method uses an adaptive grid division method to divide sparse regions and dense regions, solves the noise error caused by data collection, and on this basis, an adaptive hierarchical decomposition method is carried out. The quadtree index and frequency estimation method are used to further decompose the dense regions. Finally, an efficient post-processing technology is carried out for the frequencies of each node to provide the precision of data query. Through experiments and comparison with three algorithms, namely GT-R, PrivAG, and ASDQT, the results show that when facing datasets with uneven or dense data distributions, the LDP-ASDT algorithm has better query precision and running efficiency.

[0147] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and the examples shown and described herein.

Claims

1. A privacy protection method for user location data query, characterized in that: include: Step 1: Obtain user location data through sensors and group them based on random sampling technology so that the grouped user data meets the parallel combination of ε-localized differential privacy; Step 2: Decompose the spatial data based on the hierarchical decomposition method according to the spatial data distribution. In the first layer of decomposition, the space is divided into sparse areas and dense areas according to the set density threshold; Among them, the sparse area will no longer be divided. Each group of users in the sparse area will use the local random response method to perturb their location data, and send the perturbation results to the data collector to complete the privacy protection of the user location in the sparse area; For dense areas, a quadtree is constructed and post-processed. The dense areas are further divided by the quadtree index. Local random perturbations are performed on users at each layer, and the perturbation results are sent to the data collector for frequency construction. When the estimated frequency of the leaf node is less than the set frequency threshold or reaches the set tree height, the iteration is stopped to complete the privacy protection of the user location in the dense area. Step 3: Based on the RQT algorithm, complete the post-disturbance data aggregation and spatial data query.

2. The privacy protection method for user location data query according to claim 1, characterized in that: The step 2 also includes: Set the dataset to D, the privacy budget to ε, the number of grid cells to m, the tree height to h, the density threshold to 1, and the frequency threshold to θ; Randomly divide the data set D into n subsets to form multiple data blocks; Adaptively divide sparse and dense areas according to density threshold; Construct the root node V0, insert the divided network unit child nodes into the tree, and set the frequency estimate of the root node V0 to 1; Each user group locally perturbs its own location data using a local random response method and sends these perturbed results to the data collector; Compare the user density in the area with the density threshold. When the user density is greater than the density threshold, it is a dense area; When the user density is less than or equal to the density threshold, it is a sparse area; Perform quadtree index decomposition on dense areas, recursively until the estimated frequency of the byte node is less than or equal to the frequency threshold or reaches the set tree height; The tree is post-processed, the query results are returned, and the query frequency is returned.

3. The privacy protection method for user location data query according to claim 2, characterized in that: Also includes: In dense areas, each interval is divided into four parts by comparing the estimated value with the threshold. In the process of spatial decomposition, the decomposition of each interval is accompanied by noise error and non-uniform error. The frequency of the estimated node is evenly distributed to each child node, and the expected estimation error can be obtained, which satisfies the formula: Where X is the added noise, f is the estimated frequency, and Var is the variance of the estimated frequency of the leaf node; The non-uniform error is:

4. The privacy protection method for user location data query according to claim 3, characterized in that: The frequency threshold is calculated as: In the formula, h is the tree height, N is the range of data values, and e is the exponent.

5. The privacy protection method for user location data query according to claim 4, characterized in that: Also includes: Based on the optimal unary coding protocol, the perturbed vector b' is generated by perturbing the coding vector b. The perturbation rule is: In the formula, Pr[] is the probability of an event, p and q are probabilities, and b i For different value situations.

6. The privacy protection method for user location data query according to claim 5, characterized in that: In step 2, the data collector aggregates all the disturbance vectors collected and estimates the frequency of each value. Based on the encoding and disturbance method of OUE, the calculation formula for estimating the variance is: