Method for Querying Geometric Range of Encrypted Spatial Data Connection Based on Learning Index

By processing geometric range query of encrypted spatial data based on learning index, the problems of large index storage overhead and low query efficiency in the prior art are solved, and more efficient query performance and lower storage overhead are achieved.

CN116628298BActive Publication Date: 2025-05-27HEBEI UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310399488.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-14
Publication Date
2025-05-27
Estimated Expiration
2043-04-14

AI Technical Summary

Technical Problem

When handling encrypted spatial data, the existing geometric range query scheme has problems such as large storage overhead of index data and low query efficiency, especially in high-dimensional spatial data.

Method used

The encrypted spatial data connection geometric range query method based on learning index is adopted. By calculating the Z code of each spatial point, generating attribute vectors and index vectors, and encrypting the index vectors using predicate encryption, a privacy-protected learning index model is established, reducing the storage overhead of indexes and improving query efficiency.

Benefits of technology

While ensuring the security of index data and query privacy, it reduces the storage overhead of indexes and significantly improves query efficiency, which improves query efficiency by 4.1-5.2 times and 39.8-52.9 times compared with traditional solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116628298B_ABST
    Figure CN116628298B_ABST
Patent Text Reader

Abstract

The present invention provides a method for encrypted spatial data connection geometric range query based on a learned index. The method includes the following steps: First, the data owner calculates the Z-code of each spatial point and encrypts it, then generates an attribute value, and generates an index from the attribute value and the Z-code; at the same time, learns the data distribution of the encrypted Z-code, establishes a privacy-preserving learned index model, and sends the model and the index to the cloud server. During retrieval, the user generates a geometric vector based on the Z-codes within the geometric range, then generates a query attribute vector based on the retrieved attributes, connects the geometric vector and the query attribute vector, encrypts the generated query vector using predicate encryption, and finally generates a trapdoor using the decryption key and the encrypted Z-code of the midpoint within the geometric range, and sends it to the cloud server. The cloud server matches the trapdoor with the index and returns the query result to the user. Compared with the existing geometric range query methods, the method of the present invention not only improves the query efficiency, but also reduces the storage overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of geometric range queries for spatial data, and more particularly to a method for encrypted spatial data join geometric range query based on a learned index. Background Art

[0002] The popularity of location-based services has led to an explosive growth of spatial data, as smartphones and wearable devices can generate a large amount of geospatial data every day through daily applications such as Uber, Google Maps, and Facebook. To effectively process the massive spatial data, many enterprises and organizations provide users with geometric range query (GRQ) services that can return spatial data falling within geometric objects.

[0003] Existing geometric range query solutions can return spatial data falling within geometric objects. For example, users can find friends in a circular area around their location. For the storage of spatial data, it is common practice to outsource it to cloud service providers. However, outsourcing spatial data to cloud service providers can improve the search efficiency of the original spatial data but may lead to the leakage of users' sensitive information (such as location information). Therefore, to prevent data privacy leakage, data owners encrypt the data before outsourcing it to the cloud server. This limits the availability of spatial data because plaintext retrieval techniques cannot be used to query encrypted data. Subsequently, scholars have conducted a series of studies to design geometric range query solutions for encrypted spatial data to ensure data confidentiality without affecting data availability.

[0004] Most existing solutions only support geometric range queries, but users often want to query data points that satisfy certain attributes within the geometric range, so they cannot accurately express users' search intentions in practical applications. Recently, Kermanshahi et al. proposed two geometric range query solutions I and II by constructing a binary tree index and the Geo-DRS solution by constructing an R+ tree, achieving content privacy through symmetric homomorphic encryption and secret sharing. However, using traditional tree indexes makes the index data larger than the underlying data set, and this problem becomes more serious when the data has more spatial dimensions. In addition, a large number of internal tree nodes need to be accessed during the query process, affecting the search efficiency. Summary of the Invention

[0005] The object of the present invention is to provide a method for encrypted spatial data join geometric range query based on a learned index, which can reduce the storage overhead of the index and improve the query efficiency while ensuring the security of the index data and the privacy of the query.

[0006] The present invention is implemented as follows:

[0007] The geometric range query method for encrypted spatial data connection based on learning index provided by the present invention includes an initialization stage and a query stage.

[0008] In the initialization stage, first, the spatial data owner calculates the Z-code of each spatial point, then generates an attribute vector according to the attribute values of the spatial object, generates an index vector from the attribute vector and the Z-code, and encrypts the index vector using predicate encryption; encrypts the Z-code, learns the data distribution of the encrypted Z-code, and establishes a privacy-preserving learning index model. Finally, the model and the index vector are sent to the cloud server. In the query stage, the user generates a geometric vector based on the Z-codes within the geometric range, then generates a query attribute vector according to the searched attributes. After that, the geometric vector and the query attribute vector are concatenated to generate a query vector and encrypted using predicate encryption to obtain a decryption key. Finally, a trapdoor is generated from the decryption key and the encrypted Z-code of the midpoint of the geometric range and sent to the cloud server. After receiving the trapdoor, the cloud server matches the trapdoor with the index and finally returns the query result to the user. Compared with the existing geometric range query methods, the method of the present invention not only improves the query efficiency but also reduces the storage overhead.

[0009] In the above solution, the initialization stage includes two steps: spatial data preprocessing and index generation; the query stage includes two steps: trapdoor generation (GenTrap) and trapdoor retrieval (Query), which are specifically as follows:

[0010] a. Spatial data preprocessing;

[0011] a-1. Calculate the Z-code of the spatial data points;

[0012] Introduce the Z-order space-filling curve to reduce the n-dimensional space to a one-dimensional space, and convert each spatial data point into a unique Z-code; specifically: for a given geometric range, use the crossover operation to sort the spatial data points therein to obtain a space-filling curve; for a certain spatial data point, use the space-filling curve encoding of this point as the Z-code of this point;

[0013] a-2. Generate the attribute vector;

[0014] Generate an attribute vector according to the attribute values of the spatial object;

[0015] a-3. Encrypt the spatial data;

[0016] Use AES to encrypt the original spatial data and then upload it to the cloud server;

[0017] b. Generate the index;

[0018] b-1. Encrypt the index vector;

[0019] Generate an index vector by connecting the Z - code and the attribute vector of spatial data points, encrypt the index vector using predicate encryption, and then upload it to the cloud server;

[0020] b - 2. Construct a privacy - protected learning index model;

[0021] Encrypt the Z - code of spatial data points using a pseudo - random function, use an artificial neural network model to learn the data distribution of the encrypted Z - codes of spatial data points, construct a privacy - protected learning index model, and then upload it to the cloud server;

[0022] c. Generate a trapdoor;

[0023] c - 1. Generate a query vector;

[0024] A user with a query requirement generates a geometric vector based on the Z - codes within a geometric range and a query attribute vector based on the searched attributes; connect the geometric vector and the query attribute vector to obtain a query vector;

[0025] c - 2. Encrypt the query vector;

[0026] Encrypt the query vector using predicate encryption to obtain a decryption key and upload it to the cloud server;

[0027] c - 3. Generate a trapdoor through the decryption key and the encrypted Z - codes of the mid - points within the geometric range, and send it to the cloud server;

[0028] d. Trapdoor retrieval;

[0029] The cloud server matches the query results according to the trapdoor and returns the query results to the user. The user decrypts the returned vector using the decryption key to obtain the plain - text query results.

[0030] In the above solution, the specific process of generating the index in step b is as follows:

[0031] First, generate the master key and the message space, where the message space contains all possible plaintext messages to be encrypted. For the spatial data set in the geometric range, which contains spatial data points and the corresponding attribute values for each point, sort the spatial data points in the set using the Z-order space filling curve to obtain the Z-codes of each point; generate an attribute vector based on the attribute values, and concatenate it with the Z-code to obtain an index vector of length t + s, where t is the length of the Z-code and s is the dimension of the attribute vector. Then, use predicate encryption to encrypt the index vector. Input the master key, the index vector, and a plaintext message in the message space, and output the encrypted index vector. Input the key and the Z-code, use a pseudorandom function to encrypt the Z-code, and then construct a privacy-preserving learning index model, namely the PM learning index model, by learning the CDF (Cumulative Distribution Function) distribution of the encrypted Z-code.

[0032] Construct a privacy-preserving learning index model as follows:

[0033] By learning the data distribution of the encrypted Z-code, establish a hierarchical learning model, where each layer in the model is an artificial neural network; use the encrypted Z-code as the input and output the predicted position;

[0034] where L 1 is the loss function of the first layer, f 1 represents the model of the first layer, x is the Z-code to be queried, and y is the position corresponding to x;

[0035] where L i is the loss function of the j-th model in the i-th layer, the number of models in the i-th layer is |M| i ; N is the total number of spatial data points, N ∈ (0, ∞);

[0036] Each layer uses the loss function L i for recursive training, and finally obtain a complete model.

[0037] Add perturbations to the loss function as follows:

[0038] Expand the loss function using a polynomial and add noise to each coefficient; given the privacy parameter ε i , add the amount of noise η sampled from p(η), and the probability density function is:

[0039] where Φ i (x) is the obtained polynomial, and λ p is p(x) in the polynomial Φ iThe coefficient in (x), where N is the total number of spatial data points and N ∈ (0, ∞).

[0040] In the above solution, during the process of generating the trapdoor in step c, the query contains two elements: the geometric range and the query attribute set. For the geometric range, the present invention performs partitioning. Specifically: First, partition the entire plane region into multiple non-overlapping partitions, and then determine the point at the upper right corner and the point at the lower left corner in the geometric range. If they are in the same partition, encrypt the Z-codes of the two points at the upper right corner and the lower left corner to obtain the encrypted Z-code of the point at the upper right corner and the encrypted Z-code of the point at the lower left corner The encryption method is as follows:

[0041]

[0042]

[0043] where F is a pseudo-random function and K is a given secret key.

[0044] If they are not in the same partition, divide the query geometric range to ensure that the points at the upper right corner and the lower left corner in the geometric range corresponding to the sub-query are in the same partition, and then encrypt the Z-code.

[0045] Generate a query attribute vector according to the searched attributes, concatenate the geometric vector and the query attribute vector to obtain a query vector. Finally, through predicate encryption, input the master key and the query vector to output the decryption key SK. Encrypt the Z-codes of the points at the upper right corner and the lower left corner in the geometric range using a pseudo-random function, and use the encrypted Z-codes and the decryption key SK to generate the trapdoor TQ.

[0046]

[0047] Finally, send the generated trapdoor TQ to the cloud server.

[0048] In the above solution, the trapdoor retrieval in step d is specifically as follows: After receiving the trapdoor sent by the user, the cloud service first parses the trapdoor and inputs the obtained and into the PM learning index model, and uses the point query algorithm to obtain the predicted range, with the starting position being Pos start , and the ending position being Pos end .

[0049]

[0050]

[0051] Through the method of predicate encryption, the encrypted index vector within the prediction range is decrypted using the decryption key SK. If the decryption is successful, the match is successful, the encrypted index vector is added to the result list, and finally the spatial data corresponding to the result list is returned to the user.

[0052] The present invention introduces a learning index to design a connection geometric range query scheme based on the learning index. By learning the cumulative distribution function distribution (CDF distribution) of the encrypted Z-codes of spatial data points, a learning index model is established. The established PM learning index model is a hierarchical model, and there is no search process between layers. The position of the next layer is predicted through the input of the previous layer. When indexing, this model can be applied to execute an efficient point query algorithm to predict the query position; a differential privacy mechanism is also deployed in the model. By adding interference to the loss function, while ensuring the correctness of the output result, the leakage of effective information is reduced. After inputting the same encrypted Z-code, different predicted positions will be output, confusing the access pattern of the point query and improving the security of the query. Compared with traditional index structures such as R-trees, quadtrees, and binary trees, the learning index used in the present invention ECGRQ-LI does not search for child nodes through a large number of non-leaf nodes, thereby reducing the storage overhead of the index.

[0053] The present invention divides the query area and further improves the query efficiency by ensuring that the range of each query is within the same partition, reducing the access to the Z-codes of unnecessary data points; compared with some geometric range query schemes based on SSW encryption, the present invention ECGRQ-LI uses symmetric key predicate encryption to avoid time-consuming pairing operations; the query efficiency of ECGRQ-LI is 4.1 - 5.2 times and 39.8 - 52.9 times faster than the two most effective current schemes that support connection geometric range queries, respectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a system model diagram of the present invention.

[0055] Figure 2 It is a flowchart of the method for encrypted spatial data connection geometric range query based on the learning index of the present invention.

[0056] Figure 3 It is an execution flowchart of the method for encrypted spatial data connection geometric range query based on the learning index of the present invention among three entities.

[0057] Figure 4 It is a schematic diagram of the Z-order space filling curve of the present invention.

[0058] Figure 5 It is a schematic diagram of calculating Z-codes through crossover operation of the present invention.

[0059] Figure 6It is an example diagram for the index and query vector generation of the present invention.

[0060] Figure 7 It is a schematic diagram of the hierarchical learning index model of the present invention.

[0061] Figure 8 It is an execution flow chart between the user and the cloud server during the query phase of the present invention.

[0062] Figure 9 It is an example diagram of the conjunctive geometric range query of the present invention.

[0063] Figure 10 It is a schematic diagram of the training process of the PM learning index model of the present invention under different data sets.

[0064] Figure 11 It is a curve graph of the query accuracy of the present invention under different training rounds.

[0065] Figure 12 It is a curve graph of the search time of the present invention under different numbers of partitions.

[0066] Figure 13 It is a comparison graph of the search time between the ECGRQ-LI scheme of the present invention and the non-partitioned scheme.

[0067] Figure 14 It is a comparison graph of the index construction time between the ECGRQ-LI scheme of the present invention and the previously most effective schemes PBRQ-L and PBRQ-Q.

[0068] Figure 15 It is a comparison graph of the trapdoor generation time between the ECGRQ-LI scheme of the present invention and the previously most effective schemes PBRQ-L and PBRQ-Q. Detailed implementation manners

[0069] First, in order to improve the efficiency of the conjunctive geometric range query of encrypted spatial data and reduce the index storage overhead, the present invention proposes a privacy-preserving learning index structure. By expanding the objective function with polynomials and adding noise to each coefficient, this structure protects data privacy through the access pattern of fuzzy range queries and can avoid reconstruction attacks.

[0070] Then, an efficient conjunctive geometric range query (ECGRQ-LI) scheme based on the learning index is further designed. In particular, in order to reduce the search space and matching time, the concept of space partitioning is introduced to avoid accessing a large number of irrelevant Z-codes.

[0071] Finally, through simulation experiments and comparisons with the PBRQ-Q and PBRQ-L schemes that support the most effective solutions for connection range queries, it is concluded that the present invention has higher query efficiency and significantly reduced storage overhead.

[0072] The method for encrypted spatial data connection geometric range query based on learning index provided by the present invention (i.e., ECGRQ-LI) is applicable not only to traditional geometric range query scenarios, but also to query data points with certain attributes within the geometric range, and can more accurately express the search intention of users. In addition, the present invention also proposes an index model with higher query performance, uses differential privacy to improve the security of the model, and further reduces the query time by partitioning the geometric range.

[0073] The method for encrypted spatial data connection geometric range query based on learning index provided by the present invention starts from the perspective of predicting the query location through artificial intelligence, learns the distribution of data, trains the learning index model, and executes an efficient point query algorithm through this model, thereby improving the retrieval efficiency. Different from most geometric range query methods, it does not use the traditional tree-type index that needs to access a large number of nodes during query, but builds a hierarchical learning index, and selects the model of the next layer through the output of the previous layer, greatly reducing the retrieval time. To ensure the privacy of the query process, the Laplace mechanism that satisfies differential privacy is added during the training of the model; to facilitate the query, the plane area is partitioned to avoid accessing unnecessary data points; to speed up the matching process, predicate encryption is used to match the index vector and the query vector.

[0074] As Figure 1 shown, the solution of the present invention (ECGRQ-LI) mainly consists of three entities: the spatial data owner (SDo), the cloud service provider (CSP, or cloud server for short), and the spatial data user (SDu, or user for short). The SDo is mainly responsible for submitting spatial data to the CSP in ciphertext form. In addition, the SDo generates a secure learning index and sends it to the CSP, improving the retrieval efficiency. The CSP has powerful computing and storage capabilities and can store encrypted spatial data and indexes. When the CSP receives a connection geometric range query request from a legitimate SDu, it calculates the encrypted data and returns the corresponding search results to the SDu. The SDu generates a trapdoor and sends it to the CSP. After receiving the search results, the SDu uses the key sent by the SDo to decrypt these query results.

[0075] As Figure 2 shown, in order to protect the confidentiality of spatial data, the original spatial data of the SDo is encrypted using the AES algorithm and outsourced to the cloud server. In addition, the SDo generates the corresponding index information as follows: 1) Calculate the Z-code Z {p} of each spatial point p, and encrypt it as [Z {p}; 2) Generate an attribute vector A based on the attribute values of each spatial object p ; 3) For vector I p = Z {p} || A p Encrypt it through the PE.Encrypt algorithm to obtain the encrypted index IP * , and send it to the CSP; 4) Learn the data distribution of the encrypted Z - code, generate a PM learning index model, and upload it to the CSP. When SDu queries, first generate a query vector, and then determine whether the query range is in the same partition. If so, directly generate a query trapdoor and upload it to the CSP; if not, first partition the query range, and then generate a query trapdoor and upload it to the CSP; finally, the CSP performs result matching.

[0076] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below.

[0077] Combined with Figure 3 , in the initialization stage, the spatial data owner pre - processes the spatial data and uploads it to the cloud server. Specifically: The SDo converts the spatial data into a vector and encrypts it, and then uploads it to the CSP; The SDo generates a PM learning index and encrypts it, and then uploads it to the CSP. The initialization stage specifically includes the following steps:

[0078] Step 1: Fill the geometric space using the Z - order space - filling curve, calculate the Z - code of the spatial data points, and then encrypt the Z - code using the following encryption method:

[0079] [Z {p} = F(K, Z {p} ), where F represents a pseudo - random function, K represents a given key, Z {p} represents the Z - code of point p, and [Z {p} represents the encrypted value of the Z - code of point p.

[0080] The Z - order curve is a mapping between the n - dimensional space R n and the one - dimensional space R, which is defined as R n →R. If point p ∈ R n , then Z {p} ∈R, that is, Z {p} is the space - filling curve encoding of point p.

[0081] As Figure 4 shown, each spatial data point can be converted into a unique Z - code. Given a T×T grid, the length of the Z - code is where the Z - code is sorted and calculated by performing a cross - operation on the grid coordinates. Taking the point (3, 4) as an example, the transformation process is as Figure 5As shown. Therefore, the spatial data objects can be sorted according to the corresponding Z-codes. In particular, an artificial neural network model can be used to learn the data distribution of the Z-codes to construct a learning index.

[0082] Step 2: Generate a corresponding attribute vector according to the attribute value of the spatial object.

[0083] Step 3: Connect the Z-code and the attribute vector to obtain an index vector by the following method:

[0084] I p = Z {p} || A p , where Z {p} and A p are the Z-code and the attribute vector of point p respectively, and "‖" means connecting the two vectors of Z {p} and A p .

[0085] Suppose n attributes are extracted from the spatial dataset, and attribute i contains s i attribute values, then each spatial data object corresponds to an s-dimensional attribute vector A p , where As Figure 6 shown, when n = 3, s i = 2 (1 ≤ i ≤ 3). When the attribute value of spatial point p is the second value of the first attribute, the corresponding attribute vector is A p = (010000). In addition, if the Z-code of spatial point p is Z {p} = (011010), then the plaintext index vector I p = (011010010000).

[0086] After that, encrypt the index vector through the encryption function in predicate encryption.

[0087] The present invention uses symmetric-key predicate encryption as a cryptographic tool to protect the data privacy of the ECGRQ-LI scheme. The predicate encryption scheme can support conjunctive queries on encrypted data without revealing privacy, and it consists of four basic algorithms.

[0088] 1) PE.Setup(1 λ ) → MSK: Input the security parameter λ, and output the master key MSK and the message space χ.

[0089] 2) PE.Encrypt(MSK, Ip ∈ Γ t+s , μ ∈ χ) → Ip * : Input the master key MSK and the vector Ip = (Ip 1,..., Ip (t+s) ) ∈ Γ t+s , output the ciphertext Ip related to (Ip, μ) * , where Γ t+s ∈ {0, 1} is a finite set and does not contain wildcards.

[0090] 3) Input the key MSK and the query vector Output the decryption key SK, where contains wildcards.

[0091] 4) PE.Query(Ip * , SK) → μ(⊥): Input the ciphertext Ip corresponding to the index vector Ip * and the decryption key SK corresponding to the query vector Q. If then output μ; otherwise output ⊥. For each Ip ∈ Γ t+s , there is

[0092]

[0093] Step 4: Learn the cumulative distribution function of the encrypted Z - code based on the RIM model in deep learning, and improve the RIM model by using exponential queries to locate the correct position, and construct a learning index model. Through the function mechanism and Laplace mechanism in differential privacy, add random noise conforming to the Laplace mechanism to the loss function to construct the PM learning index model.

[0094] The present invention proposes a new privacy - protected learning index model, called the PM learning index model. Its core idea is to establish a learning model by learning the data distribution of the encrypted Z - code, predict the positions of the corresponding spatial data, and achieve efficient search and small storage overhead. For output perturbation, differential privacy technology is used to add noise to the loss function. Compared with directly adding noise to the output result, the present invention can use the optimization and adaptability of the learning model to correct the deviation of the optimal solution caused by noise addition. While ensuring the correctness of the output result, it interferes with the effective information contained in the output result as much as possible to reduce the leakage of private information.

[0095] As Figure 7 shown, the PM learning index model is a hierarchical learning model and does not cover the same number of records like a B - tree. The model of each layer is an artificial neural network with better performance.

[0096] The PM learning index model consists of the following three layers:

[0097] Input layer: The input layer is used to receive inputs such as feature vectors or keywords of data items. In the present invention, the input layer receives encrypted Z-codes.

[0098] Hidden layer: The hidden layer is trained using an artificial neural network and is the core part of the learning index model for learning the distribution and features of data items. The hidden layer contains multiple neuron nodes, which are connected and calculated through weights and biases in the neural network to obtain output results.

[0099] Output layer: The output layer is used to output the results of queries, usually including information such as the location or probability of data items. In the present invention, the output layer is used to output the predicted location.

[0100] There is a ReLU layer in the hidden layer, which is an activation function layer. The main role of the ReLU layer is to increase the non-linear ability of the learning index model.

[0101] The PM learning index model takes the encrypted Z-code as input, selects the model for the next layer, and repeats this process until the location is predicted for the last layer.

[0102] Assume that x is the encrypted Z-code to be queried, y is the location corresponding to x (y ∈ (0, m)), and the model in the first stage is denoted as f 1 , then the loss function L 1 is expressed as:

[0103]

[0104] Assume that the number of models in the i-th layer (i > 1) is |M| i , the j-th model in the i-th layer is denoted as f i (j), and the loss function of the model f i (j) is:

[0105]

[0106] where N represents the total number of spatial data points, N ∈ (0, ∞), l is the number of layers of the learning index model, and for 1 ≤ i ≤ l, each layer can be recursively trained using the loss function L i to establish a complete model.

[0107] To protect the privacy of spatial data, a direct approach is to add noise to the input and output layers of the PM learning index model to achieve differential privacy. However, the method of perturbing the input data generally uses differential privacy technology to synthesize data, which consumes a large amount of computing resources. In addition, directly interfering with the output results may lead to a large error rate. Therefore, a functional mechanism is considered to solve the above problems. The main idea is not to directly add noise perturbation to the output results, but to add it to the objective function of the optimization problem. The optimization and adaptation of the learning model can be used to correct the deviation of the optimal solution caused by noise addition. In addition, the functional mechanism is an extension of the Laplace mechanism and can be directly deployed on various objective functions. Since the loss function is the objective function of the optimization problem and plays an important role in the privacy of the entire model, the Laplace mechanism is used to perturb the function L i 。

[0108] The specific method is to expand the objective function with a polynomial and add noise to each coefficient. Suppose the obtained polynomial is Φ i (x), and the sensitivity is Δf, then where λ p represents the coefficient of the function p(x) in the polynomial Φ i (x), N represents the total number of spatial data points, N ∈ (0, ∞). In addition, given the privacy parameter ε i , the noise distribution is:

[0109]

[0110] The definition of differential privacy in the PM learning index model is: for any encrypted Z-code set of length m, [Z {P} ∈ Δ m , any adjacent training data sets D and D′ ∈ 2 2Δ* and any leakage L, if the following formula holds, then the PM learning index model satisfies ε-DP where

[0111]

[0112] Use the PM learning index model to perform secure and efficient spatial data point queries, which is the basis for spatial data range queries. For the point p, calculate the corresponding encrypted Z-code [Z {p} and input it into the PM learning index model for search. The estimated position is obtained at each layer for selecting the model of the next layer. There is a certain error in the position estimation process, that is, the position estimated in the last layer may not be the exact position of the query. However, it is very close to the true position of the data object, which can ensure the given search [Z {p}It must be within the range of [pos–min_error, pos+max_error], where min_error represents the minimum error and max_error represents the maximum error. Due to the use of differential privacy technology, the output position ranges for two identical inputs are different, confusing the access pattern of point queries.

[0113] In addition, compared with traditional index trees, there is no search process between the layers of the PM learning index model, which can avoid the low search efficiency caused by accessing a large number of internal tree nodes. The memory size of the PM learning index model has nothing to do with the dataset size m, which can solve the problem of large storage overhead of traditional tree indexes. Suppose the number of layers of the PM learning index model is l, and this model is a feedforward neural network with hidden layers, where the hidden layer has h neurons. The memory size of the PM learning index model is O(h).

[0114] Step Five: Upload the encrypted index vector and the PM learning index model to the cloud server for storage and query by users.

[0115] Combined Figure 3 with Figure 8 , the user includes the following steps in the query phase:

[0116] Step One: A user with a query requirement first generates a geometric vector according to the Z - codes within the query range.

[0117] Step Two: The user generates a query attribute vector according to the searched attributes.

[0118] Step Three: Connect the geometric vector and the query attribute vector in the following way to generate a query vector:

[0119] Q = Q G ||Q q , where Q G represents the geometric vector generated from the Z - codes, and Q q represents the query attribute vector.

[0120] Use predicate encryption to encrypt the query vector to generate a decryption key SK that can only be decrypted when the query vector matches the index vector.

[0121] When performing a conjunctive geometric range query, the user generates a geometric vector Q G according to the Z - codes within the geometric range G, generates a query attribute vector Q q according to the searched attributes, and then generates a query vector Q, as Figure 6 shown. When the user wants to search for the spatial object with the first value of the third attribute in the geometric range query G 1 , he / she first converts the query G 1 into Q G1=(010∪011)||(100∪101)=(01*10*), and then generate the query attribute vector Q q =(****10), obtaining the plaintext query vector Q = Q G1 ||Q q =(01*10*****10). Finally, use PE.GenToken to encrypt the query vector Q. The query is a process of matching the index vector with the query vector through PE.Query.

[0122] Step Four: Divide the planar region, divide the x-axis and y-axis into T x and T y parts, obtaining P D =[P D1 ,...,P Dθ , where θ = T x ×T y . If the query range is in the same partition, use the pseudorandom function to encrypt the Z-codes of the upper-right and lower-left points of the query range; if the query range is not in the same partition, divide the query range into multiple sub-query ranges in the same partition, and then use the pseudorandom function to encrypt the Z-codes of the upper-right and lower-left points of all sub-query ranges.

[0123] The present invention designs an efficient join geometric range query scheme. In particular, to reduce the search space and matching time, a PM learning index model is proposed to predict the location of spatial data, and a spatial segmentation method is proposed to avoid accessing unnecessary Z-codes during the query process.

[0124] Step Five: Generate a query trapdoor from the query vector and the encrypted Z-codes of the upper-right and upper-left points of the query geometric range, and upload it to the cloud server.

[0125] Step Six: After receiving the query trapdoor, the cloud server parses the encrypted Z-code pair of the upper-left and upper-right points in the trapdoor, inputs the encrypted Z-code pair into the PM learning index model, and reduces the search space of the query by executing the point query algorithm. Use the decryption key pair generated by the query vector to decrypt the encrypted index vector corresponding to the searched location. If the decryption is successful, put the encrypted index vector into the result list, and finally return the spatial data corresponding to the result list to the user.

[0126] The specific design details of the present invention are as follows:

[0127] 1) Setup(1 λ ) → Key: Given the security parameter λ, SDo runs PE.Setup(1 λ ), outputs the master secret key MSK, obtains the key K of the PRF pseudorandom function F, and sets Key = {MSK, K}.

[0128] 2) IndexBuild(Key, D) → {IP * , [I]}: SDo uses the master key MSK and the spatial data set D to generate the encrypted index vector Ip * and the encrypted PM learning index model [I]. The specific process is as follows:

[0129] Let the spatial data set D = {D 1 ,..., D i ,..., D m} (1 ≤ i ≤ m), where D i = {p i , a i}, p i is the spatial data point, is the attribute value of the spatial data object. For each D i = {p i , a i} ∈ D, SDo calculates the Z - code Z i of the spatial data point p {pi} , generates the attribute vector A i according to a Pi , and calculates for output and uploads it to the CSP.

[0130] Given the key K, for each point p i (1 ≤ i ≤ m), SDo runs F(K, Z {pi} ) to obtain the encrypted Z - code [Z {pi} . Then, constructs the PM learning index model [I] by learning the CDF distribution of [Z {pi} and uploads it to the CSP.

[0131] 3) TrapGen(SK, G, q) → TQ: Run by SDu, generates a trapdoor TQ for the query {G, q} through the master key MSK, where G is the geometric range and q ∈ A D is the query attribute set.

[0132] Sdu first converts the query G into a vector, and then uses Q q to generate the query attribute vector, thereby obtaining the plain - text query vector Q = Q G ||Q q . Finally, SDu encrypts Q to generate the decryption key SK ← PE.GenToken(MSK, Q).

[0133] To reduce the access to unnecessary Z - codes during the query process, first for the entire plane region P DPartition it. Specifically, divide the x-axis and y-axis into T x and T y parts respectively to obtain P D = [P D1 ,..., P Dθ , where θ = T x ×T y .

[0134] For the geometric range GS, if the upper-right point p rh and the lower-left point p ll of GS are in the same partition, then directly encrypt the Z code as and Otherwise, divide the query GS into {G S1 ,..., G Sγ} (1 ≤ γ ≤ θ) to ensure that the upper-right point p Si and the lower-left point p rh of each sub-query range G ll in GS are in the same partition. For each sub-query range G Si , SDu calculates the encrypted Z codes of point and point as and Finally, let As shown in Figure 9 , Figure 9 divides the plane into 4 4*4 partitions, T x = T y = 2. For the query G 1 , the query tokens and can be directly generated. For the query G 2 , it is divided into two sub-queries G S1 and G S2 , and then and

[0135] SDu generates the trapdoor and sends it to the CSP.

[0136] 4) Query(IP * , [I], TQ) → RList: The query process is as shown in the algorithm in Figure 11 . The CSP takes the encrypted index {Ip * , [I]} and the trapdoor TQ as inputs and outputs the query result list RList.

[0137] After receiving the trapdoor TQ sent by SDu, the CSP parses the TQ to obtain and For For each pair in and CSP uses the PM learning index model to perform point queries to narrow down the search space for join geometric range queries.

[0138] CSP performs queries through the query vector Q within the geometric range * Specifically, for each CSP performs flag← If Then add to the query result list Rlist Finally, CSP returns the spatial data corresponding to Rlist to SDu.

[0139] The present invention will be described in detail below with reference to actual examples. In this embodiment, the locations of post offices in the POST dataset are used as spatial data objects.

[0140] I. The spatial data owner generates an index.

[0141] Step 1: Use the AES algorithm to encrypt the information in the POST dataset and upload it to CSP.

[0142] Step 2: From the location information POS = {pos 1 , pos 2 ,..., pos n} in the POST dataset, generate the corresponding encrypted Z-code And use the Enron dataset to extract attribute values for each spatial data, thereby generating an attribute vector

[0143] Step 3: Connect the Z-code and the attribute vector of each spatial data object to generate an index vector, and encrypt the index vector in the following way:

[0144] PE.Encrypt(MSK, Ip, μ) → Ip * , where MSK represents the master key, Ip represents the index vector, μ represents a plaintext message in the message space, PE.Encrypt represents the encryption algorithm in symmetric key predicate encryption, and Ip * Is the generated encrypted index vector.

[0145] Step 4: Train the PM learning index model according to the encrypted Z-code, use a hidden layer model with 128 neurons, connect the ReLU layer for nonlinear transformation, the output layer represents the corresponding location of the spatial data, the learning rate is set to 0.0001, the batch training scale is 1000, and when deploying the differential privacy mechanism, the privacy budget ε is set to 0.4.

[0146] II. Query based on the encrypted data and index stored in the above description.

[0147] Step 1: Given the user's query request, which includes the query area and the attributes of the spatial data object. The size of the query area is 64×64 and there are three attribute values.

[0148] Step 2: Convert the user's query request into a query vector, and encrypt the query vector in the following way to obtain the decryption key for matching the results:

[0149] PE.GenToken(MSK,Q)→SK, where MSK represents the master key, Q represents the query vector, PE.GenToken represents the token generation algorithm in symmetric key predicate encryption, and SK represents the generated decryption key.

[0150] Step 3: Divide the planar area.

[0151] Step 4: Divide the query areas that are not in the same partition into sub-queries in the same partition, and encrypt the Z-codes of the upper-right corner point and the lower-left corner point; for queries in the same partition, directly encrypt the Z-codes of the upper-right corner point and the upper-left corner point.

[0152] Step 5: Generate a query trapdoor in the following way and send it to the cloud server.

[0153] where SK represents the decryption key, represents the encrypted Z-code of the lower-left corner point, represents the encrypted Z-code of the upper-right corner point, and TQ represents the query trapdoor.

[0154] Step 6: After receiving the query trapdoor, the cloud server parses the and in the trapdoor and inputs them into the PM learning index model to predict the location, and then uses the following method to determine whether there is a match:

[0155] flag←PE.Query(Ip * ,SK), where Ip * represents the encrypted index vector, SK represents the decryption key, PE.Query represents the query algorithm in symmetric key predicate encryption, and flag is used to indicate whether the match is successful.

[0156] If then add the corresponding encrypted index vector to the result list.

[0157] Step 7: The cloud server returns the corresponding spatial data in the result list obtained in Step 6 above to the user.

[0158] Next, the security of the ECGRQ-LI solution of the present invention is analyzed.

[0159] 1. Definition of leakage functions

[0160] Before giving the formal security definition of ECGRQ-LI, four leakage functions L size 、L precision 、L access and L search are first defined. L size ={n I , n Q} is the leakage of the size pattern, where n I represents the total number of encryption indices, and n Q represents the total number of connected geometric range queries (i.e., the trapdoor). L precision ={g, g'} is the leakage of the precision pattern, where g and g' are the code lengths of the spatial data points and the trapdoor. L access =Access(Q) is the leakage of the access pattern, where Access(Q)={D' 1 , D' 2 ,..., D' i}(i∈(1, m)) is the encrypted data returned for the encrypted connected geometric range query Q. L search =Search(Q) is the leakage of the search pattern, where Search(Q)={e 1 , e 2 ,..., e g’} represents the search pattern, and the adversary can learn whether the encrypted data points are returned by different query trapdoors. The value of e i is 1 or 0, indicating whether two queries retrieve the same key. For the index data set I and the query Q, let L(I, Q)={L size , L precision , L access , L search}.

[0161] The main security requirement of ECGRQ-LI is to protect the privacy of the query and index data stored on the untrusted CSP. On this basis, the security definitions of index privacy and query privacy of ECGRQ-LI with the above leakage functions are given, strictly following the Selective Chosen-Selective Chosen-Plaintext Attacks (IND-SCPA)

[0162] 2. Index privacy game

[0163] In the index privacy game, for two plaintext index data sets I 0and I 1 , a computationally bounded adversary A (i.e., an honest but curious server) can adaptively select a series of BuildIndex and TrapGen requests strictly following the leakage function L, but cannot distinguish the index set I 0 and I 1 . Similarly, for two query traps T q0 and T q1 , a computationally bounded adversary A can adaptively select a series of GenIndex and GenTrapoor requests strictly following the leakage function L, but cannot distinguish the two query traps T q0 and T q1 . Suppose Π=(Setup, IndexBuild, TrapGen, Query) is a probabilistic ECGRQ-LI scheme based on the security parameter λ. The security game between the adversary A and the challenger C is as follows:

[0164] Index privacy game Index A (λ)

[0165] Initialization: The challenger C runs Setup(1 λ ) to generate the key Key and keeps it secret. The adversary A sends two plaintext index data sets D 0 ={D 0,1 ,..., D 0,m} and D 1 ={D 1,1 ,..., D 1,m}, which have the same number of index records as the challenger C.

[0166] Request: The adversary A adaptively selects multiple query requests and sends them to the challenger C. These requests include the following two types:

[0167] Ciphertext: For the j-th ciphertext request, A sends D j ={D j,1 ,..., D j,i ,..., D j,m} to the challenger C, and the challenger responds

[0168] Trapdoor: For the j-th trapdoor request, A sends the connected geometric range query Q j ={G j , q j}. The challenger C responds with TQ j ←TrapGen(Key, Q j ). The query needs to satisfy the following condition: L(D 0 , Q j ) = L(D 1 , Qj ). For 1 ≤ i ≤ m, there is (D 0,i ∈ Q j ) ∧ (D 1,i ∈ Q j ) or

[0169] Challenge: For D 0 and D 1 selected during initialization, the challenger C randomly selects b ∈ {0, 1}, and then computes and sends it to A. The adversary A continues to adaptively submit a series of requests.

[0170] Guess: The adversary A guesses the value of b as b' ∈ {0, 1}.

[0171] Query Privacy Game Query A (λ):

[0172] Initialization: The challenger C runs Setup(1 λ ) to generate the key Key and keeps it secret. The adversary A sends two plaintext-connected geometric range queries Q 0 and Q 1 of the same dimension to C.

[0173] Request: The adversary A adaptively selects multiple query requests and sends them to the challenger C. These requests include the following two types:

[0174] Ciphertext: For the j-th ciphertext request, A sends D j = {D j,1 ,..., D j,i ,..., D j,m} to the challenger C. Subsequently, the challenger responds with . The query needs to satisfy the following conditions: L(D 0 , Q j ) = L(D 1 , Q j ). For 1 ≤ i ≤ m, there is (D j,i ∈ Q 0 ) ∧ (D j,i ∈ Q 1 ) or

[0175] Trapdoor: For the j-th trapdoor request, A sends the connected geometric range query Q j = {Gj, qj}. The challenger C responds with TQ j ← TrapGen(Key, Q j ).

[0176] Challenge: For Q 0 and Q selected during initialization1 , the challenger C randomly selects \(b\in\{0,1\}\) and computes and then sends it to A. The adversary A continues to adaptively submit a series of requests.

[0177] Guess: The adversary guesses the value of \(b\) as \(b'\in\{0,1\}\).

[0178] If the adversary A in the above game has at most the following negligible advantage, then the ECGRQ-LI scheme is defined to guarantee index privacy under IND-SCPA.

[0179]

[0180] Under IND-SCPA, when the adversary A in the game has at most the following negligible advantage, the ECGRQ-LI scheme can guarantee query privacy.

[0181]

[0182] 3. IND-SCPA Proof of the ECGRQ-LI Scheme

[0183] Since predicate encryption and pseudorandom functions are the underlying structures of ECGRQ-LI, the IND-SCPA security of ECGRQ-LI can be derived from the IND-SCPA of predicate encryption and pseudorandom functions.

[0184] Theorem 1: As long as predicate encryption is IND-SCPA secure, the ECGRQ-LI scheme is IND-SCPA secure.

[0185] Proof: Use the above security game to analyze the security of the ECGRQ-LI scheme. The main idea is to show that the security of ECGRQ-LI is equivalent to the security of the PE algorithm.

[0186] Initialization: The challenger C runs PE.Setup(1 λ ) to generate the predicate encryption master key MSK and the key K of the pseudorandom function, and keeps them secret. The adversary A sends two plaintext index data sets D 0 =\(\{D\) 0,1 ,..., D0,m \} and D 1 =\(\{D\) 1,1 ,..., D 1,m \} to the challenger C, where D 0,i =\(\{p\) 0,i , a 0,i \}, D 1,i =\(\{p\) 1,i , a 1,i \} (1\leq i\leq m).

[0187] Request: The adversary A adaptively selects multiple query requests and sends them to the challenger C. The requests are of two types:

[0188] Ciphertext: For the j-th ciphertext request, A outputs D j and sends it to the challenger C, where D j ={D j,1 ,..., D j,i ,..., D j,m}. For each D j,i , the challenger C generates Ip j,i ={Z {pj,i} ||A pj,i}(1 ≤ i ≤ n), and runs and [Z {p} ←F(K, Z {p} ), and obtains the learning index [I {p} by learning the data distribution of [Z j . Finally, the challenger C responds with the encrypted index vector and [I j .

[0189] Trapdoor: For the j-th trapdoor request, A outputs the connection geometric range query Q j ={G j , q j}, where G j and Q j can be represented as {(p rh , p ll ), Q j}. The challenger C encrypts Q j to obtain the query trapdoor SK j ←PE.GenToken(MSK, Q j ). In addition, the challenger C encrypts (p rh , p ll ) to obtain and and uses as the response. Moreover, the query needs to satisfy the following conditions. D 0,i ∈(p rh , p ll )∧D 1,i (p rh , p ll ) or For 1 ≤ i ≤ m, (<I p0,i , Q j > = 0)∧(<I p1,i , Q j>= 0) or (<I p0,i , Q j >!= 0) ∧ (<I p1,i , Q j >!= 0). For 1 ≤ i ≤ m, (D 0,i ∈Q j ) ∧ D 1,i ∈Q j ) or

[0190] Challenge: For D 0 and D 1 chosen during initialization, the challenger C randomly selects B ∈ {0, 1} and returns DB = (D b,1 ,..., D b,i ,..., D b,n ) to the adversary A. The adversary A continues to submit multiple requests adaptively.

[0191] Guess: The adversary A guesses that b is b'.

[0192] Therefore, the index security problem of ECGRQ-LI is successfully reduced to the data security problem of predicate encryption, and the probability that the adversary A can distinguish D 0 and D 1 is

[0193]

[0194] Therefore, as stated in Theorem 1, as long as predicate encryption is INDSCPA secure, ECGRQ-LI is index secure. Similarly, since the trapdoor is generated by predicate encryption, Theorem 2 is given.

[0195] Theorem 2: As long as predicate encryption and the pseudorandom function are IND-SCPA secure, the ECGRQ-LI scheme is IND-SCPA query confidential.

[0196] Proof: Similar to Theorem 1, the index security problem of ECGRQ-LI can be reduced to the data security problem of predicate encryption and the pseudorandom function. To avoid repetition, no detailed description is given.

[0197] 4. Security of the ECGRQ-LI Scheme under Reconstruction Attacks

[0198] The reconstruction attack is discussed here. By discussing the information leakage of a specific attack on the access pattern in the SSE scheme, its security is further elaborated. The attacker samples a sufficient number of queries and returns all index subsets that match the queries with high probability to determine the spatial data set i 1 , i 2 ,..., i n, where n represents the database size. Specifically, the attacker determines i through the maximum proper subset of all index sets 1 , and further determines i by finding the minimum proper superset of i 1 2 . Generally speaking, given i 1 , i 2 ,..., i j-1 , the attacker determines i by searching for the minimum proper superset of i 1 , i 2 ,..., i j-1 j . Experiments show that the attack can effectively recover the secret attributes of each spatial data, but it is carried out on the premise of the access pattern of the public query process. In the solution of the present invention, since the data user can hide the access pattern through differential privacy technology and balance security and query performance by adjusting the value of ε, for the same search request, the search results are different. Therefore, the attacker cannot judge whether the retrieval results come from the same request by analyzing the access pattern, and the ECGRQ-LI scheme can resist the reconstruction attack.

[0199] The performance of the ECGRQ-LI solution of the present invention is analyzed below.

[0200] 1. Theoretical analysis

[0201] First, the performance of the ECGRQ-LI scheme is analyzed theoretically, and compared with the existing schemes in terms of index construction overhead, trapdoor generation overhead, and query overhead.

[0202] Index construction overhead. For the IndexBuild algorithm of ECGRQ-LI, it first uses the predicate encryption algorithm (i.e., SHVE) to obtain Ip * = {I* p1 ,..., I* π ,..., I * π}. In addition, ECGRQ-LI trains the PM learning index model by learning the data distribution of the Z codes corresponding to the above data vectors, and reduces the search space. Therefore, the index generation time overhead of ECGRQ-LI is O(m)(T Enc + T PM ), where T Enc is the execution time of predicate encryption, m is the database size, and T PM is the generation time of the PM learning index model.

[0203] Trapdoor generation overhead. For the TrapGen algorithm of ECGRQ-LI, it concatenates the vector Q G and the vector Q qThus, the plaintext query vector Q is obtained, and the query vector Q is encrypted using the GenToken algorithm of predicate encryption to obtain SK, and by generating ([Z pll , [Z prh ) and using the PM learning index model to reduce the search space. Therefore, the cost of trapdoor generation is T Gen +2γT PRF , where t is the length of the Z code, s is the length of the attribute query vector, T Gen is the execution time of the predicate encryption trapdoor generation algorithm, and γ is the number of sub-query ranges.

[0204] Query overhead. For the query algorithm of ECGRQ-LI, CSP uses the PM learning index model to estimate the range of positions to be searched. For each index vector I * pi within the range, CSP executes PE.Query(I * pi , SK) to obtain the final query result. Therefore, the query cost of ECGRQ-LI is O(κ)T Dec , where κ is the number of Z codes within the position range, and TDec is the execution time of the predicate encryption query algorithm.

[0205] The solution of the present invention has great advantages in terms of query efficiency and storage overhead compared with the existing solutions. Specifically, GRSE-tree and FastGeo are based on SSW encryption. ECGRQ-LI mainly uses symmetric predicate encryption and pseudo-random functions to avoid time-consuming pairing operations and exponentiation operations. Therefore, the efficiency of the solution of the present invention is significantly higher than the above two solutions. In addition, due to the use of traditional index structures (i.e., R-tree, quadtree, and binary tree), the corresponding solutions search for appropriate child nodes through a large number of non-leaf nodes. ECGRQ-LI uses the learned index PM instead of the traditional index tree structure to perform queries. Since there is no search process between each stage of the learned index, the query cost is independent of the database size m. Therefore, the query cost of the above existing solutions is higher than that of the ECGRQ-LI solution of the present invention. Similarly, the use of traditional tree indexes also results in higher storage overhead for these solutions than the ECGRQ-LI solution of the present invention. In addition, compared with the existing solutions, only the solution of the present invention supports the connected geometric range query of n-dimensional spatial data and protects the access patterns of data users.

[0206] 2. Experimental result analysis

[0207] Implement the ECGRQ-LI scheme using Java and compare its performance with existing schemes. The scheme of the present invention realizes the dimensionality reduction preprocessing of spatial data and utilizes the code of the Z-order curve. The learning model is based on the code of the RIM model. The difference lies in using exponential queries to locate the correct position to improve efficiency. In addition, the JPBC library is used to implement the pairing operations of GRSE-tree and FastGeo, which are evaluated on the curve y 2 = x 3 + x, and the PRF evaluation is performed using SHA-256. The height of the BQ tree in the PBRQ-Q scheme and the R tree in the GRSE tree scheme is set to 16. All experiments are carried out on a computer with an Intel i7-6300U processor and an ubuntu20.04 operating system. The metric is the running time of the algorithm, and the average value of an algorithm running independently 30 times is taken as the experimental result. The POST dataset is used to demonstrate the performance of the scheme, and this dataset contains the locations of 123,593 post offices in the northeastern United States. Since the objects in the POST dataset do not contain text keywords, the Enron dataset is also used in this paper to extract attribute values for each spatial data object.

[0208] Index model training. First, test the performance of the PM index structure. After debugging, the model with a hidden layer has a better test effect. Therefore, in the experiment, a model with a hidden layer of 128 neurons is used, and the ReLU layer is connected for nonlinear transformation. The output layer represents the corresponding position of the spatial data. The learning rate is 0.0001, and the batch training scale is 1000. As Figure 10 shown, as the training process progresses, the learned model finally converges to a fixed loss value, which fits the data distribution. The training process is related to the dataset size m. The model can learn faster on a small data scale, and the larger the value of m, the slower the convergence speed of the model.

[0209] Further test the performance of performing point queries on the PM learning index model. For m = 105, the influence of the number of training rounds on the point query accuracy is as Figure 11 shown. As the number of training rounds increases, the query accuracy of the model under different noises gradually increases and tends to converge after 40 rounds. In addition, the smaller the ε value, the more noise is added. Therefore, as ε increases, the degree of privacy protection also increases. On the contrary, the accuracy of point queries decreases. Users can balance security and query performance by adjusting the value of ε.

[0210] Segmentation algorithm evaluation. Let the value of ε be 0.4. In this case, the values of min_error and max_error are 92 and 65 respectively. When the dataset size m is 2×10 4 ~1×10 5In the case of, the query performance of the ECGRQ-LI scheme was tested and compared with existing schemes. Before presenting the comparison results, it is necessary to verify the effectiveness of spatial partitioning and query partitioning. When the query region is a 64×64 square, the ECGRQ-LI scheme was debugged on different dataset sizes to obtain an appropriate number of partitions. As Figure 12 shown, for datasets of different sizes, the number of partitions corresponding to the optimal query performance is different. In addition, the query time of ECGRQ-LI does not always decrease as the number of partitions increases. The reason is that a large number of partitions will lead to an increase in the number of subqueries γ, that is, it will increase the number of accesses to the learned index PM. Therefore, a method for determining the number of partitions is provided. As shown in the following formula, when the number of partitions increases, the corresponding number of subqueries changes from γ to γ′, and the number of Z-codes accessed changes from to Their ratio R P can be used to evaluate the impact of the number of partitions on efficiency.

[0211]

[0212] Compare the query efficiency of the ECGRQ-LI scheme with the corresponding non-partitioned scheme. As Figure 13 shown, since the number of queries for Z-codes can be reduced after partitioning, when s = 256, K = 100, and m is in the range of 2×10 4 to 1×10 5 range, the ECGRQ-LI scheme has less query time.

[0213] Search performance. On the dataset, the ECGRQ-LI scheme and previous schemes, where the dataset size m varies from 2×10 4 to 1×10 5 . The size of the bloom filter of the GRSE-tree scheme was set to 1×10 3 . The query region is a 64×64 square with three attribute values. As shown in Table 1, the ECGRQ-LI scheme is much more efficient than the previous scheme. Due to the use of the learned index, the query efficiency of ECGRQLI is 4.1 to 5.2 times faster and 39.8 to 52.9 times faster than the most effective PBRQ-Q scheme and PBRQ-L scheme that support join geometric range queries, respectively. In addition, the present invention also removes the attribute vectors in ECGRQ-LI and compares it with existing geometric range query schemes. The query time of FastGeo is at least 200 times that of the ECGRQ-LI scheme of the present invention because FastGeo needs to perform a large number of pairing and exponentiation operations.

[0214] Table l

[0215]

[0216] ECGRQ-LI * Represents the scheme after deleting the attribute vector in the EcGRQ-LI scheme

[0217] Index construction time. Among Figure 14 the index construction time is compared with the previous most effective scheme. As can be seen from Figure 14 (a), as m increases, the index construction time of the three schemes all increases. This is because the larger m is, the longer the index vector will be, which affects the generation time of the encrypted index vector. Figure 14 (b) shows that the index construction time also increases linearly with the size of the attribute set, where m = 1×10 5 . The reason is that the number of bits of each attribute query vector increases with the increase of s, which affects the generation time of the encrypted index vector. In addition, the running time of the ECGRQ-LI scheme of the present invention is longer than that of PBRQ-L and PBRQ-Q because the learned index model requires additional training time. However, this does not affect the application of ECGRQ-LI because data users are more concerned about the response time of the query process, and the extra construction time compared with PBRQ-L and PBRQ-Q is tolerable.

[0218] The schemes in Table 1 are from the following literatures respectively:

[0219] The GRSE-tree scheme comes from the literature: B. Wang, M. Li, and H. Wang, “Geometric range searchon encryp ted spatial data,” IEEE Transactions on Information Forensics andSecurity, vol. 11, no. 4, pp. 704-719, 2015.

[0220] The FastGeo scheme comes from the literature: B. Wang, M. Li, and L. Xiong, “Fastgeo: Efficientgeometric range queries on encrypted spatial data,” IEEE transactions ondependable and secure computing, vol. 16, no. 2, pp. 245–258, 2017.

[0221] Schemes I and II are from the literature: S.K. Kermanshahi, S.-F. Sun, J.K. Liu, R. Steinfeld, S. Nepal, W.F. Lau, and M. Au, “Geometric range search on encrypted data with forward / backward security,” IEEE Transactions on Dependable and Secure Computing, pp. 698–716, 2022.

[0222] Schemes PBRQ-L and PBRQ-Q are from the literature: X. Wang, J. Ma, X. Liu, R.H. Deng, Y. Miao, D. Zhu, and Z. Ma, “Search me in the dark: Privacy-preserving boolean range query over encrypted spatial data,” in IEEE INFOCOM 2020 - IEEE Conference on Computer Communications. IEEE, 2020, pp. 2253–2262.

[0223] Trapdoor generation time. In Figure 15 the trapdoor generation time is compared with the previously most efficient scheme. As Figure 15 (a) and Figure 15 (b) show, the trapdoor generation time increases with the increase of m and s. This is because the number and length of query vectors also increase with the increase of the values of m and s, which increases the computational cost of the GenToken algorithm. In addition, the trapdoor generation time is related to the value of k, which is the number of range tokens in the PBRQ-L and PBRQ-Q schemes and represents the number of subqueries in the proposed scheme of the present invention. Therefore, when the values of m and s are fixed, the trapdoor generation times of the tested ECGRQ-LI, PBRQ-L, and PBRQ-Q schemes are shown in Figure 15 (c). It can be seen that the trapdoor generation times of the three schemes all increase with the increase of the k value. However, since the increase of the k value does not increase the length of the query vector and the predicate encryption time, the growth rate of the running time of TrapGen of the proposed ECGRQ-LI scheme of the present invention is slower.

[0224] To achieve efficient connected geometric range queries for encrypted spatial data, the present invention proposes the ECGRQ-LI algorithm. The novelty of the ECGRQ-LI scheme lies in the use of the proposed privacy-preserving learning index to reduce the search space and storage overhead. In particular, ECGRQ-LI expands the objective function by polynomial and adds noise to each coefficient to protect data privacy by interfering with the output results of the learning index. In addition, the concept of spatial segmentation is introduced to reduce the access to Z-codes during the query process. A large number of simulation results show that the search efficiency of the ECGRQ-LI scheme of the present invention is significantly higher than that of previous schemes, and the index storage overhead is also significantly reduced. Security analysis shows that ECGRQ-LI meets the IND-SCPA security requirements and can ensure data security and query privacy during the query process.

[0225] Contents not described in detail in this specification belong to the prior art well known to those skilled in the art.

[0226] The above description of the specific embodiments of the present invention is made in conjunction with the accompanying drawings and technical solutions. The above embodiments are only preferred embodiments given to fully illustrate the present invention. The protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.

Claims

1. A method for querying the geometric range of encrypted spatial data connections based on a learning index, characterized in that, it includes the following steps: a. Spatial data preprocessing; a-1. Calculate the Z-code of spatial data points; Introduce the Z-order space-filling curve to reduce the n-dimensional space to a one-dimensional space, and convert each spatial data point into a unique Z-code; Specifically: For a given geometric range, use the crossover operation to sort the spatial data points therein to obtain a space-filling curve; for a certain spatial data point, use the space-filling curve encoding of this point as the Z-code of this point; a-2. Generate an attribute vector; Generate an attribute vector according to the attribute values of spatial objects; a-3. Encrypt spatial data; Use AES to encrypt the original spatial data and then upload it to the cloud server; b. Generate an index; b-1. Encrypt the index vector; Connect the Z-code of spatial data points and the attribute vector to generate an index vector, use predicate encryption to encrypt the index vector, and then upload it to the cloud server; b-2. Construct a privacy-preserving learning index model; Use a pseudorandom function to encrypt the Z-code of spatial data points, use an artificial neural network model to learn the data distribution of the encrypted Z-codes of spatial data points, construct a privacy-preserving learning index model, and then upload it to the cloud server; c. Generate a trapdoor; c-1. Generate a query vector; A user with a query requirement generates a geometric vector according to the Z-codes within the geometric range, and generates a query attribute vector according to the searched attributes; connect the geometric vector and the query attribute vector to obtain a query vector; c-2. Encrypt the query vector; Use predicate encryption to encrypt the query vector to obtain a decryption key and upload it to the cloud server; c-3. Generate a trapdoor through the decryption key and the encrypted Z-code of the midpoint within the geometric range, and send it to the cloud server; d. Trapdoor retrieval; The cloud server matches the query results according to the trapdoor and returns the query results to the user. The user decrypts the returned vector with the decryption key to obtain the plaintext query results; In step b-2, to construct a privacy-preserving learning index model, specifically as follows: By learning the data distribution of encrypted Z-codes, establish a hierarchical learning model, and each layer in the model is an artificial neural network; use the encrypted Z-code as the input and output the predicted position; Among them, L 1 is the loss function of the first layer, f 1 represents the model of the first layer, x is the Z code to be queried, and y is the position corresponding to x; where L i is the loss function of the j-th model in the i-th layer, the number of models in the i-th layer is |M| i ; N is the total number of spatial data points; The loss function L is used for each layer i to perform recursive training and finally obtain a complete model.

2. The method for querying the geometric range of encrypted spatial data connections based on a learning index according to claim 1, characterized in that, in step b-2, add perturbations to the loss function, specifically as follows: Expand the loss function as a polynomial and add noise to each coefficient; given the privacy parameter ε i , add the amount of noise η sampled from p(η), where the probability density function is: Among them Φ i (x) is the obtained polynomial, and λ p is the coefficient of p(x) in the polynomial Φ i (x), and N ∈ (0, ∞).

3. The method for querying the geometric range of encrypted spatial data connections based on a learning index according to claim 1, characterized in that, in step c-1, to generate a geometric vector according to the Z-codes within the geometric range, specifically as follows: Divide the plane area into multiple non-overlapping partitions. For the queried geometric range, if the points in the upper right corner and the lower left corner of the geometric range are in the same partition, directly encrypt the Z-codes of these two points; If the points in the upper right corner and the lower left corner are not in the same partition, the geometric range of the query is divided into multiple geometric ranges corresponding to sub-queries to ensure that the points in the upper right corner and the lower left corner in the geometric range corresponding to the sub-query are in the same partition. Then, the Z-codes of the points in the upper right corner and the lower left corner in the geometric range corresponding to the sub-query are encrypted respectively.

4. The method for querying the geometric range of encrypted spatial data connection based on a learning index according to claim 1, characterized in that, in step d, the cloud server matches the query result according to the trapdoor, specifically as follows: The cloud server parses the encrypted Z-code in the trapdoor, inputs the encrypted Z-code into the privacy-preserving learning index model to obtain the predicted location; decrypts the encrypted index vector using the decryption key corresponding to the query vector, and if the decryption is successful, returns the encrypted index vector to the user.

Citation Information

Patent Citations

  • Spatial range query method based on extensible learning index

    CN112035586A

  • Database query method and system having access control function

    WO2018113563A1