Encrypted data query method

By 0-1 encoding and Bloom filter processing of the MBR of the R-tree index, a secure R-tree index is generated, and a search token is generated using the key of the pseudo-random seed, the security problem of multi-dimensional data query on the cloud platform is solved, and data privacy protection and security query are achieved.

CN120354439APending Publication Date: 2025-07-22JIUJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311505300.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-13
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing encrypted data query methods have poor security on cloud platforms, especially in multi-dimensional data query. Traditional solutions are difficult to effectively protect data privacy and have a risk of leakage.

Method used

The Bloom filter with 0-1 encoding and pseudo-random seeds encrypted and encoded the minimum boundary rectangle MBR of the R tree index to generate a secure R tree index, and generate a search token through the key of the pseudo-random seed to ensure the security of the query process.

Benefits of technology

It improves the privacy protection and security of multi-dimensional data in the cloud computing environment, ensures that data is not easily leaked when stored and retrieved on the cloud platform, and realizes secure data query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354439A_ABST
    Figure CN120354439A_ABST
Patent Text Reader

Abstract

The invention discloses an encrypted data query method, and relates to the technical field of data encryption. Comprising the steps that a data owner generates a secret key containing a pseudo-random seed and sends the secret key to a data user; the data owner constructs a common R-tree index, converts the common R-tree index into a safe R-tree index by using an index construction algorithm, and sends the safe R-tree index to the cloud platform; the data user obtains a search token according to the received secret key and the data query range needing to be queried, and sends the search token to the cloud platform; the cloud platform searches on the security R tree index according to the search token, and sends a search result to the data user; and the data user decrypts the search result to obtain plaintext data. According to the method, privacy protection and security of multi-dimensional data in the cloud computing environment are improved, so that a data owner and a user can safely store, retrieve and process sensitive data on a cloud platform, and data privacy and security challenges in the cloud computing environment can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data encryption, and particularly relates to an encrypted data query method. Background Art

[0002] Cloud computing has the characteristics of flexibility, scalability, reliability, convenience and low cost in providing computing resources. The powerful computing and storage capabilities and the reduced cost of cloud servers have led to a large amount of personal, organizational and enterprise data being stored in the cloud. These data usually contain users' private information, often including sensitive information, which inevitably raises concerns about data security issues. Encrypting data before outsourcing is one of the most direct methods, that is, encrypting the data before storing it in the cloud. Traditional encryption schemes are difficult to support some basic data operations. Although some novel encryption schemes have been proposed to support querying encrypted data, some of them can only process the ciphertext of one-dimensional data, while others are less efficient. Currently, there are mainly three types of secure and fast range query schemes for encrypting multi-dimensional data: order-preserving encryption schemes, bucketing schemes, and secure index schemes.

[0003] Agrawal et al. first proposed an order-preserving encryption (OPE) scheme. In the OPE scheme, the partial order among plaintexts is preserved in the ciphertexts, i.e., larger plaintexts correspond to larger ciphertexts. The main idea of the OPE scheme is to embed the order information into the ciphertexts of the data so that the order of the ciphertexts is consistent with that of the plaintexts. Specifically, for any data x > y, Enc(x) > Enc(y), where Enc represents the encryption algorithm in the OPE scheme. Therefore, the cloud can use the OPE scheme to support efficient range queries on ciphertexts. However, most current OPE schemes only support ciphertext queries for one-dimensional data and rarely involve ciphertext queries for multi-dimensional data. In addition, since the order information in the ciphertexts of the OPE scheme is revealed, attackers can use this information to accurately infer the corresponding plaintexts, leading to potential data security problems. This property of OPE allows for effective comparison and search of ciphertexts without decryption. However, this scheme lacks a formal security definition or analysis. Subsequently, Boldyreva et al. proposed a strict security definition for OPE. This security definition requires that the OPE scheme should not leak any information other than the partial order relation. However, Boldyreva et al. have shown that no OPE scheme can satisfy this strict security definition. Therefore, they proposed a weaker security definition, i.e., the ciphertexts are indistinguishable from the values computed by a random increasing function, and constructed an OPE scheme that satisfies this weaker security definition. After that, many researchers have conducted extensive related research based on the work of Boldyreva et al. However, most of these studies only consider OPE schemes for one-dimensional data and ignore the fact that there is a large amount of multi-dimensional data in the real world. Recently, Zhan et al. proposed an efficient multi-dimensional order-preserving encryption scheme, namely MDOPE. This scheme uses a network data structure to organize multi-dimensional data and uses prefix coding and Bloom filter techniques to process the values stored in the network data structure, thus realizing queries on encrypted multi-dimensional data.

[0004] Since the leakage of order information in the OPE scheme may be inevitable, in order to protect the order information and support ciphertext queries, a bucketing scheme is proposed to minimize such leakage. In the bucketing scheme, the data is grouped into buckets. All data in the same bucket are regarded as a whole. When a query matches a bucket, all data in that bucket are retrieved as the query response. Since the data within each bucket is indistinguishable, the query does not leak any order information inside the bucket. The researchers discussed the advantages of the bucketing scheme in terms of query security and efficiency. After that, many works have improved the bucketing scheme in many aspects. Lee proposed an ordered bucketing scheme to improve the search efficiency. In Lee's scheme, all buckets are organized in a certain order, i.e., all data in one bucket are less than all data in another bucket.

[0005] In previous bucketing schemes, the buckets had to be stored and searched on the data user side, which was very inconvenient. To solve this problem, Peng Wang et al. used Asymmetric Scalar Product Preserving Encryption (ASPE) to process the buckets and construct a secure index called . It can be stored on a remote cloud server for searching. Mei et al. also developed a secure bucket-based n-ary tree index to support range queries on ciphertexts of multi-dimensional data. However, their scheme is only applicable to uniformly distributed data sets.

[0006] Although some new encryption schemes have been proposed to support data retrieval on ciphertexts, these schemes still have their limitations. Some encrypted data query methods may fail to provide complete data protection, with vulnerabilities or weaknesses. For example, the same Minimum Bounding Rectangle (MBR) may have the same encrypted representation in different queries, which may allow attackers to obtain information about the data by monitoring query frequencies and patterns, resulting in data leakage and privacy violations. Summary of the Invention

[0007] The present invention provides an encrypted data query method, which can solve the technical problem of poor security when retrieving encrypted data on a cloud platform in the prior art.

[0008] The present invention provides an encrypted data query method, including:

[0009] The data owner generates a key containing a pseudo-random seed and sends the key to the data user;

[0010] The data owner constructs a normal R-tree index T, uses the index construction algorithm IndexGen to transform the normal R-tree index T into a secure R-tree index T*, and sends the secure R-tree index T* to the cloud platform;

[0011] The data user obtains a search token token according to the received key and the data query range Q to be queried Q , and sends the search token token Q to the cloud platform;

[0012] The cloud platform searches on the secure R-tree index T* according to the search token token Q , and sends the search result I * to the data user; the data user decrypts the search result I * to obtain the plaintext data;

[0013] The index construction algorithm IndexGen includes:

[0014] Encrypt the multi-dimensional data included in all leaf nodes of the normal R-tree index T;

[0015] Encrypt and encode the boundary information of each minimum bounding rectangle MBR of the ordinary R-tree index T using 0-1 encoding and introducing a pseudo-random seed to obtain the encoded form C of the minimum bounding rectangle MBR MBR .

[0016] Furthermore, the data owner generates a key containing a pseudo-random seed, including:

[0017] The data owner constructs a secure encryption scheme SE:

[0018] SE = (SE.Gen, SE.Enc, SE.Dec)

[0019] Among them, the SE.Gen module is used to generate keys, the SE.Enc module is used to encrypt data, and the SE.Dec module is used to decrypt data;

[0020] Call the SE.Gen module in the secure encryption scheme SE to generate the first key sk1:

[0021] sk1 = SE.Gen(1 λ )

[0022] Among them, λ is the security parameter and is the input of the SE.Gen module;

[0023] Select k pseudo-random seeds sd1, sd2,..., sd k as the second key sk2:

[0024] sk2 = (sd1, sd2,..., sd k );

[0025] Obtain the key SK according to the first key sk1 and the second key sk2:

[0026] SK = (sk1, sk2).

[0027] Furthermore, the ordinary R-tree index T includes:

[0028] Several leaf nodes, the minimum bounding rectangles MBR corresponding to several leaf nodes, and buckets connected to several leaf nodes;

[0029] The bucket is used to store multi-dimensional data to be encrypted, and the minimum bounding rectangle MBR is used to represent the range of the multi-dimensional data to be encrypted stored in the corresponding bucket.

[0030] Furthermore, the encryption of the multi-dimensional data included in all leaf nodes of the ordinary R-tree index T includes:

[0031] Call the SE.Enc module in the secure encryption scheme SE to encrypt the multi-dimensional data to be encrypted stored in the bucket connected to several leaf nodes.

[0032] Furthermore, encrypt and encode the boundary information of each minimum bounding rectangle MBR of the ordinary R-tree index T using 0-1 encoding and Bloom filters to obtain the encoded form C of the minimum bounding rectangle MBR MBR , including:

[0033] Obtain the boundary information of each minimum bounding rectangle MBR of the ordinary R-tree index T:

[0034] MBR = [a1, b1] × [a2, b2] ×... × [a d , b d

[0035] Among them, [a1, b1], [a2, b2], …, [a d , b d are the rectangular coordinates in the d-dimensional space, and [a i , b i is the range of the i-th dimension of the minimum bounding rectangle MBR, and a i and b i are the lower and upper limits of the range of [a i , b i ;

[0036] Convert the boundary information of each minimum bounding rectangle MBR of the ordinary R-tree into a binary string:

[0037]

[0038] Among them, is the form of encoding a i into a binary string, and is the form of encoding b i into a binary string;

[0039] Fill the random binary string of length l respectively after the binary string corresponding to each minimum bounding rectangle MBR to obtain After that, obtain

[0040] Use 0-1 encoding technology to encode the binary string with the random binary string to obtain The lower bound 1-type encoding of The upper bound 0-type encoding of

[0041] The second key sk2 = (sd1, sd2, …, sd​k ) and range lower bound type-1 encoding Input into a hash function to calculate a set of lower bound hash values:

[0042]

[0043] The second key sk2 = (sd1, sd2, …, sd k ) and range upper bound type-0 encoding Input into a hash function to calculate a set of upper bound hash values:

[0044]

[0045] Create two bit arrays and Set and each bit to 0 (i ∈ [1, d]), where d is the spatial dimension;

[0046] Set the bit at position in to 1, and set the bit at position in to 1, and output

[0047] Furthermore, the data user obtains a search token token according to the received key and the data query range Q to be queried Q , including:

[0048] The data owner constructs a search token generation algorithm TokenGen; the data user inputs the key SK = (sk1, sk2) and the query range Q = [p1, q1] × [p2, q2] × … × [p d , q d into the search token generation algorithm TokenGen to obtain a search token token corresponding to the query range Q = [p1, q1] × [p2, q2] ×... × [p d , q d ; the search token generation algorithm TokenGen includes: Q ;

[0049] Encode the query range lower bound p i and the query range upper bound q i into binary string forms

[0050] respectively and then append a random binary string of length l to obtain

[0051] Obtain 0 - encoding form of and 1 - encoding form of 0 - encoding form of and 1 - encoding form of

[0052] Input the second key sk2=(sd1, sd2,..., sd k ) and 0 - encoding form of into the hash function, input the second key sk2=(sd1, sd2,..., sd k ) and 1 - encoding form of into the hash function, and obtain:

[0053]

[0054]

[0055] Input the second key sk2=(sd1, sd2,…, sd k ) and 0 - encoding form into the hash function, input the second key sk2=(sd1, sd2,…, sd k ) and 1 - encoding form into the hash function, and obtain:

[0056]

[0057]

[0058] Output as the search token token for the query range Q Q .

[0059] Furthermore, the cloud platform searches on the secure R - tree index T* according to the search token token Q , including:

[0060] Obtain the inclusion relationship between the query range Q corresponding to the search token token Q and each minimum bounding rectangle MBR corresponding to this query range Q;

[0061] Search in the secure R - tree index T* according to the inclusion relationship:

[0062] If The secure R - tree index T *All the encrypted multi-dimensional data in the minimum bounding rectangle MBR is added to the search result set I * ;

[0063] If and then continue to iterate the minimum bounding rectangle MBR of the descendant nodes of the secure R-tree index T * ;

[0064] When Q intersects or covers the MBR of the leaf node, all the encrypted multi-dimensional data in the minimum bounding rectangle MBR of the secure R-tree index T * is added to the search result set I * ;

[0065] Furthermore, the data user decrypts the search result I * to obtain the plaintext data, including:

[0066] The data user inputs the first key sk1 and the search result set I returned by the cloud platform * into the SE.Dec module in the secure encryption scheme SE for decryption to obtain the plaintext I:

[0067] I = SE.Dec(I * , sk1).

[0068] The embodiment of the present invention provides an encrypted data query method. Compared with the prior art, its beneficial effects are as follows:

[0069] In the present invention, the data owner uses a Bloom filter with 0-1 encoding and a pseudo-random seed to encrypt and encode the boundary information of each minimum bounding rectangle MBR of the ordinary R-tree index T to obtain the encoded form C MBR . Since random numbers are introduced, the structure and key information of the index are not easily cracked, protecting the privacy of the index; and the search token sent by the data user to the cloud platform is generated based on the key containing the pseudo-random seed, which means that only the user with the pseudo-random seed key can submit a valid search request, ensuring the security of the query; the cloud platform can use the search token to search on the secure R-tree index T*. This process is based on the encrypted and encoded index, which enables the cloud platform to perform the search without knowing the plaintext data because both the search token and the index are encrypted, providing a guarantee for data privacy and security.

[0070] The present invention improves the privacy protection and security of multi-dimensional data in the cloud computing environment, enabling the data owner and user to securely store, retrieve, and process sensitive data on the cloud platform, ensuring that the data is not easily leaked when stored and processed on the cloud platform, which helps to protect the privacy of the data owner. Description of the Drawings

[0071] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention.

[0072] In the accompanying drawings:

[0073] Figure 1 is a schematic diagram of the system model provided in this specification;

[0074] Figure 2 is a schematic diagram of the MBR encoding provided in this specification;

[0075] Figure 3 is a schematic diagram of the construction of the secure R-tree index provided in this specification;

[0076] Figure 4 is a schematic diagram of the judgment of the query range and the MBR provided in this specification;

[0077] Figure 5 is the change in the number of index constructions for different schemes provided in this specification;

[0078] Figure 6 is the change in the search token generation time for different schemes provided in this specification;

[0079] Figure 7 is the change in the search time for different schemes provided in this specification. Detailed implementation manners

[0080] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. However, it should be understood that the protection scope of the present invention is not limited by the specific implementation manners.

[0081] Embodiment 1

[0082] Figure 1 is a schematic diagram of the system model of a secure and fast range query system for encrypting multi-dimensional data in this specification.

[0083] In the solution of the present invention, the data owner first constructs an ordinary R-tree index T covering all multi-dimensional data, and then performs 0-1 encoding and Bloom filter processing on the minimum bounding box MBR corresponding to each node in the R-tree. This can safely and effectively determine whether each processed MBR intersects with the query range. For each leaf node of the R-tree, its MBR corresponds to a bucket. Use a secure encryption algorithm to encrypt all multi-dimensional data within the MBR and store the encrypted data in the corresponding bucket. For convenience, the index described in the present invention is called a secure R-tree index. In addition, to ensure data security, the data owner uses a secure encryption scheme to encrypt all multi-dimensional data. When all MBRs are processed and all multi-dimensional data are encrypted, the data owner obtains a secure R-tree index and all encrypted multi-dimensional data.

[0084] The data owner outsources the secure R-tree index and the encrypted multi-dimensional data to the cloud center. For a query range, the data user uses his own key and the hash function in the Bloom filter to generate multiple hash values as search tokens and sends the search tokens to the cloud center.

[0085] After receiving the search tokens, the cloud center performs a range search on the secure R-tree index in a top-down order and returns the search results to the data user. Finally, the data user uses the same key as the data owner to decrypt all the ciphertexts in the search results. Specifically, it includes the following steps:

[0086] Step 1. The data owner establishes a secure R-tree index through the index construction algorithm IndexGen and encrypts all multi-dimensional data;

[0087] Index construction algorithm IndexGen(T) → T*: This algorithm takes the ordinary R-tree T as input and constructs a secure R-tree index T as output. This algorithm is executed by the data owner. The index construction algorithm IndexGen first calls the algorithm MBR encoding algorithm MBREncoding to process the MBRs of all nodes in the ordinary R-tree index T. Then, the index construction algorithm IndexGen calls the algorithm SE.Enc to encrypt all multi-dimensional data in the group under the leaf nodes of the ordinary R-tree T. Finally, the index construction algorithm IndexGen outputs the secure R-tree index T*. Therefore, in the secure R-tree index T*, each node contains a processed MBR, and each leaf node points to the group of encrypted multi-dimensional data covered by the MBR of this leaf node. Specifically, it includes the following steps:

[0088] Step 1.1. Use SE.Enc to encrypt all multi-dimensional data in the group under the leaf nodes of the ordinary R-tree index T. Key generation algorithm KeyGen(1 λ) → SK: This algorithm takes the security parameter λ as input and calculates a secret key SK as output. This algorithm is executed by the data owner. The detailed process of the algorithm is as follows.

[0089] Step 1.1.1, construct the secure encryption scheme SE:

[0090] SE = (SE.Gen, SE.Enc, SE.Dec)

[0091] Among them, the SE.Gen module is used to generate keys, the SE.Enc module is used to encrypt data, and the SE.Dec module is used to decrypt data.

[0092] Step 1.1.2, the key generation algorithm KeyGen calls SE.Enc to generate the key sk1:

[0093] sk1 = SE.Gen(1 λ )

[0094] Among them, λ is the security parameter, which is the input of the SE.Gen module, and sk1 = SE.Gen(1 λ ) is used to encrypt all the outsourced data before the data owner outsources the multi-dimensional data.

[0095] Step 1.1.3, select k pseudo-random seeds sd1, sd2,..., sd k as the second key sk2:

[0096] sd1, sd2,..., sd k

[0097] The key generation algorithm KeyGen selects k pseudo-random seeds sd1, sd2,..., sd k as the second part of the key sk2 = (sd1, sd2,..., sd k ), and these seeds are used for the hash functions in the Bloom filter. It should be noted that each hash function in the Bloom filter requires a value and a pseudo-random seed as input and outputs a hash value. Without sk2 = (sd1, sd2,..., sd k ), it is impossible to correctly calculate the hash value output by the hash function. Therefore, as long as all the MBRs in the Bloom filter are processed using this type of hash function, the security of the secure R-tree index can be ensured. For the same reason, as long as the query range is processed using this type of hash function in the Bloom filter, the security of the search token can be ensured.

[0098] Step 1.1.4, the key generation algorithm KeyGen outputs SK = (sk1, sk2) as the key.

[0099] Step 1.2: Use the MBR encoding algorithm MBREncoding to process the MBRs of all nodes in the ordinary R-tree T.

[0100] As Figure 2 shown, the MBR encoding algorithm MBREncoding(MBR) → C MBR : This algorithm takes an MBR as input and outputs the encoded form C of the MBR MBR . This algorithm is a sub-algorithm called by the index construction algorithm IndexGen. First, the MBR encoding algorithm MBREncoding extracts the boundary information of the MBR in the plane coordinate system, that is, R = [a1, b1] × [a2, b2]. Second, the MBR encoding algorithm MBREncoding calculates the binary strings of a1, b1, a2, and b2 respectively, selects four random binary strings, and fills these binary strings behind a1, b1, a2, and b2. Then, the MBR encoding algorithm MBREncoding calculates the range lower bound type 1 encoding of a1 and a2, and calculates the range upper bound type 0 encoding of b1 and b2. Next, the MBR encoding algorithm MBREncoding processes the type 0 encoding and type 1 encoding using a Bloom filter. Finally, the MBR encoding algorithm MBREncoding calculates the encoded form of R = [a1, b1] × [a2, b2], that is Specifically, it includes:

[0101] Step 1.2.1: The data owner constructs an ordinary R-tree index.

[0102] The R-tree is a balanced tree and can be used as an index structure for multi-dimensional data. Each node of the R-tree contains a minimum bounding rectangle MBR. The MBR of an internal node covers the union of the MBRs of its child nodes. Each leaf node is linked to a bucket, and all data covered by its MBR is stored in this bucket.

[0103] Let MBR = [a1, b1] × [a2, b2] ×... × [a d , b d , where [a i , b i is the range on the i-th dimension, d is the number of dimensions, a i and b i are the range lower bound and range upper bound of [a i , b i .

[0104] Step 1.2.2: The data owner converts each boundary information of the MBR into a binary string.

[0105] The MBR encoding algorithm MBREncoding converts a iEncoded in its binary string form Encode b i in its binary string form a i and b i are set to have length l. In particular, if or has a length less than l, the MBR encoding algorithm MBREncoding pads with 0s at the high - order positions of or .

[0106] Step 1.2.3: The data owner pads a random number after the binary string.

[0107] For security reasons, the MBR encoding algorithm MBREncoding randomizes and in their binary string forms. That is, the MBR encoding algorithm MBREncoding pads a random binary string of length l after i.e., That is By performing the same process, the algorithm also pads another random binary string of length l after i.e., That is Since and are both random binary strings of length l, and a i < b i , thus the value of .

[0108] Step 1.2.4: Use 0 - 1 encoding technology to encode the binary string with the random binary string to obtain the lower bound of the range of the type - 1 encoding of the upper bound of the range of the type - 0 encoding of

[0109] The definition of 0 - 1 encoding is: Given a binary string s = s n s n-1 ...s1 ∈ {0, 1} n , its 0 - encoding form is defined as the set Similarly, its 1 - encoding form is defined as Given two integers x and y, their 0 - 1 encoding forms can be used to determine whether x > y or x << y as follows. Let be the 0 - encoding form of x, be the 1 - encoding form of x, be the 0 - encoding form of y, is the 1-encoding form of y. If and only if , then x > y. On the contrary, if and only if , then x << y. For example: Given two data 11 and 6, their 4-bit binary strings are (1011)2 and (0110)2 respectively. It is easy to calculate the 0-1 encoding forms of 11 and 6, that is Since Therefore 11 > 6. On the contrary, since Therefore 6 ≤ 11.

[0110] Step 1.2.5, Create two bit arrays and where and Each bit in is initialized to 0 (i ∈ [1, d]). The data owner processes the 0-type encoding and 1-type encoding using a Bloom filter to determine the values in the two bit arrays and .

[0111] A Bloom filter is a probabilistic data structure used to test whether an element belongs to a set. It consists of three parts, including (i) a bit array A containing n bits, (ii) k independent hash functions h1, h2,..., h k , where h i : {0, 1} * → [1, n], i ∈ [1, k], (iii) a data set D = {d1, d2,..., d m} containing m different data. Given a data d′, through the following method, the Bloom filter can determine whether d′ belongs to D or d′ does not belong to D:

[0112] (1) The Bloom filter initializes all bits of the bit array A to 0.

[0113] (2) To add an element d j to the Bloom filter, the filter calculates the hash value h i (d j ) and sets the bit at position h i (d j ) in the bit array A to 1, where i ∈ [1, k], j ∈ [1, m].

[0114] (3) To test whether an element d′ is in the data set D, the Bloom filter calculates the hash values h i (d′) for k independent hash functions, where i ∈ [1, k]. If the bit at position h iIf the bit corresponding to (d′) is 1, then the data d′ is considered a member of the data set D. However, if the bit at any of these positions is 0, then d′ is definitely not in D.

[0115] It should be noted that the Bloom filter may produce false alarms, that is, incorrectly identify an element as a member of the set. However, through analysis, the probability of false positives can be minimized by setting appropriate parameters. Specifically, when k = (n / m)ln2, the Bloom filter has the minimum false positive probability of 2 -k . Where n is the number of elements in the data set and m is the number of bits in the bit array.

[0116] In the present invention, for data comparison, the MBR encoding algorithm MBREncoding calculates the type-1 encoding, denoted by , and then uses the hash functions h1, h2,..., h k in the Bloom filter and the second part of the key sk2 = (sd1, sd2,... sd k ) to process the binary string in. Specifically, the MBR encoding algorithm MBREncoding calculates a set of hash values and then sets the bits at the positions in to 1. Through a similar process, the MBR encoding algorithm MBREncoding calculates the type-0 encoding, denoted by . Then, the MBR encoding algorithm MBREncoding calculates a set of hash values and finally sets the bits at the positions in to 1.

[0117] Step 1.2.6, the MBR encoding algorithm MBREncoding calculates the encoded form of the MBR and outputs

[0118] Figure 3 shows the generation process of the secure R-tree index, as shown in Figure 3As shown in the figure, on the left (1), it indicates that all outsourced data is two-dimensional and distributed in a plane coordinate system, represented by hollow circles. In the middle (2), it shows that the data owner has established an ordinary R-tree index on these two-dimensional data. In an ordinary R-tree index, each node contains an MBR. Each leaf node points to a set of two-dimensional data. Specifically, node N1 contains MBR R1, leaf node N2 contains MBR R2 and points to a set of two-dimensional data D2, and leaf node N3 contains MBR R3 and points to a set of two-dimensional data D3. Then, the data owner runs the algorithm IndexGen. The algorithm IndexGen calls the sub-algorithm MBREncoding to process MBR R1, R2, and R3, and calls the algorithm R3 to encrypt all the two-dimensional data in D2 and D3. On the right (3), the processed MBRs are represented by and represent, the encrypted two-dimensional data set in D2 is represented by represent, and the encrypted two-dimensional data set in D3 is represented by represent. Finally, the algorithm IndexGen outputs a secure R-tree index.

[0119] Step 2: The data owner outsources the secure R-tree index and the encrypted multi-dimensional data to the cloud platform.

[0120] Step 3: The data owner distributes a key SK = (sk1, sk2) to the data user, and the data user uses this key to generate a search token for the query range.

[0121] Search token generation algorithm TokenGen(SK, Q) → token Q : This algorithm takes the key SK and the query range Q as inputs and outputs the search token token for the query range Q Q . The data user runs this algorithm to generate the search token token for the query range Q Q , and sends it to the cloud. Let the query range Q = [p1, q1] × [p2, q2] ×... × [p d , q d , where d represents the dimension, and [p i , q i is the range on the i-th dimension, and i ∈ [1, d]. The detailed steps of this algorithm are as follows:

[0122] Step 3.1: The algorithm TokenGen encodes the range lower bound p i into its binary string form, denoted as .

[0123] Step 3.2: After , append a random binary string of length l . Specifically, the algorithm converts the lower bound p of the range i to Then, using the same method, the upper bound q of the range is converted to i to where is the binary string form of p i , and is a random binary string of length l. Since and are both random binary strings of length l, and p i < q i , so is less than .

[0124] Step 3.3: The algorithm calculates the 0-1 encoding forms of , which are respectively denoted as and

[0125] Step 3.4: By using the hash functions h1, h2,..., h k in the Bloom filter and the second part of the key sk2 = (sd1, sd2,..., sd k ), the algorithm calculates and Calculate and

[0126] Step 3.5: The algorithm outputs and uses it as the search token token for the query range Q Q .

[0127] Step 4: The data user sends the search token to the cloud platform. After receiving the search token, the cloud platform performs a range search on the secure R-tree index and sends the search results to the data user.

[0128] Range search algorithm RangeSearch(token, T * ) → I * : This algorithm takes the search token token and the secure R-tree index T * as inputs and outputs the retrieved encrypted multi-dimensional data as the search result I * . The cloud runs the RangeSearch algorithm to perform a range search and returns the retrieved encrypted multi-dimensional data as a response to the data user.

[0129] First, introduce how to determine whether the query range Q intersects with the MBR. Then, introduce how to perform operations on the secure R-tree index T *Perform a range search on it.

[0130] Let Q = [p1, q1] × [p2, q2] ×... × [p d , q d be a query range, and MBR = [a1, b1] × [a2, b2] ×... × [a d , b d be the MBR in the secure R-tree index, where [p i , q i and [a i , b i are the ranges on the i-th dimension respectively, d represents the dimension, and i ∈ [1, d].

[0131] Step 4.1. Determine whether the query range Q intersects with the MBR

[0132] To support range search using the secure R-tree index, the algorithm should determine whether Specifically, if q i < a i or b i < p i , then That is On the contrary, Due to the adoption of 0-1 coding technology, if satisfies or then q i < a i or b i < p i , that is On the contrary, In addition, the algorithm also determines a special intersection, that is Specifically, if p i < a i and b i < q i , then That is On the contrary, Due to the adoption of 0-1 coding technology, if satisfies and then p i < a i and b i < q i , that is On the contrary, then

[0133] For example Figure 4As shown, the query range is Q = [p1, q1] × [p2, q2], and the MBR is R = [a1, b1] × [a2, b2]. When R is at position 1, since q1 < a1 (according to the 0-1 encoding technique, q1 < a1 means ), so there is That is Therefore, if Then there is Similarly, when R is at position 2, position 3, and position 4, since b1 < p1 (i.e., and ), b2 < p2 (i.e., and ), and q2 < a2 (i.e., and ). Except for the above four cases, Q and R have an intersection, so there is In addition, there is a special intersection between Q and R, that is When R is at position 5, since p1 < a1, b1 < q1, p2 < a2, and b2 < q2 (according to the 0-1 encoding technique, these four inequalities mean and ), so if it satisfies and Then there is Therefore, by using the above method, the algorithm RangeSearch can determine and

[0134] Note that to ensure the security of the 0-1 encoding, a Bloom filter with a special hash function is adopted. Since all MBRs and query ranges Q in the secure R-tree index T * have been processed by the 0-1 encoding technique and then by the Bloom filter, the algorithm RangeSearch can use the corresponding binary array and hash value to determine whether an MBR intersects or is covered by the query range. The specific details are as follows.

[0135] For the query range Q = [p1, q1] × [p2, q2] ×... × [p d , q d and MBR = [a1, b1] × [a2, b2] ×... × [a d , b d , if There exists a tuple in such that the binary array at h1(s, sd1), h2(s, sd2), …, h k (s, sdk All bits at the position of ) are 1, which means q i <a i . If There exists a tuple in such that the binary array at the positions of h1(s, sd1), h2(s, sd2), …, h k (s, sd k ) are all 1, which means p i >b i . If the algorithm RangeSearch determines that q i <a i or p i >b i , then there is That is On the contrary, if there exists then the algorithm RangeSearch can determine. If There exists a tuple in such that the binary array at the positions of h1(s, sd1), h2(s, sd2), …, h k (s, sd k ) are all 1, which means p i <a i . If There exists a tuple in such that the binary array at the positions of h1(s, sd1), h2(s, sd2), …, h k (s, sd k ) are all 1, which means q i >b i . If the algorithm RangeSearch determines that p i <a i and q i >b i , then there is That is On the contrary, if not, then

[0136] According to the above method, by determining whether all bits at the positions of the hash values in the corresponding Bloom filter array are 1, the algorithm RangeSearch can determine and

[0137] Step 4.2, Perform a range search on the secure R-tree index T * and

[0138] For ease of explanation, assume that N is a node associated with the MBR in the secure R-tree index. If the algorithm RangeSearch determines that then the algorithm adds all the encrypted multi-dimensional data in the MBR to the result set. If the algorithm RangeSearch determines that and then the algorithm continues to iteratively check the MBRs of the descendant nodes of N. When Q intersects or covers the MBR of a leaf node, the algorithm adds all the encrypted multi-dimensional data in the leaf node MBR to the result set. By using the search token the RangeSearch algorithm performs a range search in the secure R-tree index T * in a top-down manner. Finally, the RangeSearch algorithm returns the search result I * (i.e., the result set) as a response to the data user.

[0139] Step 5: The data user decrypts the ciphertext in the search result using the same key as the data owner.

[0140] Dec(SK, I * ) → I: This algorithm takes the retrieved ciphertext I * as input and decrypts the ciphertext using the key sk1 of the first part, i.e., I = SE.Dec(I * , sk1). The data user runs the algorithm Dec to decrypt the retrieved ciphertext I * and finally obtains the search result in plaintext form.

[0141] In the system model of the present invention, it is assumed that the cloud platform is semi-trusted, which means that the cloud platform follows the specified protocols and procedures but may be curious for various reasons, including being compromised to act on behalf of a third party.

[0142] Definition 1: Correctness. Given a query range Q, the cloud platform performs a range search using the secure R-tree index and returns the search result C * , where for each search result if the decrypted data d j falls within the minimum bounding rectangle (MBR) that intersects the query range Q, the range search scheme based on bucket classification is considered correct.

[0143] Definition 2: Security. Given a leakage function F, if all adversaries A cannot reveal more information than the leakage function F, then the SFRQ range search method is considered secure. The leakage function F(x, y) = positiondiff(x, y), where positiondiff(x, y) returns the position of the first difference between x and y.

[0144] Simulation experiment

[0145] As Figures 5-7 shown, the effects of the present invention can be further illustrated by the following test steps:

[0146] In the experiment, the scheme, the MDOPE scheme, and the SFRQ scheme of the present invention were compared. These schemes were implemented in Java on a personal computer equipped with an AMD Ryzen 5 2500U CPU and 8G RAM.

[0147] In the scheme, asymmetric scalar product preserving encryption (ASPE) was adopted and the scheme was implemented using Jama Library version 1.0.3. In the scheme experiment, some uniformly random two-dimensional data were selected to test the efficiency of the above schemes. In the scheme and the SFRQ scheme, the fan-out of the index was set to 6. This means that each two-dimensional range was divided into at most 6 smaller two-dimensional ranges. To achieve fairness in the experimental comparison, in the MDOPE scheme, each node on the first dimension contained only 1 split data, and each node on the second dimension contained 2 split data. This is because the range on the first dimension was divided into 2 smaller ranges by using 1 split data, and the range on the second dimension was divided into 3 smaller ranges by using 2 split data. According to the Cartesian product, in the MDOPE scheme, a two-dimensional range was divided into 6 smaller two-dimensional ranges. In addition, MDOPE supports exact range search. To fairly compare the MDOPE scheme, the scheme, and the SFRQ scheme, it was set that the MBR of each leaf node in the scheme and the SFRQ scheme contained only 1 data.

[0148] Index construction

[0149] As Figure 5 (a) shown, when the index height increases, the number of index constructions in the scheme, the MDOPE scheme, and the SFRQ scheme increases exponentially. As Figure 5 (b) shown, when the data quantity increases, the number of index constructions in the scheme, the MDOPE scheme, and the SFRQ scheme increases linearly. Compared with Compared with the MDOPE scheme, the index construction in the SFRQ scheme is more efficient.

[0150] The indexes in the [scheme], MDOPE scheme, and SFRQ scheme are all in tree structure. When the height of these indexes increases, the number of index nodes grows exponentially, resulting in an exponential growth in the index construction time. When the amount of data increases, more index nodes are needed to index this data. In the MDOPE scheme, the index needs to process each data. According to the experimental settings, in the [scheme] and SFRQ scheme, the index needs to process each MBR that contains only one data. Therefore, the index construction time in the [scheme], MDOPE scheme, and SFRQ scheme grows linearly. In addition, the index construction in the SFRQ scheme is the most efficient because the 0-1 coding in the SFRQ scheme is more efficient than the ASPE in the [scheme]. Although the prefix coding in MDOPE is also very efficient, it needs to process more segmented data. Therefore, the SFRQ scheme is more efficient than the MDOPE scheme.

[0151] Search token generation.

[0152] As Figure 6 shown, when the length of the bit string increases, the time for search token generation in the MDOPE scheme grows exponentially, while the [scheme] and SFRQ scheme have a very slow growth in the time for search token generation. The search token generation in the SFRQ scheme is the most efficient.

[0153] In the MDOPE scheme, the query range is first converted into a bit string. Assume the length of the bit string is l. Then, the MDOPE scheme fills an additional bit string after the original bit string. The length of the new bit string is 2l + 2. Next, the MDOPE scheme calculates the prefix coding of the new bit string. Finally, the MDOPE scheme processes the prefix coding of the new bit string using a Bloom filter to obtain the search token for the query range. In the above process, the slowest step is for the MDOPE scheme to calculate the prefix coding of the new bit string with a length of 2l + 2. In this step, each different bit string needs to be compared, and all 2 2l+2 different bit strings need to be merged into several bit strings. Therefore, the time for generating search tokens in the MDOPE scheme grows exponentially with the increase in the length of the bit string. In In the scheme, the query range is encrypted by using ASPE. The encrypted form of the query range is its search token. Since the efficiency of ASPE is independent of the bitstring length, the generation time of the search token is a constant. In the SFRQ scheme, the query range is first converted into a bitstring. Then, the SFRQ scheme pads an additional bitstring after the original bitstring. The length of the new bitstring is 2l. Next, the SFRQ scheme calculates the 0-1 encoding of the new bitstring. The total number of 0-1 encodings does not exceed 2l. Since the total number of bitstrings is very small, the search token generation in the SFRQ scheme is very efficient.

[0154] Range search.

[0155] As Figure 7 (a) shows, when the number of data is fixed at 10000 and the bitstring length increases, the search time remains almost unchanged. As Figure 7 (b) shows, when the number of data increases, the search times of the SFRQ scheme, scheme and the MDOPE scheme increase. Compared with the scheme and the MDOPE scheme, the SFRQ scheme is the most efficient.

[0156] As Figure 7 (a) shows, since the bitstring length is independent of the underlying encryption method ASPE, the scheme's search time remains almost unchanged. In the SFRQ scheme and the MDOPE scheme, when the bitstring length increases, the additional computational overhead is very low, which results in almost no increase in the search time. As Figure 7 (b) shows, when the number of data increases, the height of the index increases, resulting in scheme, the MDOPE scheme, and the SFRQ scheme need to perform more range searches on these indexes. Therefore, the search times of these schemes increase with the increase in the number of data. In the MDOPE scheme, many split data are inserted into the internal nodes of the index to support range search. Many comparison operations lead to low efficiency of range search. In addition, range search should be performed separately for each dimension. Therefore, the range search in the MDOPE scheme is not very efficient. Since the underlying hash value comparison method in the SFRQ scheme is more efficient than the scheme's ASPE scheme, the SFRQ scheme is more efficient than the scheme.

[0157] Theorem 1. The SFRQ scheme satisfies the correctness property defined in Definition 1.

[0158] Proof 1. Suppose there is an MBR = [a1, b1] × [a2, b2] ×... × [a d , b d, and there is a query range \(Q = [p_1, q_1]\times[p_2, q_2]\times\cdots\times[p d , q d . If holds, then there is In addition, the following equation holds.

[0159]

[0160]

[0161]

[0162] Therefore, if the query range \(Q\) intersects with the MBR, the MBR will be retrieved. Finally, all the ciphertexts in the leaf nodes of the secure index will be returned as search results. In conclusion, the SFRQ scheme is correct.

[0163] Theorem 2. The SFRQ scheme satisfies the security property defined in Definition 2.

[0164] Proof 2. Since the data is encrypted using a secure encryption method, the security of the data can be guaranteed by the encryption method.

[0165] In the SFRQ scheme, the data is encrypted by using a secure encryption scheme SE. The security of the MBR in the secure R-tree index can be analyzed as follows. The MBR is processed through padding, 0-1 encoding, and Bloom filters. The security of the data can be guaranteed by the security of the secure encryption scheme SE. Assume that (i) \(x = x_1x_2\cdots x n represents the boundary information of the MBR after filling with random values (as described in Section 5), and (ii) \(y = y_1y_2\cdots y n represents the boundary information of the query range after filling with random values (as described in Section 5). If and the member in the intersection of is \(t\), where the length of \(t\) is \(m\), it can be inferred that \(x_1 = y_1\), \(x_2 = y_2\), \(\cdots\), \(x m-1 = y m-1 , \(x m \neq y m . Therefore, the cloud only knows the leakage function \(F(x, y)=\text{position} diff (x, y)\). Therefore, the SFRQ scheme satisfies the security property defined in Definition 2.

[0166] Conclusion

[0167] The present invention proposes a range search scheme SFRQ. In the SFRQ scheme, a secure index for encrypted multi-dimensional data is constructed by using R-tree index, Bloom filter and 0-1 coding technology. Each node of the secure index is associated with a minimum bounding rectangle (MBR), and this secure index enables the cloud to quickly perform range queries on ciphertexts. The boundary information of the MBR is processed by 0-1 coding. By utilizing the characteristics of 0-1 coding, it can be determined whether the query range intersects with the MBR of the secure index node. The hash function in the Bloom filter is used to ensure the security of the query range and the MBR of the secure index node. Therefore, the proposed SFRQ scheme can support efficient range search for encrypted multi-dimensional data. A large number of simulation experiments have been carried out, and the results show that the cloud server can securely and efficiently perform range queries on the secure R-tree index in a top-down manner, and the results show its efficiency in procedures such as secure index construction, search token generation and range query execution. In addition, security analysis shows that no external entity including the cloud server can obtain additional information during the entire query process.

[0168] The embodiments described above merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent shall be subject to the appended claims.

Claims

1. A method for querying encrypted data, characterized in that, Including: The data owner generates a key containing a pseudo-random seed and sends the key to the data user; The data owner constructs a normal R-tree index T, uses the index construction algorithm IndexGen to convert the normal R-tree index T into a secure R-tree index T*, and sends the secure R-tree index T* to the cloud platform; The data user obtains a search token token according to the received secret key and the data query range Q that needs to be queried Q , and sends the search token token Q to the cloud platform; The cloud platform searches based on the search token Q on the secure R-tree index T*, and sends the search result I * to the data user; The data user decrypts the search result I * to obtain the plaintext data; The index construction algorithm IndexGen includes: Encrypting the multi-dimensional data contained in all leaf nodes of the normal R-tree index T; Encrypt and encode the boundary information of each minimum bounding rectangle MBR of the ordinary R-tree index T using 0-1 encoding and introducing a pseudo-random seed to obtain the encoded form C of the minimum bounding rectangle MBR MBR .

2. The encryption data query method according to claim 1, wherein The data owner generates a key containing a pseudo-random seed, including: The data owner constructs a secure encryption scheme SE: SE = (SE.Gen, SE.Enc, SE.Dec) Among them, the SE.Gen module is used to generate a key, the SE.Enc module is used to encrypt data, and the SE.Dec module is used to decrypt data; Call the SE.Gen module in the secure encryption scheme SE to generate the first key sk1: sk1 = SE.Gen(1 λ ) Among them, λ is a security parameter and is the input of the SE.Gen module; Select k pseudo-random seeds sd1, sd2,..., sd k as the second secret key sk2: sk2 = (sd1, sd2,..., sd k ); Obtain the key SK according to the first key sk1 and the second key sk2: SK = (sk1, sk2).

3. The encrypted data query method according to claim 1, wherein The normal R-tree index T includes: Several leaf nodes, minimum bounding rectangles MBR corresponding to several leaf nodes, and buckets connected to several leaf nodes; The bucket is used to store multi-dimensional data to be encrypted, and the minimum bounding rectangle MBR is used to represent the range of multi-dimensional data to be encrypted stored in the corresponding bucket.

4. The encryption data query method according to claim 2, wherein The encrypting the multi-dimensional data contained in all leaf nodes of the normal R-tree index T includes: Call the SE.Enc module in the secure encryption scheme SE to encrypt the multi-dimensional data to be encrypted stored in the bucket connected to several leaf nodes.

5. The encryption data query method according to claim 2, wherein Encrypt and encode the boundary information of each minimum bounding rectangle MBR of the ordinary R-tree index T using 0-1 coding and Bloom filter to obtain the encoded form C of the minimum bounding rectangle MBR MBR , including: Obtain the boundary information of each minimum bounding rectangle MBR of the normal R-tree index T: MBR = [a1, b1] × [a2, b2] ×... × [a d , b d ​ Among them, [a1, b1], [a2, b2], …, [a d , b d are rectangular coordinates in the d-dimensional space, and [a i , b i is the range of the i-th dimension of the minimum bounding rectangle MBR. a i and b i are the lower bound and upper bound of the range of [a i , b i ; Convert the boundary information of each minimum bounding rectangle MBR of the normal R-tree into a binary string: Among them, encodes a i into the form of a binary string, encodes b i into the form of a binary string; Random binary strings of length l are respectively filled in the binary strings corresponding to each minimum bounding rectangle MBR After that, obtain Encode the binary string with a random binary string using 0-1 coding technique to obtain Lower bound of the range of Type 1 coding Upper bound of the range of Type 0 coding Input the second key sk2 = (sd1, sd2,..., sd k ) and the lower bound 1-type encoding into the hash function to calculate a set of lower bound hash values: Input the second key sk2 = (sd1, sd2,..., sd k ) and the range upper bound type-0 encoding into a hash function to compute a set of upper bound hash values: Create two bit arrays and Initialize each bit in and to 0 (i ∈ [1, d]), where d is the spatial dimension; Set the bit at the position in to 1, set the bit at the position in to 1, and output 6. The encryption data query method according to claim 2, wherein The data user obtains a search token token according to the received secret key and the data query range Q to be queried Q , including: The data owner constructs a search token generation algorithm TokenGen; the data user inputs the secret key SK = (sk1, sk2) and the query range Q = [p1, q1] × [p2, q2] ×... × [p d , q d into the search token generation algorithm TokenGen to obtain a search token token corresponding to the query range Q = [p1, q1] × [p2, q2] ×... × [p d , q d ; the search token generation algorithm TokenGen includes: Q ; Lower limit p of the query range is separately i , upper limit q of the query range i is encoded in the form of a binary string respectively add a random binary string of length l obtain Obtained 0-encoded form and 1-encoded form 0-encoded form and 1-encoded form Input the second key sk2 = (sd1, sd2,..., sd k ) and in its 0-encoded form into the hash function, and input the second key sk2 = (sd1, sd2,..., sd k ) and in its 1-encoded form into the hash function to obtain: The second secret key sk2 = (sd1, sd2,..., sd k ) and the encoded form are input into a hash function, and the second secret key sk2 = (sd1, sd2,..., sd k ) and the encoded form are input into a hash function to obtain: Output As a search token token for the query range Q Q .

7. The encryption data query method according to claim 6, wherein The cloud platform searches based on the search token Q on the secure R-tree index T*, including: Obtain the inclusion relationship between the query range Q corresponding to the search token Q and each minimum bounding rectangle MBR corresponding to the query range Q; Search in the secure R-tree index T* according to the inclusion relationship: If add all the encrypted multi-dimensional data in the minimum bounding rectangle MBR of the secure R-tree index T * to the search result set I * ; If And then continue to iterate the minimum bounding rectangle MBR of the descendant nodes of the secure R-tree index T * ; When Q intersects or covers the MBR of a leaf node, add all the encrypted multi-dimensional data in the minimum bounding rectangle MBR of the secure R-tree index T * to the search result set I * .

8. The method for querying encrypted data according to claim 2, wherein The data user decrypts the search result I * to obtain plaintext data, including: The data user inputs the first secret key sk1 and the search result set I returned by the cloud platform * into the SE.Dec module in the secure encryption scheme SE for decryption to obtain the plaintext I: I = SE.Dec(I * , sk1).