Order-preserving ciphertext indexing method and system supporting forward security

By optimizing segmentation functions and building a new index structure, the problem of inefficient order-preserved ciphertext indexing under forward security is solved, efficient data query and security improvement is achieved, and it is suitable for order-preserved encryption schemes with forward security.

CN120281482APending Publication Date: 2025-07-08SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510368045.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing order-preserving encryption scheme cannot support order-preserving ciphertext indexing with forward security, resulting in inefficient data query, especially the time to conduct range query on cloud servers.

Method used

A new segmentation function fsplit(data) is designed to optimize the uniformity of subset size and segmentation intervals, and a new index structure is constructed, and a forward-safe order-preserved ciphertext index is achieved by adding a token state key chain between the hash table and the table jump structure.

Benefits of technology

It improves the query efficiency of order-preserved ciphertext data, can distinguish historical data from new data, improves security performance, and flexibly weighs the size of subsets and segmentation intervals by adjusting parameters, suppressing the influence of noise and outliers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281482A_ABST
    Figure CN120281482A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data encryption processing, and provides an order-preserving ciphertext indexing method and system supporting forward security. According to the technical scheme, firstly, a new segmentation function is designed, the uniformity of the size of a subset and the rationality of a segmentation interval are balanced through a global optimization objective function, and a new segmentation function is designed; a segmentation strategy can be adjusted according to data density and abrupt change points, and the size of a subset and a segmentation interval can be flexibly weighed through lambda; secondly, a new index structure is constructed, the order-preserving ciphertext can be inquired in a personalized mode according to the state of the multi-user token, and the index function of forward security is supported; the method comprises the following steps: mapping ciphertext data and token states in a user query token to a target address on a token state key chain, forming a search token, sending the search token to carry out data retrieval, carrying out data retrieval according to the target address, and carrying out retrieval on a multi-level index to finally obtain a result set of target data. And the retrieval efficiency is considered while the security is considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data encryption processing, and particularly relates to an order-preserving ciphertext index method and system supporting forward security. Background Art

[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.

[0003] Currently, many order-preserving encryption schemes do not support order-preserving ciphertext indexing, which reduces the practicality and query efficiency of data. When a user initiates a range query (such as: salary>8000 and salary<15000), if the data set in the cloud server is very large, this undoubtedly greatly increases the time overhead for the user to retrieve the target data, bringing a very poor user experience. Implementing an order-preserving ciphertext index structure on the cloud server will greatly improve the user experience. It is necessary to propose an efficient index structure supporting order-preserving ciphertext to improve the retrieval experience of order-preserving ciphertext.

[0004] The simplest and most direct way to implement the index function of the cloud server for order-preserving ciphertext is to refer to the implementation of commercial databases. Currently, the mainstream index types of databases include B+ tree index, hash index, etc. Among them, the B+ tree index is the most commonly used index type in relational databases, and it is the default index implementation method in mainstream databases such as MySQL and Oracle; as Figure 1 shown in the index structure based on the B+ tree, the B+ tree index is suitable for range queries and sorting operations. Its core principle is a multi-way balanced tree structure. All data is stored in leaf nodes, and the leaf nodes are connected by a linked list. Non-leaf nodes only store index key values, and it also supports creating a B+ tree index structure for a specified field.

[0005] The hash index is the most commonly used index type in non-relational databases and is implemented based on a hash table. Its principle is to store the index key value after being calculated by a hash function. A significant feature is that the equal-value query is extremely fast, but it does not support range queries. There are relevant literatures proposing to achieve the aggregation of order-preserving ciphertext by using a special constructed hash function based on the skiplist data structure, as Figure 2 shown in the hash index structure based on the skiplist. In this way, the hash function retrieves data in a pseudo-binary tree search manner, improving the index of the hash structure. The skiplist structure is combined with the hash table, and a (max, min) data block is maintained under each skiplist structure, which improves the retrieval efficiency of the equal-value query while solving the problem of not supporting range queries.

[0006] The inventor found that the above solution cannot be applied to the ordered ciphertext data with forward security, and it is impossible to distinguish historical data from newly added data. Therefore, in the forward-secure ordered encryption scheme, this index structure cannot improve the retrieval efficiency. Summary of the Invention

[0007] In order to solve at least one of the technical problems existing in the above background technology, the present invention provides a forward-secure ordered ciphertext index method and system, which designs a forward-secure ordered ciphertext index method that simultaneously considers security and efficiency, and improves the security and efficiency of retrieval.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] The first aspect of the present invention provides a forward-secure ordered ciphertext index method, including the following steps:

[0010] Obtain a given sorted data set, determine the optimal splitting point by combining the interval of the data, the subset size, and the constructed data set splitting function, generate a hash table according to the optimal splitting point, where the key in the hash table is the splitting point hash value, and the value in the hash table is a pointer pointing to the head node of the token state key chain;

[0011] Construct a token state key chain, where each node stores the hash value of the token, and at the same time each node is associated with a secondary index skip list structure;

[0012] Calculate the split sub-data set to which the ciphertext belongs according to the user query ciphertext and the token, and obtain the split point hash value according to the split point to which the sub-data set belongs; obtain the head node address of the token state key chain according to the split point hash value, and calculate the token hash value according to the state of the token;

[0013] Traverse and retrieve according to the head node address of the token state key chain, find the node corresponding to the token hash value, and obtain the down pointer of the node;

[0014] For the query ciphertext, query the secondary index skip list structure in combination with the down pointer to obtain the query result.

[0015] Further, the step of obtaining a given sorted data set and determining the optimal splitting point by combining the interval of the data, the subset size, and the constructed data set splitting function includes:

[0016] For the sorted data, select the minimum value and the maximum value in the data set;

[0017] Generate k + 2 equally spaced integers within the interval of the minimum value and the maximum value to generate a set of pre-splitting points. After removing the endpoints, take the middle k integers as the pre-splitting points;

[0018] Use the greedy strategy to uniformly sample between the left and right boundaries of the current pre-split point integer candidate points to calculate the total cost of the dataset splitting function. When the total cost is minimized, the corresponding split point is the optimal split point. Further, the dataset splitting function is:

[0019]

[0020] Among them, the first term |S i | represents the size of the i-th subset after splitting. The second term represents the optimization of the interval between split points. By adjusting the hyperparameter λ, the trade-off between the split point interval and the subset size is controlled. arg represents finding the independent variable at which the function obtains a certain specific value, represents finding the minimum value of the function in a certain set or range; therefore represents the independent variable at which the function obtains the minimum value. The independent variable is the set of split points S = {x1, x2,... x k},S i represents the i-th subset. Among them, x j and x j+1 represent the positions of the j-th and j + 1-th split points, and B n represents the expected size of each subset.

[0021] Furthermore, in the secondary index skip list structure, the node value in the skip list is the ciphertext. First, the elements of the split sub-dataset are encrypted in order, and then XORed according to the hash value of the token node to which they belong. All elements are linked into a bottom-level linked list in an ordered form, and a secondary index is generated by taking points at intervals on the bottom-level linked list, and a primary index is generated by taking points at intervals on the secondary index.

[0022] Furthermore, for querying the ciphertext, the secondary index skip list structure is queried in combination with the down pointer to obtain the query result, including:

[0023] Find all values greater than the query ciphertext. Find the last node less than or equal to the ciphertext x in the primary index, find the secondary index through the down pointer, and then find the last node less than the ciphertext x in the secondary index. Through the downward pointer, come to the bottom-level linked list. In the bottom-level linked list, find the first node target that satisfies being greater than the ciphertext x, and then return the data corresponding to all nodes after the target.

[0024] Furthermore, calculate the split sub-dataset to which the ciphertext belongs through the aggregation function. The aggregation function is:

[0025]

[0026] Among them, x is the query ciphertext, and x i is the split-point ciphertext, and B c is the size of the sub-dataset set after the dataset is split. BinarySearch(x, 1, B c - 1) is a recursive function.

[0027] The second aspect of the present invention provides an order-preserving ciphertext index system that supports forward security, including:

[0028] A data splitting module, which is used to obtain a given sorted dataset, determine the optimal split point by combining the interval of the data, the subset size, and the constructed dataset splitting function, generate a hash table according to the optimal split point, where the key in the hash table is the split-point hash value, and the value in the hash table is a pointer pointing to the head node of the token state key chain;

[0029] A data mapping module, which is used to construct a token state key chain, and each node stores the hash value of the token, and at the same time each node is associated with a secondary index skip list structure;

[0030] A data retrieval module, which is used to calculate the split sub-dataset to which the ciphertext belongs according to the user query ciphertext and the token, and obtain the split-point hash value according to the split point to which the sub-dataset belongs; obtain the head node address of the token state key chain according to the split-point hash value, and calculate the token hash value according to the state of the token; traverse and retrieve according to the head node address of the token state key chain, find the node corresponding to the token hash value, and obtain the down pointer of the node; for the query ciphertext, query the secondary index skip list structure in combination with the down pointer to obtain the query result.

[0031] The third aspect of the present invention provides a computer-readable storage medium.

[0032] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the order-preserving ciphertext index method that supports forward security as described above.

[0033] The fourth aspect of the present invention provides a computer device.

[0034] A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the order-preserving ciphertext index method that supports forward security as described above.

[0035] The fifth aspect of the present invention provides a computer device.

[0036] A program product includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the above-described method for supporting forward-secure ordered ciphertext indexing.

[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0038] 1. The present invention is applicable to ordered ciphertext data with forward security, which can distinguish historical data and newly added data, thereby improving the retrieval efficiency of this indexing structure in the forward-secure ordered encryption scheme.

[0039] 2. The present invention designs a new splitting function. By globally optimizing the objective function, it balances the uniformity of subset sizes and the rationality of splitting intervals. It can adjust the splitting strategy according to data density and mutation points, and can flexibly weigh the subset size and splitting interval by adjusting parameters. Moreover, it can suppress the influence of noise and outliers on the splitting result.

[0040] 3. The present invention constructs a new indexing structure, which can perform personalized queries on ordered ciphertext according to the multi-user token status and support the forward-secure indexing function. A token status key chain is added between the hash table and the skiplist structure. In this way, for a hash function using a special structure, ordered ciphertext data within a certain range is aggregated together. The proxy server maps the ciphertext data and the token status in the user query token to the target address on the token status key chain, and then forms a search token and sends it to the cloud server for data retrieval. The cloud server performs data retrieval according to the target address and retrieves on the multi-level index, and finally obtains the result set of the target data, improving the security performance.

[0041] The advantages of the additional aspects of the present invention will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0043] Figure 1 is an index structure based on a B+ tree;

[0044] Figure 2 is a hash index structure based on a skiplist;

[0045] Figure 3 is a flowchart of the method for supporting forward-secure ordered ciphertext indexing provided by an embodiment of the present invention;

[0046] Figure 4 is an improved hash index structure provided by an embodiment of the present invention;

[0047] Figure 5 is the process of a user retrieving ciphertext provided by an embodiment of the present invention;

[0048] Figure 6 are four experimental data sets provided by an embodiment of the present invention, where (a) is the California civil servant salary data set, (b) is the uniform distribution data set, (c) is the normal distribution data set, and (d) is the non-uniform non-continuous distribution data set;

[0049] Figure 7 is the original segmentation of the California salary table provided by an embodiment of the present invention;

[0050] Figure 8 is the segmentation effect of the California salary table segmentation with l = 10 provided by an embodiment of the present invention -9 ;

[0051] Figure 9 is B provided by an embodiment of the present invention n = 4096 for the segmentation effect;

[0052] Figure 10 is B provided by an embodiment of the present invention n = 4096 original segmentation box plot;

[0053] Figure 11 is B provided by an embodiment of the present invention n = 4096, l = 10 -9 ;

[0054] Figure 12 is B provided by an embodiment of the present invention n = 4096, l = 10 -9 box plot;

[0055] Figure 13 is the comparison result of the range query time overhead provided by an embodiment of the present invention. Detailed implementation manners

[0056] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0057] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0058] Note that the terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly dictates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0059] The present invention designs an order-preserving ciphertext index scheme that supports forward secrecy. First, a new splitting function fsplit(data) is designed. It balances the uniformity of subset sizes and the rationality of splitting intervals through a global optimization objective function. It can adjust the splitting strategy according to data density and mutation points and flexibly weigh the subset size and splitting interval through λ, and can suppress the influence of noise and outliers on the splitting results. Secondly, a new index structure is constructed, which can query order-preserving ciphertexts personalized according to the multi-user token status and support the index function of forward secrecy. The main idea of the index structure is to add a token status key chain OTL between the hash table and the skiplist structure. In this way, for a certain range of order-preserving ciphertext data aggregated together using a hash function with a special structure, the proxy server maps it to the target address on the token status key chain OTL according to the ciphertext data and token status in the user query token, and then forms a search token and sends it to the cloud server for data retrieval. The cloud server retrieves data according to the target address and retrieves it on the multi-level index, and finally obtains the result set of the target data.

[0060] Embodiment 1

[0061] As Figure 3 shown, this embodiment provides an order-preserving ciphertext index method that supports forward secrecy, including the following steps:

[0062] Step 1: Obtain a given sorted data set, determine the optimal splitting point in combination with the interval of the data, the subset size, and the constructed data set splitting function, and generate a hash table according to the optimal splitting point;

[0063] The existing method combines the aggregation function LH(x) and the cumulative distribution function f cd (x), and balances the aggregation hash function H oa (x) within the y range through the cumulative distribution function, where H oa (x) = Hash(LH(x), sk) makes each set have the same approximate size, sk represents the user's private key, and the aggregation function LH(x) can be expressed as:

[0064]

[0065] Among them, Count(D) represents the number of different elements in the data set D, and B n represents the expected size of each subset. For an ordered data set D = {1, 2, 5, 15, 16, 17, 18, 19}, if B n = 3, the data set can be divided into {1, 2}, {5, 15, 16}, {17, 18, 19}. Then when the data set is large, the cumulative distribution function f cd (x) will have difficulties in fitting and low efficiency. The implementation essence is based on the equal-frequency uniform segmentation of the sample quantity, that is, the data set is divided into subsets with sizes close to B n through quantiles. This segmentation ensures the balance of the subset sample sizes, but the segmentation interval is completely determined by the data distribution, which may lead to uneven intervals.

[0066] Therefore, the expression of LH(x) is as follows:

[0067]

[0068] Among them x i is the segmentation point in the data set D.

[0069] For the above method of determining the segmentation point only through uniform segmentation (fixed subset size), there are the following problems:

[0070] (1) Ignoring the natural intervals of the data distribution: The data may be dense (small intervals) in some areas and sparse (large intervals) in some areas; forced uniform segmentation will result in too small intervals of segmentation points in dense areas and too large intervals of segmentation points in sparse areas.

[0071] (2) Forcing the division of a fixed number of subsets: The segmentation function divides the data set through the data set, which may result in the last subset being much smaller than the target size (for example, the total amount of data cannot be divided evenly by B n ); ignoring the natural segmentation boundaries in the data (such as data mutation points).

[0072] (3) Unable to adjust the segmentation strategy according to the data characteristics: The segmentation effect of the segmentation function completely depends on B n , when there are noises or outliers in the data, uniform segmentation will cause the outliers to destroy the continuity of the subsets.

[0073] To solve the above problems, this embodiment provides a new segmentation method, which specifically includes:

[0074] In this embodiment, the given sorted data set D sorted , the goal is to find a set of segmentation points S = {x1, x2,... x k}, divide D into k + 1 subsets {S1, S2, … S k+1}, satisfying:

[0075] First, the size of each subset |S i | is as close as possible to the expected size B representing each subset n ;

[0076] Second, give priority to splitting at positions with large data intervals.

[0077] Therefore, in this embodiment, the splitting point is optimized by combining the data interval and the subset size. Specifically, a loss function is defined to measure the difference between the size of each subset and B n , as well as the size of the data interval within the subset, and find the splitting point that minimizes the total loss. Therefore, the specific splitting process includes the following steps:

[0078] Step 101: For the sorted data, select the minimum value min and the maximum value max in the dataset;

[0079] Step 102: Pre-generate a set of splitting points based on the uniform splitting method;

[0080] In this embodiment, k + 2 equally spaced integers (including endpoints) are generated within the interval [min, max] of the minimum and maximum values to generate a pre-splitting point set. After removing the endpoints, the middle k integers are taken as the pre-splitting points;

[0081] Step 103: Search for the splitting point with the lowest loss cost of the splitting function near the splitting point, and adjust the pre-generated set of splitting points to obtain an adjusted set of splitting points;

[0082] In this embodiment, the greedy strategy is used to search for the splitting point with the lowest cost near the splitting point (B n / 2);

[0083] Specifically, the greedy strategy is used to uniformly sample integer candidate points between the left and right boundaries of the current pre-splitting point (the left boundary is the previous splitting point, the right boundary is the next splitting point, and if it does not exist, the data extreme value is taken) to calculate the total cost total cost = sum abs + λ·sum inv , where the block size deviation: |S i | is the size of the i-th subset; the splitting point density penalty The smaller the spacing, the greater the penalty;

[0084] Search for the splitting point with the lowest cost and adjust the splitting point;

[0085] In this embodiment, a new splitting function f split (data) is proposed, which can maximize the data interval at the splitting point while minimizing the difference between the subset size and the expected threshold B n , expressed as:

[0086]

[0087] where arg represents finding the independent variable at which the function attains a certain specific value, represents finding the minimum value of the function within a certain set or range; thus represents the independent variable at which the function attains the minimum value, and this independent variable is the splitting point set S = {x1, x2, … x k}, where x j and x j+1 represent the j-th and (j + 1)-th splitting points;

[0088] The first term |S i | represents the size of the i-th subset after splitting, so this term represents the deviation between the subset size after splitting the dataset and the expected size B n , and in some cases, a slight excess beyond B n is allowed.

[0089] The second term represents the optimization of the interval between splitting points. Minimizing this term is equivalent to maximizing the interval between splitting points. By adjusting the hyperparameter λ, the trade-off between the splitting point interval and the subset size can be controlled.

[0090] f split (data) finds the splitting points through global optimization, that is: allowing the subset size to deviate from B n within a certain range, but the deviation amplitude can be controlled through the deviation term; placing splitting points preferentially near data mutation points (such as numerical jumps, density changes) to improve the splitting rationality; f split (data) provides flexibility through the hyperparameter λ, that is: increasing λ can preferentially ensure the uniformity of the splitting intervals, which is suitable for scenarios with sparse data distributions; decreasing λ can preferentially ensure the uniformity of the subset sizes, which is suitable for scenarios with dense data distributions; in addition, the interval penalty term can naturally inhibit placing splitting points near noise points, improving robustness.

[0091] For example: for the dataset D = {1, 2, 3, 4, 5, 100}, the target subset size B n = 3; by the interval penalty term, splitting points are avoided from being placed between 5 and 100, and subsets {1, 2, 3}, {4, 5}, {100} may be generated, sacrificing the uniformity of the subset sizes, but retaining semantic rationality.

[0092] Step 104: Generate a hash table based on the optimal split point, where the key in the hash table is the split point hash value H oa (x i ), the value in the hash table is pointer st, which points to the first node of the token state key chain;

[0093] Step 2: Build a token state key chain, each node of which stores the hash value of the token, and each node is associated with a secondary index jump table structure;

[0094] The present invention improves the skip table structure and adds a token state key chain OT between the aggregation function mapping to the skip table structure. L , so that forward security can be met when the user sends a data query with a historical token. Specifically, the proxy server first maps the query ciphertext to a certain value through an aggregate hash function, and then calculates its token state key chain OT based on the token state and the aggregate value. L Due to the unidirectional nature of the single linked list, only the data on the chain after the current address can be retrieved. The specific structure is as follows Figure 4 As shown, the specific steps include:

[0095] Step 201: The token state key chain is a single linked list structure, and each node stores the hash value H(k s ,OT i ), i represents the token state, when new data is inserted, i+1; each node also has a pointer associated with the first address of the skip list, and each node in the token chain is associated with a skip list structure;

[0096] Step 202: Initialize the token state key chain. Each token chain node is H(k s ,OT i ), where i is 1, because it is the initialization token, the first token node, the node successor pointer points to the NIL value, and there is also a down pointer pointing to the first address of the first level index of the jump table;

[0097] Step 203: In the skip list structure of the secondary index, the node value in the skip list is the ciphertext. The ciphertext form is related to the token state. j , whose ciphertext form is associated with the token state;

[0098] That is: the sub-dataset element x after segmentation j First perform order-preserving encryption to obtain MCORE.Enc(x j ), and then XOR it with the hash value of the token node to which it belongs, and get The ordered form links all elements into an underlying linked list, generates a secondary index by taking points at intervals from the underlying linked list, and generates a primary index by taking points at intervals from the secondary index.

[0099] For example, for all ciphertexts with a token status of 1, a skip list structure with a secondary index is generated. The head address of this skip list structure is associated with the L1 node in the token chain. The underlying is a single linked list structure of ciphertexts. The specific process of generating the skip list is as follows: For each sub-dataset, first store the plaintexts in the sub-dataset in a single linked list structure of ciphertexts, which is an ordered increasing original single linked list. Then, based on the original linked list, take every other node as a node of the secondary index (that is, take values at intervals from the original linked list), so that the secondary index is obtained. On the basis of the secondary index structure, using the same method, take points at intervals to generate the primary index.

[0100] Step 3: Calculate the split sub-dataset to which the ciphertext belongs according to the user's query ciphertext and token, and obtain the split point hash value according to the split point to which the sub-dataset belongs; obtain the head node address of the token status key chain according to the split point hash value, and calculate the token hash value according to the status of the token;

[0101] As Figure 5 shown in the process of the user retrieving the ciphertext, due to the addition of the index structure and the aggregate hash function, the token status information table W needs to be reconstructed in the proxy server. As Figure 4 known, the aggregate hash value H oa (x) and the token status OT are in a "one-to-many" mode, that is: in the subset S i within a split point, the key-value pair form in the hash table is <H oa (x i ), ptr>, where H oa (x i ) is the hash value of the split point, and ptr is the pointer to the token key chain. Figure 4 The corresponding token status information table W is shown in Table 1, where

[0102] Table 1 Token Status Information Table W

[0103]

[0104] After receiving the query token OT and ciphertext t of the data user user u, the proxy server verifies the identity legality of user u in the query complementary key table U q ; for a user with legal identity, recalculate its ciphertext t to obtain the complete query ciphertext c target .

[0105] In this embodiment, the sub - dataset to which the ciphertext belongs is calculated through an aggregation function, which specifically includes:

[0106] Analyze the original aggregation function LH(x), whose main function is to judge the size relationship between the value x and many split points x i and only count the number of values greater than the split point x i . Finally, the value of the LH(x) function is obtained to locate the set of sub - datasets to which it belongs. Its time complexity is O(B c ); Considering that the comparison operation of order - preserving ciphertext is a very time - consuming operation, a better performance requirement should be imposed on the aggregation function during data retrieval. Therefore, this embodiment proposes an efficient aggregation function LH new (x). The mathematical definition of the LH(x) function is as follows:

[0107] Given the input value x, the function LH(x) returns the smallest index i that satisfies x i >x, that is:

[0108] LH(x)=min{i|x i >x},

[0109] If all x i ≤x, then return B c - 1.

[0110] The formal formula of the efficient aggregation function LH new (x) in this embodiment:

[0111] Assume that the split - point sequence is in ascending order Define the recursive function BinarySearch(x, L, R):

[0112]

[0113] When L is the left boundary of the binary search algorithm, initially 1, and R is the right boundary of the binary search algorithm, initially B c - 1, the expression of the efficient aggregation function LH new (x) is obtained as follows:

[0114]

[0115] Among them, x is the query ciphertext, x i is the split - point ciphertext, and B c is the size of the set of sub - datasets after dataset splitting.

[0116] In this way, LH newThe time complexity of (x) is O(log|S|), which is a huge performance improvement compared to the time complexity of O(|S|) for the original aggregation function LH(x) to locate the aggregation position of x. In this way, the time overhead of the aggregation function changes from O(|S|) to O(log|S|), where |S| is the size of the set of split points.

[0117] Step 4: Traverse and retrieve according to the head node address of the token status key chain to find the node corresponding to the token hash value, and obtain the down pointer of this node; for the query ciphertext, query the secondary index skip list structure in combination with the down pointer to obtain the query result.

[0118] Specifically, it includes:

[0119] After the user locates the sub-dataset of the split to which x belongs through the LH new (x) function, according to the split point x to which it belongs i , calculate its key value H oa (x i ) in the hash table, and then according to the status H(k s , OT i ) of the token, find which node in the token chain it belongs to according to the token status i. After finding the node, this node is associated with the head address of the first-level index of the skip list structure.

[0120] After locating the head address of the specific first-level index through the token chain, start data retrieval. For example, to find all values greater than the ciphertext x, find the last node less than the ciphertext x in the first-level index, then find the secondary index through the downward pointer, and then find the last node less than the ciphertext x in the secondary index. Through the downward pointer, come to the bottom linked list. Continuing from the previous node, in the bottom linked list, find the first node target that satisfies being greater than the ciphertext x, and then return all the nodes after target (all the nodes after target satisfy being greater than x).

[0121] If in the token chain in the second step, assuming that the traversed one is L2 in the token chain and its successor node is L1, after completing the skip list traversal associated with L2 above, a similar traversal process for the skip list associated with L1 is also required, and finally all the data is returned.

[0122] Step 5: When the data is dynamically updated, perform corresponding operations according to different types;

[0123] When inserting data, for newly added data, which is similar to data initialization, the data owner needs to divide the newly added data set according to the existing index split points, then calculate and generate ciphertext and the corresponding skip list index structure, forward the ciphertext and the index structure to the cloud server, and update the token status information table W of the proxy cloud server. Considering that the subsequent inserted data may not conform to the split point scheme calculated from the initial data set, resulting in the size of a split point set exceeding the set block size threshold, denoted as B th . To solve this problem, an index reconstruction mechanism is defined.

[0124] Index Reconstruction. Assume that the split function generates B c block subsets, and the value range of the data set is A n . Therefore, define Since the value range A n > Count(D), then there must be B th > Bn. The data owner will maintain a table locally to record the number of nodes under each block. The key-value pair is (LH(x i ), size), and this table has a total of B c entries. Before inserting new data, the data owner will retrieve whether the size of the target block exceeds the threshold B th . For nodes that exceed the threshold, split the data of the node. The splitting method is to split the skip list data under the current split point in an evenly divided manner, and use the average value of the maximum and minimum values of the data set under this split point as the new split point. In this way, through the secondary index structure, the underlying linked list can be quickly split into two linked lists with uniform sizes. Therefore, the data distribution of the split sub-data sets has a positive effect on the effect of index reconstruction; specifically, if the data is evenly distributed, the split points after splitting can divide the new sub-data sets more evenly; on the contrary, for unevenly distributed data, the split points after splitting have an unsatisfactory splitting effect on the new sub-data sets, and subsequent inserted data will be more likely to induce index reconstruction, increasing the time overhead.

[0125] Deleting Data. For deleting data, the data owner only needs to send its ciphertext to the cloud server. If the value exists, logically delete the data.

[0126] Updating Data. After logically deleting the original data, insert the new data. In addition to updating the token status, the index also needs to be updated.

[0127] Experimental Results and Comparative Analysis

[0128] To evaluate the performance of index construction, data insertion, range query in the proposed forward-secure ordered ciphertext index scheme of the present invention, as well as the size of the storage space occupied by the data structures involved in the scheme, the FSOREL scheme in this chapter is mainly compared with the existing SOREL scheme in the following aspects, including the performance of index construction, insertion, and range query.

[0129] (1) Effect and efficiency of index construction. First, the present invention conducts experiments on the effect of index construction. For the segmented data sets, four data sets are used for experiments, as shown in Table 2, including a real data set and three random data sets. The first real data set uses the Total Pay & Benefits column of California civil servant salaries, hereinafter referred to as the salary data set. The data set is converted into positive integers, and the value range An = {0, 1863526}. The second data set is a uniform data set, with the value range An = {0, 1000000}; the third data set is a normally distributed data set, with the value range An = {0, 4000000}; the fourth is a non-uniform and non-continuous data set, with the value range An = {0, 4000000}. Among them, the column data

[0130] Table 2 Data Sets

[0131]

[0132] When segmenting the data set, calculating the standard deviation of the subset sizes and the variance of the segmentation intervals can help evaluate the quality and rationality of the segmentation results. Because the meanings of these two indicators, the standard deviation of the subset sizes and the variance of the segmentation intervals, are elaborated as follows:

[0133] (1) Standard deviation of subset sizes. The standard deviation of subset sizes measures the degree of dispersion of the sizes of each subset after segmentation; specifically: if the standard deviation value is small, it means that the segmentation algorithm evenly divides the data set into subsets with similar sizes, and the segmentation result is relatively balanced; if the standard deviation is large, it means that there are significant differences in the sizes of each subset, and there are some subsets that are too large or too small, and the segmentation result is unbalanced and unsatisfactory.

[0134] (2) Variance of segmentation intervals. The variance of segmentation intervals measures the degree of dispersion of the intervals between segmentation points; specifically: if the variance of the segmentation intervals is relatively small, it means that the segmentation points are more evenly distributed in the data set; if the variance is large, it means that the intervals between some segmentation points may be too large or too small, and the segmentation result is unsatisfactory.

[0135] Therefore, the present invention combines the two to comprehensively evaluate the quality of the segmentation results, that is: an ideal segmentation result is that the standard deviation of the subset sizes is small, and at the same time, the variance of the segmentation intervals is also small, which means that the segmentation result meets the requirements of balance and average.

[0136] The present invention conducts experiments on the Total Pay & Benefits column in the salary dataset of California public servants. The data distribution of the salary dataset is as follows Figure 6 shown in (a) below. It can be noted that its data density is 13%, and its data distribution is very uneven, mainly concentrated in the interval [200000, 250000]. The segmentation effect of the original scheme is as follows Figure 7 shown. The present invention controls the data segmentation situation when LH new , l = 10 -9 The segmentation effect is as follows Figure 8 shown. Note that the size of the segmented subsets of the latter is not a fixed size, but fluctuates around the expected subset size B n . When the expected subset size B n = 2048, calculate the standard deviation of the subset size and the variance of the segmentation interval (hereinafter simply referred to as the standard deviation and variance). The standard deviation and variance of the original segmentation scheme dataset are 166 and 165666403 respectively, while the standard deviation and variance of the present invention are 62 and 42433426 respectively. It can be seen that the segmentation effect of the present invention is superior to the original scheme both in terms of the dispersion degree of each subset size and the dispersion degree of the intervals between the segmentation points.

[0137] For the uniformly distributed dataset, its data distribution is as follows Figure 6 shown in (b) below. It is considered that using uniform segmentation (fixed subset size) to determine the segmentation is the best scheme. The LH new function of the present invention can achieve uniform segmentation by setting l = 0; for the normally distributed dataset, its data distribution is as follows Figure 6 shown in (c) below. It is also considered that using the uniform segmentation method is the best scheme. Compared with the uniformly distributed dataset, it is necessary to experiment with the value range of the expected subset size B n . The same segmentation effect can also be achieved by setting l = 0. Therefore, these two datasets will not be elaborated redundantly.

[0138] For the non-uniform and non-continuous dataset, its data distribution is as follows Figure 6 shown in (d) below. First, use the original segmentation function B n = 4096 for segmentation. The size of the segmented subsets is as follows Figure 9 shown. Calculate the standard deviation and variance of the original scheme, which are 112 and 174614499 respectively. Use a box plot to observe the data distribution in the segmented subsets. The subset data distribution is as follows Figure 10 shown. Since the dataset is non-uniform and non-continuous, it can be observed that Figure 10 a large number of outliers appear in the middle part. The outlier points indicate the existence of extreme values or outliers in the sub-datasets;

[0139] In the solution of the present invention, different values of l are taken for segmentation. For l = 10 -9 The segmentation effect is as Figure 11 and Figure 12 shown. Compared with the box plot with a large number of outliers, the subset data segmented by l = 10 -9 is more uniform. The standard deviation and variance of this segmentation scheme are calculated to be 80 and 167132643 respectively. It can be seen that the segmentation effect of the present invention is better than the original scheme in both the dispersion degree of the sizes of each subset and the dispersion degree of the intervals between the segmentation points.

[0140] In addition to being affected by the size of the data set, the time for constructing the index is also affected by the expected size B n of the sub - data set. In each data set, the expected size B n of the sub - data set ranges from 256 to 1024 (256, 512, 1024). The results show that the larger B n is, the longer the construction time required.

[0141] Since the segmentation scheme of the present invention selects the segmentation points based on global optimization, the time for constructing the index is affected by the distribution of the data set.

[0142] (2) The efficiency of data insertion. Since a different index reconstruction scheme from SOREL is adopted, the data insertion efficiency needs to be analyzed in two cases: when not exceeding the threshold and when exceeding the threshold. In the SOREL scheme, for inserting into the data set S, it is necessary to evaluate whether the existing index segmentation points will cause the size of the subset data set to exceed the threshold B th , if it does not exceed the threshold, the insertion time overhead is O(B c ), Count(D) is the number of different elements in the data set D, and B n is the expected size of the subset; if there exists a data set S i of a certain index node, and the size of the data set S i exceeds the threshold B th , then all index nodes need to be reconstructed, and its time complexity is O(|DlogD|), where |D| is the size of the data set.

[0143] In the FSOREL scheme of the present invention, similarly, for inserting into the data set S, it is necessary to evaluate whether the existing index segmentation points will cause the size of the subset data set to exceed the threshold B th , if it does not exceed the threshold, the time complexity is O(logB c ); if there exists a data set S i of a certain index node, and the size of the data set S i exceeds the threshold B th, it is necessary to split the splitting point. At this time, it is necessary to split the token chain and skiplist data under the splitting point to be split. Since there is a secondary index, only the first-level index needs to be traversed when evenly dividing the sub-datasets, and its time complexity is O(|B n / 4|). Therefore, the overall time complexity is O(logB c ) + O(|OT| * |B n |).

[0144] (3) Efficiency of range query.

[0145] The time complexity of data query is similar to that of data insertion, which is O(logB c ) + O(logB n ), where B c is the number of splitting points and B n is the expected size of the sub-dataset. In addition, it is also related to the data density in the dataset. The so-called data density means that in a set with a fixed domain size, the more elements it contains, the higher the data density, that is, the smaller the difference between adjacent elements in the dataset.

[0146] The time complexity of range query is similar to that of data insertion. It is necessary to query the boundary value x of the range query and find the first skiplist node that meets the boundary value x. Therefore, the time complexity is O(logB c ) + O(logB n ). First, use binary search in the token status information table W to find the split sub-dataset to which the boundary value x belongs, and the time complexity is O(logB c ); subsequently, perform data retrieval in the skiplist structure, and the time complexity is O(logB n ), where B n is the expected size of the sub-dataset.

[0147] Data screening was performed on the Total Pay&Benefits column in the California public employee salary dataset. Experiments were conducted using the simplest query condition (8000 < salary < 15000). There were a total of 236,178 data entries in this dataset, and the target dataset had 2,616 entries. Experiments were carried out using 4 different bit numbers: 24 bits, 32 bits, 48 bits, and 64 bits. If the index structure is not used for querying, all the data in the dataset needs to be traversed and compared, and the query time overhead is the time for comparing the ordered ciphertext of 236,178 data entries, which will be a very time-consuming operation. After splitting the dataset using the splitting function, the splitting points related to 8000 and 15000 are x1 = 4697, x2 = 11306, and x3 = 18035. Search tokens are generated by binary searching the token status information table W. The search tokens greatly reduce the traversed dataset. Therefore, data retrieval only needs to be performed within the sets of splitting points x2 and x3, and the query time overhead is greatly reduced through the secondary index. The experimental results are as follows Figure 13 As shown. The experimental results intuitively show that the time overhead of range queries drops sharply when there is an index, greatly improving the efficiency of range queries.

[0148] Example Two

[0149] This example provides an ordered ciphertext index system that supports forward security, including:

[0150] A data splitting module, which is used to obtain a given sorted dataset, determine the optimal splitting points by combining the data interval, subset size, and the constructed dataset splitting function, generate a hash table according to the optimal splitting points. The keys in the hash table are the hash values of the splitting points, and the values in the hash table are pointers pointing to the first node of the token status key chain;

[0151] A data mapping module, which is used to construct a token status key chain. Each node stores the hash value of the token, and at the same time, each node is associated with a secondary index skip list structure;

[0152] A data retrieval module, which is used to calculate the split sub-dataset to which the ciphertext belongs according to the user's query ciphertext and the token, and obtain the hash value of the splitting point according to the splitting point to which the sub-dataset belongs; obtain the address of the first node of the token status key chain according to the hash value of the splitting point, and calculate the token hash value according to the status of the token; traverse and retrieve according to the address of the first node of the token status key chain, find the node corresponding to the token hash value, and obtain the down pointer of this node; for the query ciphertext, query the secondary index skip list structure in combination with the down pointer to obtain the query result.

[0153] It should be noted that the specific implementation of the forward-security-supported ordered ciphertext indexing system in the embodiments of the present invention is similar to that of the forward-security-supported ordered ciphertext indexing method in the embodiments of the present invention. For specific details, please refer to the description in the method section. To avoid redundancy, it will not be elaborated here.

[0154] Embodiment 3

[0155] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the forward-security-supported ordered ciphertext indexing method as described above.

[0156] Embodiment 4

[0157] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the forward-security-supported ordered ciphertext indexing method as described above.

[0158] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0159] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0160] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 or steps of the functions specified in a plurality of blocks.

[0162] Those of ordinary skill in the art can understand that all or part of the processes of the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above various methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0163] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An ordered ciphertext index method that supports forward security, characterized in that, The steps are as follows: Obtain a dataset with a given sorting order, determine the optimal splitting point by combining the interval of the data, the subset size, and the constructed dataset splitting function, generate a hash table according to the optimal splitting point, where the key in the hash table is the splitting point hash value, and the value in the hash table is a pointer pointing to the head node of the token status key chain; Construct a token status key chain, where each node stores the hash value of the token, and at the same time, each node is associated with a secondary index skip list structure; Calculate the split sub-dataset to which the ciphertext belongs according to the user's query ciphertext and the token, and obtain the splitting point hash value according to the splitting point to which the sub-dataset belongs; Obtain the head node address of the token status key chain according to the splitting point hash value, and calculate the token hash value according to the status of the token; Traverse and retrieve according to the head node address of the token status key chain, find the node corresponding to the token hash value, and obtain the down pointer of the node; For the query ciphertext, query the secondary index skip list structure in combination with the down pointer to obtain the query result.

2. The method for supporting forward-secure and order-preserving ciphertext indexing according to claim 1, wherein The step of obtaining a dataset with a given sorting order, determining the optimal splitting point by combining the interval of the data, the subset size, and the constructed dataset splitting function includes: For the sorted data, select the minimum value and the maximum value in the dataset; Generate k + 2 equally spaced integers within the interval of the minimum value and the maximum value to generate a set of pre-splitting points. After removing the endpoints, take the middle k integers as the pre-splitting points; Use the greedy strategy to uniformly sample integer candidate points between the left and right boundaries of the current pre-splitting point, and calculate the total cost of the dataset splitting function. When the total cost is minimized, the corresponding splitting point is the optimal splitting point. The total cost of the dataset splitting function is calculated for the candidate points. When the total cost is minimized, the corresponding splitting point is the optimal splitting point.

3. The method for supporting forward-secure and order-preserving ciphertext indexing according to claim 1, wherein The dataset splitting function is: Among them, the first item |S i | represents the size of the i-th subset after splitting. The second item represents the optimization of the interval between splitting points. By adjusting the hyperparameter λ, the trade-off between the splitting point interval and the subset size is controlled. arg represents finding the independent variable at which the function obtains a specific value represents finding the minimum value of the function within a certain set or range. Therefore represents the independent variable at which the function obtains the minimum value. The independent variable is the set of splitting points S = {x1, x2, … x k}, and S i represents the i-th subset. Among them, x j and x j+1 represent the j-th and (j + 1)-th splitting points. B n represents the expected size of each subset.

4. The method for supporting forward-secure ordered ciphertext indexing according to claim 1, wherein In the secondary index skip list structure, the node value in the skip list is the ciphertext. First, perform order-preserving encryption on the elements of the split sub-dataset, and then perform exclusive OR according to the hash value of the token node to which it belongs. Link all elements in an ordered form into a bottom-level linked list, generate a secondary index in the form of interval sampling for the bottom-level linked list, and generate a primary index in the form of interval sampling for the secondary index.

5. The method for supporting forward-secure ordered ciphertext indexing according to claim 4, wherein The step of querying the secondary index skip list structure in combination with the down pointer for the query ciphertext to obtain the query result includes: Find all values greater than the query ciphertext. Find the last node in the primary index that is less than or equal to the ciphertext x. Through the down pointer, find the secondary index, and then find the last node in the secondary index that is less than the ciphertext x. Through the downward pointer, reach the bottom-level linked list. In the bottom-level linked list, find the first node target that satisfies being greater than the ciphertext x, and then return the data corresponding to all nodes after the target.

6. The method for supporting forward-secure ordered ciphertext indexing according to claim 1, characterized in that, Calculate the split sub-dataset to which the ciphertext belongs through an aggregation function, and the aggregation function is: Among them, x is the query ciphertext, and x i is the split-point ciphertext, and B c is the size of the set of sub-datasets after dataset splitting, and BinarySearch(x, 1, B c - 1) is a recursive function.

7. An ordered ciphertext index system supporting forward security, characterized in that It includes: A data splitting module, which is used to obtain a dataset with a given sorting order, determine the optimal splitting point by combining the interval of the data, the subset size, and the constructed dataset splitting function, generate a hash table according to the optimal splitting point, where the key in the hash table is the splitting point hash value, and the value in the hash table is a pointer pointing to the head node of the token status key chain; A data mapping module, which is used to construct a token status key chain, where each node stores the hash value of the token, and at the same time, each node is associated with a secondary index skip list structure; A data retrieval module, which is used to calculate the split sub-dataset to which the ciphertext belongs according to the user's query ciphertext and the token, and obtain the splitting point hash value according to the splitting point to which the sub-dataset belongs; Obtain the head node address of the token status key chain according to the split point hash value, and calculate the token hash value according to the status of the token; Traverse and retrieve according to the head node address of the token status key chain, find the node corresponding to the token hash value, and obtain the down pointer of the node; For the query ciphertext, query the secondary index skip list structure in combination with the down pointer to obtain the query result.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the forward-security supported ordered ciphertext indexing method described in any one of claims 1-6.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the forward-security supported ordered ciphertext indexing method described in any one of claims 1-6.

10. A program product, which is a computer program product and includes a computer program, characterized in that When the computer program is executed by the processor, it implements the steps in the forward-security supported ordered ciphertext indexing method described in any one of claims 1-6.