Distributed Concealment Method and Readable Storage Medium under Trusted Execution Environment
By providing a distributed occult Join algorithm in a trusted execution environment, using primary key and equivalent connection optimization algorithms, the use of existing algorithms is solved and security risks is achieved, and efficient and secure database connection operations are achieved.
Patent Information
- Application Number
- CN202510472807.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing distributed occult Join algorithm has limited use and data security risks. It only supports primary key connections and requires publication of the frequency information of the key, which affects its universality and data privacy.
A distributed occult method is provided in a trusted execution environment. By judging whether the connection keys in table S meet the primary key constraints, a primary key connection algorithm or an equivalent connection optimization algorithm is used. The equivalent connection optimization algorithm includes sorting, maximum prefix and calculation, connection degree calculation, obscure extension and data rearrangement to achieve obscure join operations.
It improves the computing power and performance of database connection operations in TEE environment, while enhancing data privacy protection, solves the limitations of existing algorithms and security risks, and has a wide range of application prospects.
Smart Images

Figure CN120030530B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of database information security, and particularly to a distributed obfuscation method and a readable storage medium in a trusted execution environment. Background Art
[0002] In the current environment, providing encrypted data computing for users has become one of the important services of cloud service providers. The core value of encryption lies in ensuring the confidentiality and privacy of data, especially sensitive data such as personal privacy information, financial documents, and business secrets. However, computing on encrypted data usually incurs a large overhead, and the computing speed may decrease by several orders of magnitude.
[0003] To solve this problem, a common method is to perform computing based on a trusted execution environment (TEE), that is, decrypting and computing after the data enters the trusted area and re-encrypting it before leaving the area. However, existing TEE solutions still have security risks. The host system can infer auxiliary information of encrypted data by monitoring the memory access patterns of application programs. Therefore, the computing within the security area should be made as obfuscated as possible.
[0004] Currently, the solutions to this problem are mainly divided into two ways: software and hardware.
[0005] 1) Hardware solution: Usually, the Oblivious RAM (ORAM) technology is adopted, which can effectively prevent the leakage of data access patterns, but it will bring a high performance overhead, and the computing time will increase by a factor of O(log²N).
[0006] 2) Software solution: Design obfuscation algorithms for specific computing operators (such as the Join operator). Currently, there are two typical distributed obfuscation Join algorithms - Opaque and SODA: Opaque is an obfuscated sorting algorithm based on column sorting, which can implement the Join operation, but its application scenario is limited and it is only applicable to primary key joins; SODA believes that the cost of column sorting is too high, supports conventional equality Joins, but requires the frequency information of the most popular keys in two tables to be publicly disclosed, thus having certain defects in security.
[0007] Due to the above methods still presenting significant challenges in terms of usage constraints and security, the current distributed obfuscation Join algorithms face the following problems:
[0008] 1) Usage limitation: Only supporting primary key joins affects its generality;
[0009] 2) Security risk: Requiring the public disclosure of frequencies of similar keys or other auxiliary information affects data privacy. Summary of the Invention
[0010] The purpose of the present invention is to provide a distributed anonymity method and a readable storage medium in a trusted execution environment, so as to solve the problems such as limited use and potential data security risks existing in the existing distributed anonymity Join algorithm.
[0011] To achieve the above object, the present invention provides a distributed anonymity method in a trusted execution environment for a Join operator in a database. The Join operator takes table R and table S as inputs and outputs all tuple combinations that meet the join condition through the join key B. Assume R = R(A, B) and S = S(B, C); the distributed anonymity method includes:
[0012] S1. According to the primary key join definition, determine whether the join key B in the table S meets the primary key constraint; if it meets, adopt the primary key join algorithm; otherwise, adopt the equivalent join optimization algorithm;
[0013] When adopting the equivalent join optimization algorithm, the following steps are executed:
[0014] S21. Sort the table R and the table S according to the join key B;
[0015] S22. Calculate the maximum prefix sum of the table R and the table S on the join key B respectively to obtain table R1 and table S1;
[0016] S23. Calculate the join degree of the table R1 using the table S1 to obtain table R2, and calculate the join degree of the table S1 using the table R1 to obtain table S2;
[0017] S24. Perform anonymity expansion on the table R2 and the table S2 respectively to obtain table R3 and table S3;
[0018] S25. Rearrange the table S3 to obtain table S4;
[0019] S26. Horizontally merge the data of the table R3 and the table S4 in each server, and finally extract columns A, B, and C of the merged result as the output result.
[0020] Optionally, when adopting the primary key join algorithm, the following steps are executed:
[0021] S31. Perform key value assignment on the tuples in the table S;
[0022] S32. Sort the join key B in the partitions of the table R on each server, and select the first tuple with the same value in the join key B as the representative tuple, mark other tuples as inactive tuples, and record the corresponding original server ID at the same time;
[0023] S33. Assign key values to the representative tuples according to the join key B, and randomly assign the inactive tuples to other servers;
[0024] S34. On each server, perform a join calculation between the representative tuples and the local tuples in the table S;
[0025] S35. Transmit the data after the join calculation to the original server, and distribute the results of the representative tuples to the corresponding inactive tuples to form the final result.
[0026] Optionally, in S31, the key value assignment includes:
[0027] S311. Define the tuple in the form of a key-value pair, and each tuple is denoted as x = (k, v), where k is the key and v is the value;
[0028] S312. Calculate the key value assignment operation t = h(k) through a random oracle h to determine the assignment target of the tuple;
[0029] S313. According to the calculation result of h(k), gather the tuples with the same t to the same server;
[0030] S314. Use public parameters for adjustment and padding during the key value assignment process.
[0031] Optionally, in S22, calculating the maximum prefix sum of the table R includes:
[0032] S221. For the table R, calculate the prefix sum of each element on the join key B;
[0033] S222. For the table R, add a column DR, and refresh all the elements in the corresponding column DR with the same prefix with the maximum prefix sum of the same prefix to obtain the table R1;
[0034] Calculating the maximum prefix sum of the table S includes:
[0035] S223. For the table S, calculate the prefix sum of each element on the join key B;
[0036] S224. For the table S, add a column DS, and refresh all the elements in the corresponding column DS with the same prefix with the maximum prefix sum of the same prefix to obtain the table S1.
[0037] Optionally, the algorithm for the prefix sum is as follows:
[0038] E1. The prefix sum represents the cumulative number of times a certain prefix appears in the tuple;
[0039] E2. At the beginning of the calculation, the cumulative count is set to 0;
[0040] E3. Scan the tuples in sequence. When encountering the same prefix, increment the cumulative count by 1;
[0041] E4. When encountering a new prefix, the cumulative count is recalculated starting from 1.
[0042] Optionally, in S23, calculating the join degree of the table R1 using the table S1 includes:
[0043] S231. Delete column C of the table S1 and deduplicate the data to obtain a table S with the join key B as the primary key column B ;
[0044] S232. Based on the table R1 and the table S B , use the primary key join algorithm to obtain table R L ;
[0045] S233. Delete the data in table R L where the DS degree value is 0, and delete column DR to obtain the table R2;
[0046] Calculating the join degree of the table S1 using the table R1 includes:
[0047] S234. Delete column A of the table R1 and deduplicate the data to obtain a table R with the join key B as the primary key column B ;
[0048] S235. Based on the table S1 and the table R B , use the primary key join algorithm to obtain table S L ;
[0049] S236. Delete the data in S L where the DR degree value is 0, and delete column DS to obtain the table S2.
[0050] Optionally, in S24, performing stealthy expansion on the table R2 includes:
[0051] Expand columns A and B according to the DS degree value in the table R2, and then remove column DS to obtain table R3;
[0052] Performing stealthy expansion on the table S2 includes:
[0053] Expand columns B and C according to the DS degree value in the table S2, and then remove column DR to obtain table S3.
[0054] Optionally, in S25, rearranging the table S3 includes:
[0055] S251. Add a column I to the S3. The value of the column I is the prefix sum of the connection key B, and use a column J to store the global minimum rank of the same prefix.
[0056] S252. Calculate the global sorted column L according to the column I and the column J. Perform a modulo operation on the global sorted column L to obtain the server serial number column T. Then delete the column I and the column J, and perform stealth filling to obtain S T ;
[0057] S253. Allocate the data of the table S T to the corresponding servers according to the server serial number column T, and sort the data in each server according to the server serial number column L;
[0058] S254. Delete the L column and the T column of the table S T to obtain the table S4.
[0059] Based on the same inventive concept, the present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed, it can implement the distributed stealth method in the trusted execution environment as described above.
[0060] In the distributed stealth method and the readable storage medium provided by the present invention in the trusted execution environment, at least one of the following beneficial effects is achieved:
[0061] 1) Enhanced security: The present invention introduces the definitions of communication stealth and computing stealth to ensure that in a distributed computing and secure execution environment (TEE), an attacker cannot infer information about the input data from the data exchange or memory access patterns. At the same time, a stealth extension algorithm is adopted to reduce the risk of data leakage and enhance the data privacy protection ability by filling and randomizing the data distribution;
[0062] 2) More efficient primary key join algorithm: Existing join algorithms are usually not specifically optimized for primary key joins. The present invention makes use of the primary key constraint to reduce unnecessary repeated calculations during join calculations;
[0063] 3) Implementation and performance improvement of equivalent join: When dealing with non-primary key joins, the present invention proposes a new equivalent join algorithm. By means of prefix sum calculation, join degree calculation, and stealth extension, etc., the table is filled and extended to meet the stealth requirements, and at the same time, the calculation of general equivalent joins is realized, and there is a certain degree of performance improvement;
[0064] 4) The present invention not only improves the computing power and performance of database connection operations in the TEE environment, but also ensures data privacy through security policies, and has broad application prospects. It can be applied to multiple fields such as cloud databases, data security computing, and big data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Those of ordinary skill in the art should understand that the provided drawings are used to better understand the present invention and do not constitute any limitation to the scope of the present invention. Among them:
[0066] Figure 1 It is a flowchart of a distributed concealment method in a trusted execution environment provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] Before elaborating on the present invention in detail, the basic concepts related to distributed concealment are described here first.
[0068] 1. Distributed algorithm
[0069] In distributed computing, multiple servers work together and execute computing tasks through specific algorithms, which are usually called distributed algorithms. Each server holds a part of the input data, and these data elements (for example, there are a total of N) are almost evenly distributed among the servers. Specifically, the i-th server has a subset of n = Θ(N / p) elements as its initial local data set, where N is the total data volume, p is the number of servers, and {n} represents the data volume of each server, and these are public to all servers. In the process of multiple rounds of computing, in each round, the i-th server processes its local data set X and generates a set of outputs CY = {Y1, Y2,..., Yp}, where i ∈ [p], and this stage marks the completion of the current round of computing. Then, it enters the communication stage, and all servers exchange data in the network. Specifically, the i-th server transmits its output Y to the j-th server. After receiving data from all other servers, the server merges these data and updates its local data set. Then, the updated data set will be used for the next round of computing, or as the input for subsequent tasks, or as the final output to be sent out.
[0070] To distinguish from distributed algorithms, the term "single-machine algorithm" is generally used to refer to the traditional computing setting with only one server. Usually, the single-machine algorithm only involves the local computing of the server and does not involve network communication.
[0071] 2. Encrypted state system
[0072] In a confidential computing system, data on the server is uploaded in encrypted form. When the client submits a query request, the system converts the query into a series of database operations and executes these operations according to the underlying algorithm. Finally, the operation result, i.e., the output of the query, will be returned to the client in encrypted form.
[0073] In such a system, each server is equipped with a protected trusted execution environment (TEE) to ensure that the memory of the server can hold the data elements it processes during computation. All data elements are encrypted in storage outside the TEE, i.e., the data is stored in plaintext inside the TEE and in ciphertext outside.
[0074] 3. Security Definitions
[0075] 3.1 Communication Concealment:
[0076] In a distributed environment, for any deterministic or randomized algorithm A and any input X, let S(X) = {S ij} denote a sequence of messages, where {S ij} represents the size of the message sent from the i-th server to the j-th server. Communication concealment is defined as follows:
[0077] If algorithm A has communication concealment, then there exists a probabilistic simulator Y that can generate a simulated sequence S1 = Y(n1, n2, …, np) such that no polynomial-time algorithm can distinguish the real S(X) from the simulated S1 with a success probability greater than 1 / 2. In other words, the record of any input depends only on the input size on the server, ensuring that an attacker cannot infer any information about the input from the record.
[0078] 3.2 Computational Concealment:
[0079] For any single-machine algorithm A, when it is executed on a certain server, the memory access operations of the input X can be represented as a sequence ((op1, a1, x1), …, (opk, ak, xk)), where each op represents a read or write operation. Specifically, a read operation means reading an element from memory location a, and a write operation means updating the element x to memory location a. For the input X, the memory access pattern of algorithm A is defined as A(X) = (a1, …, ak), which is all the information that a memory access pattern adversary (such as a TEE side-channel attacker) can observe. This paper defines computational concealment as follows:
[0080] If algorithm A has computational stealth, then there exists a probabilistic simulator Sim that can generate a simulated computation sequence A1 = Sim(n) such that no polynomial-time algorithm can distinguish A(X) from A1 with a success probability exceeding 1 / 2.
[0081] 4. Padding the output
[0082] The above security definition assumes that the input size N is public, but the output size M is not, because the output depends on the data content. To ensure stealth, the stealth algorithm always pads the output to the maximum possible output size (i.e., the worst-case output size when the input size is N). However, in the join operation, the worst-case output size is the product of the sizes of the two input tables. The join operation causes the computational cost to grow exponentially, leading to performance degradation or the system becoming unusable. Therefore, in practical implementations, a certain degree of sacrifice in security is often chosen in exchange for better performance. That is to say, the output size is not always padded to the worst-case output size, but rather a reasonable balance is found between security and privacy.
[0083] To make the objectives, advantages, and features of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the accompanying drawings are in a very simplified form and use non-precise scales, only for the purpose of facilitating and clearly assisting in explaining the embodiments of the present invention. To make the objectives, features, and advantages of the present invention more obvious and understandable, please refer to the accompanying drawings. It should be noted that the structures, scales, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the limiting conditions of the implementation of the present invention. Any modification of the structure, change in the proportional relationship, or adjustment of the size, in the case of being the same or similar to the effects that the present invention can produce and the objectives that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0084] As used in the present invention, the singular forms "a", "an", and "the" include plural objects unless the context clearly indicates otherwise. As used in the present invention, the term "or" is generally used in the sense of including "and / or" unless the context clearly indicates otherwise. As used in the present invention,
[0085] Please refer to Figure 1, this embodiment provides a distributed concealment method in a trusted execution environment for the Join operator in a database. The Join operator takes tables R and S as inputs and outputs all tuple combinations that meet the join condition through the join key B. Assume R = R(A, B) and S = S(B, C); the distributed concealment method in the trusted execution environment includes the following steps:
[0086] S1. According to the primary key join definition, determine whether the join key B in table S satisfies the primary key constraint; if it does, use the primary key join algorithm; otherwise, use the equijoin optimization algorithm.
[0087] When using the equijoin optimization algorithm, perform the following steps:
[0088] S21. Sort tables R and S by the join key B.
[0089] S22. Calculate the maximum prefix sums of tables R and S on the join key B respectively to obtain tables R1 and S1.
[0090] S23. Calculate the join degrees of table R1 using table S1 to obtain table R2, and calculate the join degrees of table S1 using table R1 to obtain table S2.
[0091] S24. Perform concealment expansion on tables R2 and S2 respectively to obtain tables R3 and S3.
[0092] S25. Rearrange table S3 to obtain table S4.
[0093] S26. Horizontally merge the data of tables R3 and S4 in each server, and finally extract columns A, B, and C of the merged result as the output result.
[0094] Specifically, first execute S1. According to the primary key join definition, determine whether the join key B in table S satisfies the primary key constraint, which is equivalent to determining whether the join key B is the primary key for the join of tables R and S. In database design, the primary key is a field or combination of fields used to uniquely identify each row record in a table, and it has uniqueness and non-nullability.
[0095] If it does, use the primary key join algorithm. In this embodiment, when using the primary key join algorithm, perform the following steps:
[0096] S31. Assign key values to the tuples in table S.
[0097] S32. Sort the join key B within the partitions of the table R on each server, select the first tuple with the same value in the join key B as the representative tuple, mark other tuples as inactive tuples, and record the corresponding original server ID at the same time;
[0098] S33. Allocate key values for the representative tuples according to the join key B, and randomly allocate the inactive tuples to other servers;
[0099] S34. On each server, perform a join calculation between the representative tuples and the local tuples in the table S;
[0100] S35. Transmit the data after the join calculation to the original server, and distribute the results of the representative tuples to the corresponding inactive tuples to form the final result.
[0101] In this embodiment, in S31, the steps of the key value allocation specifically include:
[0102] S311. Define the tuple in the form of a key-value pair, and each tuple is denoted as x = (k, v), where k is the key and v is the value;
[0103] S312. Calculate the key value allocation operation t = h(k) through the random oracle h to determine the allocation target of the tuple;
[0104] S313. According to the calculation result of h(k), gather the tuples with the same t to the same server;
[0105] S314. Use public parameters for adjustment and padding during the key value allocation process.
[0106] It should be noted that the random oracle is a prior art and is an idealized function model that can generate random outputs. Formally, this random oracle can be defined by the following three basic properties:
[0107] 1) Input and output: The random oracle accepts inputs of any length and returns a random value of a fixed length;
[0108] 2) Randomness: For each new input, the output value is randomly selected and uniformly distributed in the output space;
[0109] 3) Consistency: For the same input, no matter how many times it is queried, the output of the random oracle is always fixed.
[0110] In this embodiment, in S33, the steps of randomly allocating the inactive tuples to other servers specifically include:
[0111] S331. Generate a random number with the number of target servers as the upper limit;
[0112] S332. Each tuple evenly and independently determines a target server for allocation according to the random number.
[0113] In this embodiment, in S34, on each server, the representative tuple is joined with the local tuples in Table S. The specific join algorithm can use the industry - common algorithms, such as Nested Loop Join algorithm, HashJoin algorithm, Merge Join algorithm, etc. The present invention will not elaborate on this.
[0114] Of course, in addition to the above primary key join algorithm, other primary key join algorithms well - known to those skilled in the art can also be used, and the present invention does not limit this.
[0115] If the join of Table R and Table S does not satisfy the primary key constraint, then the equivalent - value join optimization algorithm is adopted. In this embodiment, when the equivalent - value join optimization algorithm is adopted, the following steps are executed:
[0116] S21. Sort Table R and Table S according to the join key B.
[0117] S22. Calculate the maximum prefix sums of Table R and Table S on the join key B respectively, to obtain Table R1 (A, B, DR) and Table S1 (B, C, DS).
[0118] S23. Use Table S1 (B, C, DS) to calculate the join degree of Table R1 (A, B, DR) to obtain Table R2 (A, B, DS), and use Table R1 (A, B, DR) to calculate the join degree of Table S1 (B, C, DS) to obtain Table S2 (B, C, DR).
[0119] S24. Perform stealthy expansion on Table R2 (A, B, DS) and Table S2 (B, C, DR) respectively to obtain Table R3 (A, B) and Table S3 (B, C).
[0120] S25. Rearrange Table S3 (B, C) to obtain Table S4 (B, C).
[0121] S26. Horizontally merge the data of Table R3 (A, B) and Table S4 (B, C) within each server, and finally extract columns A, B, and C of the merged result as the output result.
[0122] The implementation of the equivalent - value join optimization algorithm is further described below with a specific example.
[0123] First, execute S21 to sort the table R and the table S according to the join key B, and obtain the sorted table R and table S as follows:
[0124]
[0125] Then, execute S22 to calculate the maximum prefix sums of the table R and the table S on the join key B respectively, and obtain the table R1 (A, B, DR) and the table S1 (B, C, DS).
[0126] In the S22, calculating the maximum prefix sum of the table R includes:
[0127] S221. For the table R, calculate the prefix sum of each element on the join key B;
[0128] S222. For the table R, add a column DR, and refresh the corresponding column DR of all elements with the same prefix with the maximum prefix sum of the same prefix, to obtain the table R1 (A, B, DR);
[0129] Calculating the maximum prefix sum of the table S includes:
[0130] S223. For the table S, calculate the prefix sum of each element on the join key B;
[0131] S224. For the table S, add a column DS, and refresh the corresponding column DS of all elements with the same prefix with the maximum prefix sum of the same prefix, to obtain the table S1 (B, C, DS).
[0132] Among them, the algorithm of the prefix sum is as follows:
[0133] E1. Define the prefix sum: The prefix sum represents the cumulative number of times a certain prefix appears in the tuple;
[0134] E2. Initialize the cumulative count: At the beginning of the calculation, the cumulative number of times is set to 0;
[0135] E3. Traverse the tuple and accumulate: Scan the tuple in turn, and when encountering the same prefix, add 1 to the cumulative number of times;
[0136] E4. Detect the prefix change: When encountering a new prefix, the cumulative number of times starts to be calculated from 1 again.
[0137] In this embodiment, the obtained table R1 (A, B, DR) is shown in the following table:
[0138]
[0139] The obtained table S1 (B, C, DS) in this embodiment is as follows:
[0140]
[0141] Then, execute S23. Calculate the connection degree of the table R1 (A, B, DR) using the table S1 (B, C, DS) to obtain the table R2 (A, B, DS), and calculate the connection degree of the table S1 (B, C, DS) using the table R1 (A, B, DR) to obtain the table S2 (B, C, DR).
[0142] Specifically, in the S23, calculating the connection degree of the table R1 (A, B, DR) using the table S1 (B, C, DS) includes:
[0143] S231. Delete the column C of the table S1 (B, C, DS), and remove duplicates from the data to obtain a table S B (B, DS) with the connection key B as the primary key column;
[0144] S232. According to the table R1 (A, B, DR) and the table S B (B, DS), use the primary key connection algorithm to obtain the table R L (A, B, DR, DS);
[0145] S233. Delete the data in the table R L (A, B, DR, DS) where the value of DS degree is 0, and delete the DR column to obtain the table R2 (A, B, DS);
[0146] After executing S231, the obtained table S B (B, DS) is as follows:
[0147]
[0148] After executing S232, the obtained table R L (A, B, DR, DS) is as follows:
[0149]
[0150] After executing S2333, the obtained table R2 (A, B, DS) is as follows:
[0151]
[0152] It should be noted that the "degree" here should be understood as the number of occurrences of tuples in the future result, and the partition can represent different servers.
[0153] In this embodiment, calculating the connection degree of the table S1(B, C, DS) using the table R1(A, B, DR) includes:
[0154] S234. Delete column A of the table R1(A, B, DR), and remove duplicates from the data to obtain a table R B (B, DR) with the connection key B as the primary key column;
[0155] S235. According to the table S1(B, C, DS) and the table R B (B, DR), use the primary key connection algorithm to obtain a table S L (B, C, DS, DR);
[0156] S236. Delete the data in S L (B, C, DS, DR) where the value of the DR degree is 0, and delete column DS to obtain the table S2 (B, C, DR).
[0157] After executing S234, the obtained table R B (B, DR) is shown in the following table:
[0158]
[0159] After executing S235, the obtained table S L (B, C, DS, DR) is shown in the following table:
[0160]
[0161] After executing S236, the obtained table S2 (B, C, DR) is shown in the following table:
[0162]
[0163] Then execute S24 to perform stealthy expansion on the table R2(A, B, DS) and the table S2(B, C, DR) respectively to obtain a table R3(A, B) and a table S3(B, C).
[0164] In this embodiment, in S24, performing stealthy expansion on the table R2 includes:
[0165] Expand columns A and B according to the value of the DS degree in the table R2(A, B, DS), and then remove column DS. The expansion method includes but is not limited to copying. The obtained table R3(A, B) is shown in the following table:
[0166]
[0167] Performing stealth expansion on the table S2(B, C, DR) includes:
[0168] Expanding columns B and C according to the value of the DS degree in the table S2(B, C, DR), and then removing column DR, the obtained table S3(B, C) is as follows:
[0169]
[0170] Then execute S25 to rearrange the table S3(B, C) to obtain the table S4(B, C).
[0171] In this embodiment, in the S25, rearranging the table S3(B, C) includes:
[0172] S251. Add a column I to the S3(B, C), the value of the column I is the prefix sum of the connection key B, and use a column J to store the global minimum sequence number of the same prefix;
[0173] S252. Calculate the global sorted column L according to the column I and the column J, perform a modulo operation on the global sorted column L to obtain the server serial number column T, then delete the column I and the column J, and perform stealth filling to obtain S T (B, C, L, T);
[0174] S253. Allocate the data of the table S T (B, C, L, T) to the corresponding servers, and sort the data in each server according to the server serial number column L;
[0175] S254. Delete the L column and the T column of the table S T (B, C, L, T) to obtain the table S4(B, C).
[0176] In this embodiment, first execute S251, add a column I to the S3(B, C), the value of the column I is the prefix sum of the connection key B, and use a column J to store the global minimum sequence number of the same prefix, and the obtained new table is as follows:
[0177]
[0178] Then execute S252, calculate the global sorted column L according to the column I and the column J, perform a modulo operation on the global sorted column L to obtain the server serial number column T, then delete the column I and the column J, and perform stealth filling to obtain S T (B, C, L, T) as follows:
[0179]
[0180] Then execute S253, and allocate the data of the table S T (B, C, L, T) to the corresponding servers, and sort the data in each server according to the server serial number column L.
[0181] Next, execute S254 to delete the L column and T column of the table S T (B, C, L, T), and obtain the table S4(B, C) as follows:
[0182]
[0183] Finally, execute S26 to horizontally merge the data of the table R3(A, B) and the table S4(B, C) in each server, and finally extract the A, B, and C columns of the merged result as the output result, specifically as follows:
[0184]
[0185] The following illustrates the performance improvement of the primary key join algorithm and the equi-join optimization algorithm of the present invention through experiments.
[0186] 1) Primary key join algorithm
[0187] Experiment hypothesis: Two groups of experiments are conducted, using the traditional join algorithm and the optimized primary key join algorithm of the present invention (considering the primary key constraint) respectively. The present invention conducts a primary key join experiment on the famous TPC-H data set, and uses the following query to evaluate the join algorithms of the present invention and Opaque and SODA:
[0188] SELECT * FROM orders JOIN customer ON o_custkey = c_custkey;
[0189] c_custkey is the key of the customer table, the scaling factor is 10, the amount of customer data is 150,000, and the amount of orders data is 1,500,000. Through experiments, it is found that the time-consuming of the primary key join algorithm of the present invention is 6.44 s, while the time-consuming of the join algorithms of Opaque and SODA are 12.5 s and 11 s respectively.
[0190] 2) Equi-join algorithm
[0191] Experiment: Since Opaque does not support general joins, the SODA algorithm is compared with the equi-join algorithm proposed in the present invention here.
[0192] The data volumes of A and B are both 42,000, and the frequency distributions are as follows: the enumeration frequency of a1 in Table A is approximately 100, and the enumeration frequency of b1 in Table B is approximately 8. The query statement is as follows:
[0193] SELECT * FROM A a JOIN B b ON o_a.a1 = b.b1;
[0194] Through experiments, it is found that the time consumption of the equivalent join optimization algorithm of the present invention is 25 s, while the time consumption of the SODA algorithm is 41 s.
[0195] Through the above experiments and data comparison, it can be seen that the primary key join algorithm and the equivalent join algorithm provided by the present invention can significantly improve the performance of the TEE database system, especially while ensuring data privacy, optimizing the join algorithm and query efficiency.
[0196] Based on the same inventive concept, an embodiment of the present invention further provides a readable storage medium, on which a computer program is stored, and when the computer program is executed, it can implement the distributed concealment method in the trusted execution environment as described above.
[0197] A readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. For example, it can be but is not limited to an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the above. The computer programs described herein can be downloaded from the readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter or network interface in each computing / processing device receives the computer program from the network and forwards the computer program for storage in the readable storage medium in each computing / processing device. The computer program for performing the operations of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages - such as Smalltalk, C++, etc., and conventional procedural programming languages - such as the "C" language or similar programming languages. The computer program can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet). In some embodiments, by using the state information of the computer program to customize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute computer-readable program instructions to implement various aspects of the present invention.
[0198] Aspects of the present invention are described herein with reference to the flowcharts and / or block diagrams of methods, systems, and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer programs. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these programs are executed by the processor of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. These computer programs can also be stored in a readable storage medium, and these computer programs cause a computer, a programmable data processing device, and / or other devices to work in a specific manner. Thus, the readable storage medium storing the computer programs includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams.
[0199] The computer programs can also be loaded onto a computer, other programmable data processing device, or other device, such that a series of operation steps are executed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, thereby causing the computer programs executed on the computer, other programmable data processing device, or other device to implement the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams.
[0200] In summary, the embodiments of the present invention provide a distributed concealment method and a readable storage medium in a trusted execution environment. Compared with the prior art, innovations and optimizations are mainly made in aspects such as security, primary key connection algorithm, implementation and optimization of general equality connection, and performance improvement. The present invention not only improves the computing power and performance of database connection operations in the TEE environment, but also ensures data privacy through security policies, and has a wide range of application prospects and can be applied to multiple fields such as cloud databases, data security computing, and big data analysis.
[0201] In addition, it should also be recognized that although the present invention has been disclosed above with preferred embodiments, the above embodiments are not intended to limit the present invention. For any person skilled in the art, without departing from the scope of the technical solution of the present invention, many possible changes and modifications can be made to the technical solution of the present invention by using the technical content disclosed above, or modified into equivalent embodiments with equivalent changes. Therefore, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A distributed anonymity method in a trusted execution environment, for a Join operator in a database, wherein the Join operator takes a table R and a table S as inputs and outputs all tuple combinations that satisfy a join condition through a join key B, where R = R(A, B) and S = S(B, C); characterized in that: The distributed anonymity method comprises: S1. According to the primary key connection definition, determine whether the connection key B in the table S satisfies the primary key constraint; if so, use the primary key connection algorithm; otherwise, use the equivalue connection optimization algorithm; When the equivalue join optimization algorithm is used, the following steps are performed: S21, sorting the table R and the table S according to the connection key B; S22, respectively calculating the maximum prefix sum of the table R and the table S on the connection key B to obtain table R1 and table S1; S23, using the table S1 to calculate the connectivity of the table R1 to obtain table R2, and using the table R1 to calculate the connectivity of the table S1 to obtain table S2; S24, performing concealment expansion on the table R2 and the table S2 respectively to obtain table R3 and table S3; S25, rearrange the table S3 to obtain table S4; S26. Merge the data of the table R3 and the table S4 in each server horizontally, and finally extract columns A, B, and C of the merged result as output results.
2. The distributed anonymity method in a trusted execution environment according to claim 1, characterized in that: When the primary key join algorithm is used, the following steps are performed: S31, assigning key values to the tuples in the table S; S32, sorting the connection key B in the partitions of the table R on each server, and selecting the first tuple with the same value in the connection key B as the representative tuple, marking the other tuples as inactive tuples, and recording the corresponding original server ID; S33, assigning key values to the representative tuples according to the connection key B, and randomly assigning the inactive tuples to other servers; S34, on each server, performing a join calculation on the representative tuple and the local tuple in the table S; S35, transmitting the data after the connection calculation to the original server, and distributing the result of the representative tuple to the corresponding inactive tuple to form a final result.
3. The distributed anonymity method in a trusted execution environment according to claim 2, characterized in that: In the S31, the key value allocation includes: S311. Define a tuple as a key-value pair, where each tuple is denoted by x = (k, v), where k is the key and v is the value; S312, calculating the key value assignment operation t = h(k) through the random oracle h, and determining the assignment target of the tuple; S313, according to the calculation result of h(k), tuples with the same t are aggregated to the same server; S314. Use common parameters to adjust and fill in the key value allocation process.
4. The distributed anonymity method in a trusted execution environment according to claim 1, characterized in that: In S22, calculating the maximum prefix sum of the table R includes: S221. For the table R, calculate the prefix sum of each element on the connection key B; S222, for the table R, add a column DR, and use the maximum prefix sum of the same prefix to refresh all columns DR corresponding to the elements with the same prefix, to obtain the table R1; Calculating the maximum prefix sum of the table S includes: S223. For the table S, calculate the prefix sum of each element on the connection key B; S224. For the table S, add a column DS, and use the maximum prefix sum of the same prefix to refresh the columns DS corresponding to all elements with the same prefix, to obtain the table S1.
5. The distributed anonymity method in a trusted execution environment according to claim 4, characterized in that: The prefix sum algorithm is as follows: E1, the prefix sum represents the cumulative number of times a prefix appears in the tuple; E2. At the beginning of the calculation, the cumulative number is set to 0; E3. Scan the tuples in sequence. When the same prefix is encountered, the cumulative number is increased by 1. E4. When a new prefix is encountered, the cumulative number is recalculated from 1.
6. The distributed anonymity method in a trusted execution environment according to claim 4, characterized in that: In S23, using the table S1 to calculate the connectivity of the table R1 includes: S231, delete column C of the table S1, and deduplicate the data to obtain a table S1 with the connection key B as the primary key column. B ; S232, according to the table R1 and the table S B , use the primary key join algorithm to get table R L ; S233, delete the table R L The data whose DS value is 0 in the table are deleted, and the DR column is deleted to obtain the table R2; Calculating the connectivity of the table S1 using the table R1 includes: S234, delete column A of the table R1, and deduplicate the data to obtain a table R1 with the connection key B as the primary key column. B ; S235, according to the table S1 and the table R B , use the primary key join algorithm to get table S L ; S236, delete the S L The data with a DR value of 0 are removed and the DS column is deleted to obtain the table S2.
7. The distributed anonymity method in a trusted execution environment according to claim 6, characterized in that: In S24, performing a concealment extension on the table R2 includes: Expand columns A and B according to the values of DS in table R2, and then remove column DS to obtain table R3; The concealed expansion of Table S2 includes: According to the values of DS in Table S2, columns B and C are expanded, and then column DR is removed to obtain Table S3.
8. The distributed anonymity method in a trusted execution environment according to claim 1 or 7, characterized in that: In S25, rearranging the table S3 includes: S251, add an I column to S3, the value of the I column is the prefix sum of the connection key B, and use a J column to store the global minimum sequence of the same prefix; S252: Calculate a global sorting column L based on the I column and the J column, perform a modulo operation on the global sorting column L to obtain a server sequence number column T, then delete the I column and the J column, perform hidden padding, and obtain S T ; S253, according to the server sequence number column T, the table S T The data is distributed to corresponding servers, and the data is sorted in each server according to the server sequence number column L; S254, delete the table S T The L and T columns of Table S4 are obtained.
9. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the distributed anonymity method in a trusted execution environment according to any one of claims 1 to 8 can be implemented.
Citation Information
Patent Citations
Browser-oriented hidden query method and system and storage medium
CN118606539A
Privacy protection dynamic space keyword query method and device under multi-attribute cost constraint, electronic equipment and storage medium
CN118643055A