Distributed hiding method in trusted execution environment and readable storage medium
By providing a distributed occult Join algorithm in a trusted execution environment, the existing algorithms are solved for the use of limited use and security risks during primary key connections and equivalent connections, and efficient and secure database join operations are achieved, which is suitable for cloud databases and big data analysis and other fields.
Patent Information
- Application Number
- CN202510472807.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing distributed obscure Join algorithms have limited use and data security risks, especially when primary key connections and equivalent connections, the generality is insufficient and sensitive information needs to be disclosed.
A distributed occult method is provided in a trusted execution environment, and a primary key connection or equivalent connection optimization algorithm is selected by judging whether the connection key in table S meets the primary key constraints. The equivalence connection optimization algorithm includes sorting, maximum prefix and calculation, connection degree calculation and obscure extension to achieve obscure join operations.
This method improves the security and performance of database join operations, enhances data privacy protection capabilities, supports a wider range of connection scenarios, and achieves efficient computing and obscure protection in TEE environment.
Smart Images

Figure CN120030530A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of database information security, and in particular to a distributed anonymity method and a readable storage medium in a trusted execution environment. Background Art
[0002] In the current environment, providing encrypted data computing for users has become one of the important services of cloud service providers. The core value of encryption is to ensure the confidentiality and privacy of data, especially sensitive data such as personal privacy information, financial documents and business secrets. However, computing on encrypted data usually brings high overhead and the computing speed may drop by several orders of magnitude.
[0003] To solve this problem, a common method is to perform calculations based on a trusted execution environment (TEE), that is, to decrypt the data after it enters the trusted area and re-encrypt it before leaving the area. However, the existing TEE solution still has security risks, and the host system can infer auxiliary information of the encrypted data by monitoring the memory access pattern of the application. Therefore, the calculations in the secure area should be as anonymous as possible.
[0004] Currently, the solutions to this problem are mainly divided into two methods: software and hardware.
[0005] 1) Hardware solution: Usually, hidden RAM (ORAM) technology is used, which can effectively prevent the leakage of data access patterns, but it will bring higher performance overhead and the calculation time will increase by a multiple of O(log²N).
[0006] 2) Software solution: Design an anonymous algorithm for specific computing operators (such as Join operators). Currently, there are two typical distributed anonymous Join algorithms: Opaque and SODA. Opaque is an anonymous sorting algorithm based on column sorting, which can implement Join operations, but its application scenarios are limited and only applicable to primary key connections. SODA believes that column sorting is too expensive and supports conventional equal-value Join, but it needs to disclose the frequency information of the most popular keys in the two tables, which has certain security defects.
[0007] Since the above methods still have significant challenges in terms of usage constraints and security, the current distributed hidden Join algorithm faces the following problems: 1) Usage restrictions: only primary key connections are supported, which affects its versatility; 2) Security risks: The frequency of keys or other auxiliary information needs to be disclosed, which affects data privacy. Summary of the invention
[0008] The purpose of the present invention is to provide a distributed anonymity method and a readable storage medium in a trusted execution environment to solve the problems of limited use and data security risks in existing distributed anonymity Join algorithms.
[0009] In order to achieve the above object, the present invention provides a distributed anonymity method in a trusted execution environment, which is used for a Join operator in a database. The Join operator takes a table R and a table S as inputs, and outputs all tuple combinations that meet the join condition through a join key B. Assume that R = R(A, B) and S = S(B, C); the distributed anonymity method includes: S1. According to the primary key connection definition, determine whether the connection key B in the table S satisfies the primary key constraint; if so, use the primary key connection algorithm; otherwise, use the equivalue connection optimization algorithm; When the equi-join optimization algorithm is used, the following steps are performed: S21, sorting the table R and the table S according to the connection key B; S22, respectively calculate the maximum prefix sum of the table R and the table S on the connection key B to obtain the table R 1 and Table S 1 ; S23, using the table S 1 For the table R 1 Calculate the connectivity and get table R 2 , and using the table R 1 For the table S 1 Calculate the connectivity and get table S 2 ; S24, respectively, the table R 2 And the table S 2 Perform hidden expansion to obtain table R 3 and Table S 3 ; S25, the table S 3 Rearrange to get table S 4 ; S26, the table R in each server 3 and the table S 4 The data is merged horizontally, and finally the A, B, and C columns of the merged result are extracted as the output result.
[0010] Optionally, when the primary key connection algorithm is adopted, the following steps are performed: S31, assigning key values to the tuples in the table S; S32, sorting the connection key B in the partitions of the table R on each server, and selecting the first tuple with the same value in the connection key B as the representative tuple, marking the other tuples as inactive tuples, and recording the corresponding original server ID; S33, assigning key values to the representative tuples according to the connection key B, and randomly assigning the inactive tuples to other servers; S34, on each server, performing a join calculation on the representative tuple and the local tuple in the table S; S35, transmitting the data after the connection calculation to the original server, and distributing the result of the representative tuple to the corresponding inactive tuple to form a final result.
[0011] Optionally, in S31, the key value allocation includes: S311. Define a tuple as a key-value pair, where each tuple is denoted by x = (k, v), where k is the key and v is the value; S312, calculating the key value assignment operation t = h(k) through the random oracle h, and determining the assignment target of the tuple; S313, according to the calculation result of h(k), tuples with the same t are aggregated to the same server; S314. Use public parameters to adjust and fill in the key value allocation process.
[0012] Optionally, in S22, calculating the maximum prefix sum of the table R includes: S221. For the table R, calculate the prefix sum of each element on the connection key B; S222: For the table R, add a column DR, and use the maximum prefix sum of the same prefix to refresh all columns DR corresponding to the same prefix elements, so as to obtain the table R. 1 ; Calculating the maximum prefix sum of the table S includes: S223. For the table S, calculate the prefix sum of each element on the connection key B; S224. For the table S, add a column DS, and use the maximum prefix sum of the same prefix to refresh the columns DS corresponding to all elements with the same prefix, and obtain the table S. 1 .
[0013] Optionally, the prefix sum algorithm is as follows: E1, the prefix sum represents the cumulative number of times a prefix appears in the tuple; E2. At the beginning of the calculation, the cumulative number is set to 0; E3. Scan the tuples in sequence. When the same prefix is encountered, the cumulative number is increased by 1. E4. When a new prefix is encountered, the cumulative number is recalculated from 1.
[0014] Optionally, in S23, using the table S 1 For the table R 1 Connectivity calculation includes: S231, delete the table S 1 Column C of the table is obtained by removing duplicate data and obtaining a table S with the connection key B as the primary key column. B ; S232, according to the table R 1 and the table S B , use the primary key join algorithm to get table R L ; S233, delete the table R L The data with DS value of 0 in the table is obtained by deleting the DR column. 2 ; Using the table R 1 For the table S 1 Connectivity calculation includes: S234, delete the table R 1 Column A of the table is obtained by removing duplicate data and obtaining a table R with the connection key B as the primary key column. B ; S235, according to the table S 1 and the table R B , use the primary key join algorithm to get table S L ; S236, delete the S L The data with DR value of 0 is deleted, and the DS column is deleted to obtain the table S 2 .
[0015] Optionally, in S24, the table R 2 The stealth extensions include: According to the table R 2 Expand columns A and B based on the value of DS, and then remove column DS to get table R 3 ; For the table S 2 The stealth extensions include: According to the table S 2 Expand columns B and C by the value of DS, and then remove column DR to get table S 3 .
[0016] Optionally, in S25, the table S 3 Rearrangements include: S251, the S 3 Add an I column, the value of which is the prefix sum of the connection key B, and use a J column to store the global minimum sequence of the same prefix; S252: Calculate a global sorting column L based on the I column and the J column, perform a modulo operation on the global sorting column L to obtain a server sequence number column T, then delete the I column and the J column, perform hidden padding, and obtain S T ; S253, according to the server serial number column T, the table S T The data is distributed to corresponding servers, and the data is sorted in each server according to the server sequence number column L; S254, delete the table S T The L and T columns of the table S are obtained. 4 .
[0017] Based on the same inventive concept, the present invention also provides a readable storage medium on which a computer program is stored. When the computer program is executed, it can implement the distributed anonymity method in a trusted execution environment as described above.
[0018] The distributed anonymity method and readable storage medium in a trusted execution environment provided by the present invention have at least one of the following beneficial effects: 1) Security enhancement: This invention introduces the definition of communication anonymity and computation anonymity to ensure that in distributed computing and secure execution environments (TEEs), attackers cannot infer information about input data from data exchange or memory access patterns. At the same time, an anonymity expansion algorithm is used to reduce the risk of data leakage by padding and randomizing data distribution, thereby enhancing data privacy protection capabilities; 2) More efficient primary key join algorithm: Existing join algorithms are usually not optimized specifically for primary key joins. The present invention reduces unnecessary repeated calculations during join calculations by utilizing primary key constraints. 3) Implementation and performance improvement of equi-join: When processing non-primary key joins, the present invention proposes a new equi-join algorithm, which fills and expands the table through prefix sum calculation, connection degree calculation, and hidden extension, so that it meets the hidden requirements, and realizes the calculation of general equi-join, and has a certain degree of performance improvement; 4) The present invention not only improves the computing power and performance of database connection operations in the TEE environment, but also ensures data privacy through security strategies. It has broad application prospects and can be applied to cloud databases, data security computing, big data analysis and other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Those skilled in the art should understand that the drawings are provided for a better understanding of the present invention and do not constitute any limitation on the scope of the present invention. Figure 1 A flowchart of a distributed anonymity method in a trusted execution environment provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0020] Before describing the present invention in detail, the basic concepts involved in distributed anonymity are first described.
[0021] 1. Distributed Algorithm In distributed computing, multiple servers work together to perform computing tasks through a specific algorithm, which is usually called a distributed algorithm. Each server holds a portion of the input data, and these data elements (for example, a total of N) are almost evenly distributed among the servers. Specifically, the i-th server has a subset of n=Θ(N / p) elements as its initial local data set, where N is the total amount of data, p is the number of servers, and {n} represents the amount of data on each server, which is public to all servers. In the multi-round computing process, in each round, the i-th server processes its local data set X and generates a set of outputs CY={Y 1 ,Y 2 ,…,Yp}, where i∈[p], this stage marks the completion of this round of calculation. After that, it enters the communication stage, and all servers exchange data in the network. Specifically, the i-th server transmits its output Y to the j-th server. After receiving the data from all other servers, the server merges these data and updates its local data set. Then, the updated data set will be used for the next round of calculation, or as the input of subsequent tasks, or as the final output.
[0022] To distinguish it from distributed algorithms, the term "stand-alone algorithm" is generally used to refer to the traditional computing setting with only one server. Typically, a stand-alone algorithm only involves local computing on the server and does not involve network communication.
[0023] 2. Confidential system In a confidential system, data on the server is uploaded in encrypted form. When a client submits a query request, the system converts the query into a series of database operations and executes these operations according to the underlying algorithm. Finally, the operation result, that is, the output of the query, will be returned to the client in encrypted form.
[0024] In this system, each server is equipped with a protected trusted execution environment (TEE) to ensure that the server's memory can accommodate the data elements it processes during computing. All data elements are encrypted and protected in storage outside the TEE, that is, the data is stored in plain text inside the TEE and in cipher text outside.
[0025] 3. Security Definition 3.1 Communication Anonymity: In a distributed environment, for any deterministic or randomized algorithm A and any input X, let S(X) = {S ij} represents a message sequence, where {S ij} represents the size of the message sent from the i-th server to the j-th server. The communication anonymity is defined as follows: If algorithm A has communication anonymity, then there exists a probabilistic simulator Y that can generate a simulated sequence S 1 =Y(n 1 , n 2 , …, np), such that no polynomial-time algorithm can distinguish the real S(X) from the simulated S with a success probability greater than 1 / 2. 1 In other words, the record of any input depends only on the size of the input on the server, ensuring that an attacker cannot infer any information about the input from the record.
[0026] 3.2 Computational Anonymity: For any stand-alone algorithm A, when it is executed on a server, the memory access operations of input X can be represented as a sequence ((op 1 , a 1 , x 1 ), …, (opk, ak, xk)), where each op represents a read or write operation. Specifically, a read operation refers to reading an element from memory location a, while a write operation refers to updating an element x to memory location a. For input X, the memory access pattern of algorithm A is defined as A(X) = (a 1 , …, ak), which is all the information that a memory access pattern adversary (such as a TEE bypass attacker) can observe. This paper defines computational anonymity as follows: If algorithm A has computational anonymity, then there exists a probabilistic simulator Sim that can generate a simulated computation sequence A 1 = Sim(n), so that no polynomial-time algorithm can distinguish A(X) from A with probability greater than 1 / 2. 1 .
[0027] 4. Fill output The above security definition assumes that the input size N is public, but the output size M is not public because the output depends on the data content. To ensure anonymity, the anonymity algorithm always pads the output to the maximum possible output size (that is, the worst output size when the input size is N). However, in the join operation, the worst output size is the product of the sizes of the two input tables. The join operation will cause the computational cost to grow exponentially, resulting in performance degradation or system unusability. Therefore, in actual implementation, it is often chosen to sacrifice security to a certain extent in exchange for better performance. In other words, the output size is not necessarily always padded to the worst output size, but to find a reasonable balance between security and privacy.
[0028] In order to make the purpose, advantages and features of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the accompanying drawings are in a very simplified form and use non-precise proportions, which are only used to conveniently and clearly assist in explaining the purpose of the embodiments of the present invention. In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, please refer to the accompanying drawings. It should be noted that the structure, proportion, size, etc. illustrated in the drawings of this specification are only used to match the content disclosed in the specification for people familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Any modification of the structure, change in the proportional relationship or adjustment of the size, under the same or similar conditions as the effects that can be produced by the present invention and the purposes that can be achieved, should still fall within the scope of the technical content disclosed by the present invention.
[0029] As used in this disclosure, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. As used in this disclosure, the term "or" is generally used in a sense including "and / or" unless the context clearly dictates otherwise. As used in this disclosure, Please refer to Figure 1 This embodiment provides a distributed anonymity method in a trusted execution environment, which is used for a Join operator in a database. The Join operator takes a table R and a table S as inputs, and outputs all tuple combinations that meet the join condition through a join key B. Assume that R = R(A, B) and S = S(B, C); the distributed anonymity method in a trusted execution environment includes the following steps: S1. According to the primary key connection definition, determine whether the connection key B in the table S satisfies the primary key constraint; if so, use the primary key connection algorithm; otherwise, use the equivalue connection optimization algorithm; When the equi-join optimization algorithm is used, the following steps are performed: S21, sorting the table R and the table S according to the connection key B; S22, respectively calculate the maximum prefix sum of the table R and the table S on the connection key B to obtain the table R 1 and Table S 1 ; S23, using the table S 1 For the table R 1 Calculate the connectivity and get table R 2 , and using the table R 1 For the table S 1 Calculate the connectivity and get table S 2 ; S24, respectively, the table R 2 And the table S 2 Perform hidden expansion to obtain table R 3 and Table S 3 ; S25, the table S 3 Rearrange to get table S 4 ; S26, the table R in each server 3 and the table S 4 The data is merged horizontally, and finally the A, B, and C columns of the merged result are extracted as the output result.
[0030] Specifically, S1 is executed first, and according to the primary key connection definition, it is determined whether the connection key B in the table S satisfies the primary key constraint, which is equivalent to determining whether the connection key B is the primary key connecting the table R and the table S. In database design, the primary key is a field or a combination of fields used to uniquely identify each row of records in the table, and is unique and non-empty.
[0031] If the conditions are met, the primary key connection algorithm is adopted. In this embodiment, when the primary key connection algorithm is adopted, the following steps are performed: S31, assigning key values to the tuples in the table S; S32, sorting the connection key B in the partitions of the table R on each server, and selecting the first tuple with the same value in the connection key B as the representative tuple, marking the other tuples as inactive tuples, and recording the corresponding original server ID; S33, assigning key values to the representative tuples according to the connection key B, and randomly assigning the inactive tuples to other servers; S34, on each server, performing a join calculation on the representative tuple and the local tuple in the table S; S35, transmitting the data after the connection calculation to the original server, and distributing the result of the representative tuple to the corresponding inactive tuple to form a final result.
[0032] In this embodiment, in S31, the key value allocation step specifically includes: S311. Define a tuple as a key-value pair, where each tuple is denoted by x = (k, v), where k is the key and v is the value; S312, calculating the key value assignment operation t = h(k) through the random oracle h, and determining the assignment target of the tuple; S313, according to the calculation result of h(k), tuples with the same t are aggregated to the same server; S314. Use public parameters to adjust and fill in the key value allocation process.
[0033] It should be noted that the random oracle is an existing technology and is an idealized function model that can generate random output. Formally, the random oracle can be defined by the following three basic properties: 1) Input and output: The random oracle accepts input of any length and returns a random value of fixed length; 2) Randomness: For each new input, the output value is randomly selected and evenly distributed in the output space; 3) Consistency: For the same input, no matter how many times it is queried, the output of the random oracle is always fixed.
[0034] In this embodiment, in S33, the step of randomly allocating the inactive tuples to other servers specifically includes: S331, generating a random number with the number of target servers as the upper limit; S332. Each tuple is evenly and independently assigned to a target server determined according to the random number.
[0035] In this embodiment, in S34, on each server, the representative tuple is connected with the local tuple in the table S for calculation. The specific connection algorithm can use the common algorithm in the industry, such as the Nested Loop Join algorithm, the HashJoin algorithm, the Merge Join algorithm, etc., which will not be elaborated in the present invention.
[0036] Of course, in addition to the above primary key connection algorithm, other primary key connection algorithms well known to those skilled in the art may also be used, and the present invention is not limited thereto.
[0037] If the connection between the table R and the table S does not satisfy the primary key constraint, an equi-join optimization algorithm is used. In this embodiment, when an equi-join optimization algorithm is used, the following steps are performed: S21, sorting the table R and the table S according to the connection key B; S22, respectively calculate the maximum prefix sum of the table R and the table S on the connection key B to obtain the table R 1 (A, B, DR) and Table S 1 (B, C, DS); S23, using the table S 1 (B, C, DS) for the table R 1 (A, B, DR) performs connectivity calculation and obtains table R 2 (A, B, DS), and using the table R 1 (A, B, DR) for the table S 1 (B, C, DS) calculates the connectivity and obtains Table S 2 (B, C, DR); S24, respectively, the table R 2 (A, B, DS) and said Table S 2 (B, C, DR) is extended secretly to obtain table R 3 (A, B ) and Table S 3 (B, C); S25, the table S 3 (B, C) are rearranged to obtain Table S 4 (B, C); S26, the table R in each server 3 (A, B) and the table S 4 The data of (B, C) are merged horizontally, and finally the A, B, and C columns of the merged result are extracted as the output result.
[0038] The implementation of the equivalue connection optimization algorithm is further explained below with a specific example.
[0039] First, S21 is executed to sort the table R and the table S according to the connection key B, and the sorted table R and the table S are obtained as follows:
[0040] Then, S22 is executed to respectively calculate the maximum prefix sum of the table R and the table S on the connection key B to obtain the table R 1 (A, B, DR) and Table S 1 (B, C, DS).
[0041] In S22, calculating the maximum prefix sum of the table R includes: S221. For the table R, calculate the prefix sum of each element on the connection key B; S222: For the table R, add a column DR, and use the maximum prefix sum of the same prefix to refresh all columns DR corresponding to the same prefix elements, so as to obtain the table R. 1 (A, B, DR); Calculating the maximum prefix sum of the table S includes: S223. For the table S, calculate the prefix sum of each element on the connection key B; S224. For the table S, add a column DS, and use the maximum prefix sum of the same prefix to refresh the columns DS corresponding to all elements with the same prefix, and obtain the table S. 1 (B, C, DS).
[0042] The prefix sum algorithm is as follows: E1. Define prefix sum: the prefix sum represents the cumulative number of times a prefix appears in the tuple; E2. Initialize the cumulative count: At the beginning of the calculation, the cumulative count is set to 0; E3. Traverse the tuples and accumulate: Scan the tuples in sequence, and when the same prefix is encountered, the cumulative number is increased by 1; E4. Detect prefix changes: When a new prefix is encountered, the cumulative number of times is counted again from 1.
[0043] In this embodiment, the obtained table R 1 (A, B, DR) are shown in the following table:
[0044] Table S obtained in this example 1 (B, C, DS) are shown in the following table:
[0045] Then execute S23, using the table S 1 (B, C, DS) for the table R 1 (A, B, DR) performs connectivity calculation and obtains table R 2 (A, B, DS), and using the table R 1 (A, B, DR) for the table S 1 (B, C, DS) calculates the connectivity and obtains Table S 2 (B, C, DR).
[0046] Specifically, in S23, using the table S 1 (B, C, DS) for the table R 1 (A, B, DR) connectivity calculation includes: S231, delete the table S 1 (B, C, DS), and deduplicate the data to obtain a table S with the join key B as the primary key column. B (B, DS); S232, according to the table R 1 (A, B, DR) and said Table S B (B, DS), use the primary key join algorithm to get table R L (A, B, DR, DS); S233, delete the table R L The data with DS value of 0 in (A, B, DR, DS) and the DR column are deleted to obtain the table R 2 (A, B, DS); After executing S231, the table S B (B, DS) are shown in the following table:
[0047] After executing S232, the table R L (A, B, DR, DS) are shown in the following table:
[0048] After executing S2333, the table R 2 (A, B, DS) are shown in the following table:
[0049] It should be noted that the "degree" here should be understood as the number of occurrences of tuples in future results, and partitions can represent different servers.
[0050] In this embodiment, using the table R 1 (A, B, DR) for the table S 1 (B, C, DS) Connectivity calculation includes: S234, delete the table R 1 (A, B, DR), and deduplicate the data to obtain a table R with the join key B as the primary key column. B (B, DR); S235, according to the table S 1 (B, C, DS) and said table RB (B, DR), use the primary key join algorithm to get table S L (B, C, DS, DR); S236, delete the S L The data with DR value of 0 in (B, C, DS, DR) are deleted, and the DS column is deleted to obtain the table S 2 (B, C, DR).
[0051] After executing S234, the table R B (B, DR) are shown in the following table:
[0052] After executing S235, the obtained table S L (B, C, DS, DR) are shown in the following table:
[0053] After executing S236, the table S 2 (B, C, DR) are shown in the following table:
[0054] Then, S24 is executed to respectively 2 (A, B, DS) and said Table S 2 (B, C, DR) is extended secretly to obtain table R 3 (A, B ) and Table S 3 (B, C).
[0055] In this embodiment, in S24, the table R 2 The stealth extensions include: According to the table R 2 The value of DS in (A, B, DS) is used to expand columns A and B, and then remove column DS. The expansion method includes but is not limited to copying. The resulting table R 3 (A, B ) are shown in the following table:
[0056] For the table S 2 (B, C, DR) Perform hidden expansion including: According to the table S 2 The DS value in (B, C, DR) is expanded by columns B and C, and then the DR column is removed, resulting in table S 3 (B, C) are shown in the following table:
[0057] Then execute S25 to 3 (B, C) are rearranged to obtain Table S 4 (B, C).
[0058] In this embodiment, in S25, the table S 3 (B, C) Rearrangements include: S251, the S 3 (B, C) add an I column, the value of which is the prefix sum of the connection key B, and use a J column to store the global minimum sequence of the same prefix; S252: Calculate a global sorting column L based on the I column and the J column, perform a modulo operation on the global sorting column L to obtain a server sequence number column T, then delete the I column and the J column, perform hidden padding, and obtain S T (B, C, L, T); S253, according to the server serial number column T, the table S T The data of (B, C, L, T) are distributed to the corresponding servers, and the data are sorted in each server according to the server sequence number column L; S254, delete the table S T The L and T columns of (B, C, L, T) are obtained as shown in Table S 4 (B, C).
[0059] In this embodiment, S251 is first executed to 3 (B, C) adds an I column, the value of which is the prefix sum of the connection key B, and uses a J column to store the global minimum sequence of the same prefix, and the new table is as follows:
[0060] Then, S252 is executed to calculate the global sorting column L according to the I column and the J column, perform a modulo operation on the global sorting column L to obtain the server sequence number column T, and then delete the I column and the J column, perform hidden filling, and obtain S T (B, C, L, T) are as follows:
[0061] Then execute S253, according to the server sequence number column T, the table S T The data of (B, C, L, T) are distributed to corresponding servers, and the data are sorted in each server according to the server sequence number column L.
[0062] Then execute S254 to delete the table ST The L and T columns of (B, C, L, T) are obtained as shown in Table S 4 (B, C) are as follows:
[0063] Finally, S26 is executed to convert the table R in each server into 3 (A, B) and the table S 4 The data of (B, C) are merged horizontally, and finally the A, B, and C columns of the merged result are extracted as the output results, as follows:
[0064] The following experiments illustrate the performance improvement of the primary key join algorithm and the equivalue join optimization algorithm of the present invention.
[0065] 1) Primary key join algorithm Experimental hypothesis: Two sets of experiments are conducted, using the traditional join algorithm and the optimized primary key join algorithm of the present invention (taking primary key constraints into consideration). The present invention conducts primary key join experiments on the famous TPC-H dataset, and uses the following queries to evaluate the join algorithms of the present invention and Opaque and SODA: SELECT * FROM orders JOIN customer ON o_custkey = c_custkey; c_custkey is the key of the customer table, the scaling factor is 10, the customer data volume is 150000, and the orders data volume is 1500000. It is found through experiments that the primary key join algorithm of the present invention takes 6.44 seconds, while the join algorithms of Opaque and SODA take 12.5 seconds and 11 seconds respectively.
[0066] 2) Equi-connection algorithm Experiment: Because Opaque does not support general connections, the SODA algorithm is compared with the equi-join algorithm proposed in the present invention.
[0067] The data volume A and B are both 42,000, and the frequency distribution is as follows: the enumeration frequency of a1 in table A is about 100, and the enumeration frequency of b1 in table B is about 8. The query statement is as follows: SELECT * FROM A a JOIN B b ON o_a.a1= b.b1; It is found through experiments that the time consumption of the equivalent connection optimization algorithm of the present invention is 25 seconds, while the time consumption of the SODA algorithm is 41 seconds.
[0068] Through the above experiments and data comparisons, it can be seen that the primary key connection algorithm and the equivalent connection algorithm provided by the present invention can significantly improve the performance of the TEE database system, especially while ensuring data privacy, while optimizing the connection algorithm and query efficiency.
[0069] Based on the same inventive concept, an embodiment of the present invention further proposes a readable storage medium on which a computer program is stored. When the computer program is executed, it can implement the distributed anonymity method in a trusted execution environment as described above.
[0070] The readable storage medium can be a tangible device that can keep and store the instructions used by the instruction execution device, such as but not limited to an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device or any suitable combination thereof. The more specific example (non-exhaustive list) of the readable storage medium includes: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a convex structure in a groove on which instructions are stored, and any suitable combination thereof. The computer program described herein can be downloaded from the readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer program from the network and forwards the computer program for storage in the readable storage medium in each computing / processing device. The computer program for performing the operation of the present invention can be an assembly instruction, an instruction set architecture (ISA) instruction, a machine instruction, a machine-related instruction, a microcode, a firmware instruction, a state setting data, or a source code or object code written in any combination of one or more programming languages, including object-oriented programming languages-such as Smalltalk, C++, etc., and conventional procedural programming languages-such as "C" language or similar programming languages. The computer program can be executed completely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on the remote computer, or completely on the remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet). In some embodiments, by utilizing the state information of a computer program to personalize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute computer-readable program instructions to implement various aspects of the present invention.
[0071] Here, various aspects of the present invention are described with reference to the flowchart and / or block diagram of the method, system and computer program product according to the embodiment of the present invention. It should be understood that each square frame of the flowchart and / or block diagram and the combination of the square frames in the flowchart and / or block diagram can be realized by a computer program. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device, so as to produce a machine, so that these programs, when executed by the processor of a computer or other programmable data processing device, produce a device for realizing the function / action specified in one or more square frames in the flowchart and / or block diagram. These computer programs can also be stored in a readable storage medium, and these computer programs make the computer, programmable data processing device and / or other equipment work in a specific way, so that the readable storage medium storing the computer program includes a manufactured product, which includes instructions for realizing various aspects of the function / action specified in one or more square frames in the flowchart and / or block diagram.
[0072] The computer program may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are executed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the computer program executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0073] In summary, the embodiments of the present invention provide a distributed anonymity method and a readable storage medium in a trusted execution environment. Compared with the prior art, the invention mainly innovates and optimizes security, primary key connection algorithm, implementation and optimization of general equivalent connection, and performance improvement. The present invention not only improves the computing power and performance of database connection operations in the TEE environment, but also ensures data privacy through security policies. It has broad application prospects and can be applied to cloud databases, data security computing, big data analysis and other fields.
[0074] In addition, it should be recognized that although the present invention has been disclosed as a preferred embodiment, the above embodiment is not intended to limit the present invention. For any technician familiar with the art, without departing from the scope of the technical solution of the present invention, the technical content disclosed above can be used to make many possible changes and modifications to the technical solution of the present invention, or modified into equivalent embodiments of equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still belongs to the scope of protection of the technical solution of the present invention.
Claims
1. A distributed anonymity method in a trusted execution environment, for a Join operator in a database, wherein the Join operator takes a table R and a table S as inputs and outputs all tuple combinations that satisfy the join condition through a join key B, assuming that R = R(A, B) and S = S(B, C); characterized in that, The distributed anonymity method comprises: S1. According to the primary key connection definition, determine whether the connection key B in the table S satisfies the primary key constraint; if so, use the primary key connection algorithm; otherwise, use the equivalue connection optimization algorithm; When the equi-join optimization algorithm is used, the following steps are performed: S21, sorting the table R and the table S according to the connection key B; S22, respectively calculating the maximum prefix sum of the table R and the table S on the connection key B to obtain table R1 and table S1; S23, using the table S1 to calculate the connectivity of the table R1 to obtain table R2, and using the table R1 to calculate the connectivity of the table S1 to obtain table S2; S24, performing concealment expansion on the table R2 and the table S2 respectively to obtain table R3 and table S3; S25, rearrange the table S3 to obtain table S4; S26. Merge the data of the table R3 and the table S4 in each server horizontally, and finally extract columns A, B, and C of the merged result as output results.
2. The distributed anonymity method in a trusted execution environment according to claim 1, characterized in that: When the primary key join algorithm is used, the following steps are performed: S31, assigning key values to the tuples in the table S; S32, sorting the connection key B in the partitions of the table R on each server, and selecting the first tuple with the same value in the connection key B as the representative tuple, marking the other tuples as inactive tuples, and recording the corresponding original server ID; S33, assigning key values to the representative tuples according to the connection key B, and randomly assigning the inactive tuples to other servers; S34, on each server, performing a join calculation on the representative tuple and the local tuple in the table S; S35, transmitting the data after the connection calculation to the original server, and distributing the result of the representative tuple to the corresponding inactive tuple to form a final result.
3. The distributed anonymity method in a trusted execution environment according to claim 2, characterized in that: In the S31, the key value allocation includes: S311. Define a tuple as a key-value pair, where each tuple is denoted by x = (k, v), where k is the key and v is the value; S312, calculating the key value assignment operation t = h(k) through the random oracle h, and determining the assignment target of the tuple; S313, according to the calculation result of h(k), tuples with the same t are aggregated to the same server; S314. Use public parameters to adjust and fill in the key value allocation process.
4. The distributed anonymity method in a trusted execution environment according to claim 1, characterized in that: In S22, calculating the maximum prefix sum of the table R includes: S221. For the table R, calculate the prefix sum of each element on the connection key B; S222, for the table R, add a column DR, and use the maximum prefix sum of the same prefix to refresh all columns DR corresponding to the elements with the same prefix, to obtain the table R1; Calculating the maximum prefix sum of the table S includes: S223. For the table S, calculate the prefix sum of each element on the connection key B; S224. For the table S, add a column DS, and use the maximum prefix sum of the same prefix to refresh the columns DS corresponding to all elements with the same prefix, to obtain the table S1.
5. The distributed anonymity method in a trusted execution environment according to claim 4, characterized in that: The prefix sum algorithm is as follows: E1, the prefix sum represents the cumulative number of times a prefix appears in the tuple; E2. At the beginning of the calculation, the cumulative number is set to 0; E3. Scan the tuples in sequence. When the same prefix is encountered, the cumulative number is increased by 1. E4. When a new prefix is encountered, the cumulative number is recalculated from 1.
6. The distributed anonymity method in a trusted execution environment according to claim 4, characterized in that: In S23, using the table S1 to calculate the connectivity of the table R1 includes: S231, delete column C of the table S1, and deduplicate the data to obtain a table S1 with the connection key B as the primary key column. B ; S232, according to the table R1 and the table S B , use the primary key join algorithm to get table R L ; S233, delete the table R L The data whose DS value is 0 in the table are deleted, and the DR column is deleted to obtain the table R2; Calculating the connectivity of the table S1 using the table R1 includes: S234, delete column A of the table R1, and deduplicate the data to obtain a table R1 with the connection key B as the primary key column. B ; S235, according to the table S1 and the table R B , use the primary key join algorithm to get table S L ; S236, delete the S L The data with a DR value of 0 are removed and the DS column is deleted to obtain the table S2.
7. The distributed anonymity method in a trusted execution environment according to claim 6, characterized in that: In S24, performing a concealment extension on the table R2 includes: Expand columns A and B according to the values of DS in table R2, and then remove column DS to obtain table R3; The concealed expansion of Table S2 includes: According to the values of DS in Table S2, columns B and C are expanded, and then column DR is removed to obtain Table S3.
8. The distributed anonymity method in a trusted execution environment according to claim 1 or 7, characterized in that: In S25, rearranging the table S3 includes: S251, add an I column to S3, the value of the I column is the prefix sum of the connection key B, and use a J column to store the global minimum sequence of the same prefix; S252: Calculate a global sorting column L based on the I column and the J column, perform a modulo operation on the global sorting column L to obtain a server sequence number column T, then delete the I column and the J column, perform hidden padding, and obtain S T ; S253, according to the server serial number column T, the table S T The data is distributed to corresponding servers, and the data is sorted in each server according to the server sequence number column L; S254, delete the table S T The L and T columns of Table S4 are obtained.
9. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the distributed anonymity method in a trusted execution environment according to any one of claims 1 to 8 can be implemented.
Citation Information
Patent Citations
Customer information storage method and system based on information retrieval efficiency
CN118363957A
Browser-oriented hidden query method and system and storage medium
CN118606539A
Privacy protection dynamic space keyword query method and device under multi-attribute cost constraint, electronic equipment and storage medium
CN118643055A
Anonymous rating structure for database
US20200382301A1
Anonymity algorithm-based data sharing privacy protection method
US20250053686A1