Big data hierarchical encryption filtering mechanism based on TEE, query optimization method, system, equipment and medium
By building a hierarchical encryption filtering mechanism in TEE, using technologies such as HMAC, stable Bloom filter and red-black tree index, the balance of privacy protection and performance efficiency in cloud big data query is solved, and efficient ciphertext query optimization is achieved.
Patent Information
- Application Number
- CN202510417365.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-04
AI Technical Summary
The existing cloud-based big data confidential computing solution cannot meet data privacy protection and query performance efficiency at the same time. The traditional Parquet format fails after data encryption, resulting in a significant decline in query performance.
A TEE-based hierarchical encryption filtering mechanism is adopted to build a hierarchical encryption filtering mechanism through multi-grain encryption strategies at partition level, file level, row group level and column level, including HMAC encrypted partition directory name, stable Blonde filter and red-black tree index, order-preserved encryption and encrypted Blonde filter to realize ciphertext query optimization.
Without leaking metadata, the query performance is improved, the TEE internal decryption calculation overhead is reduced, and the query efficiency is achieved close to plaintext calculations is solved, which solves the metadata leakage and performance bottleneck problems.
Smart Images

Figure CN120256466A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data privacy protection, and specifically relates to a big data layered encryption filtering mechanism and a query optimization method based on TEE. Background Art
[0002] In order to effectively protect data privacy, the current mainstream solution is to build a trusted computing base (TCB) through a trusted execution environment (TEE) and remote attestation technology, such as the Enclave verification mechanism based on Intel SGX (Software Guard Extensions), which can ensure the credibility of the computing process in a hardware isolation environment. However, when implementing data encryption protection, the existing big data TEE architecture destroys the structural characteristics of the data during the encryption process, resulting in the failure of core query optimization mechanisms such as predicate pushdown and partition pruning based on the partition information, column information, and statistical information (such as min / max values, data distribution, etc.) recorded in metadata in traditional data processing scenarios. This forces a full disk data scan, which seriously affects query performance.
[0003] Around the above problems, the current academic community has proposed a variety of solutions, such as: encrypting data through a fully homomorphic encryption algorithm, converting filter conditions into homomorphic operations, so that the query engine can perform calculations directly on the ciphertext without decryption, and theoretically completely avoiding the problem of predicate pushdown failure. However, the computational complexity of the fully homomorphic encryption algorithm is extremely high, resulting in a sharp decline in its performance in large-scale data analysis scenarios. Actual measurements show that even with the latest FHE acceleration framework (such as HElib), the query latency for processing TB-level data sets is still 4 orders of magnitude higher than plaintext calculations, making it difficult to apply in practice. In addition, transparent encryption is also one of the solutions. Its core is to encrypt data outside the TEE, but save some metadata information in plaintext, so that these plaintext metadata can still be used for optimization during query. However, this solution has serious security risks: cloud service providers can infer data distribution characteristics by monitoring the access pattern of metadata outside the TEE, and even indirectly reconstruct the original data through frequent queries on certain partitions.
[0004] To solve the above problems, researchers began to try to build a lightweight query optimization layer inside the TEE, supporting predicate pushdown by caching some min / max statistics in the Enclave [Eskandarian S, ZahariaM. Oblidb: Oblivious query processing for secure databases [J]. arXiv preprintarXiv: 1710.00458, 2017.]. This move avoids the metadata leakage problem faced by transparent encryption to a certain extent, reduces the redundant decryption calculation amount, and improves query efficiency. However, due to the memory limitation of SGX EPC, the cache hit rate is often less than 40% during large-scale queries, resulting in a large number of row groups having to be frequently swapped in and out of the Enclave, resulting in serious memory paging overhead. Actual measurements show that in big data scenarios, the query latency fluctuation of this solution can reach 2.8 times the baseline value, making it difficult to guarantee stable query performance.
[0005] Recent research, such as Shield Store, proposes to build an encrypted statistical summary tree in TEE to store bucket metadata of hash tables [Kim, T., Park, J., Woo, J., Jeon, S., & Huh, J. (2019, March). Shieldstore: Shielded in-memory key-value storage with sgx. In Proceedings of the Fourteenth EuroSys Conference 2019 (pp. 1-15).]. However, as the amount of data grows, the size of the metadata may exceed the EPC capacity; in addition, each read requires layer-by-layer hash verification. In large-scale data scenarios, the verification overhead accumulates into an additional computational burden. The Intel SGX SDK native encryption library only provides basic AES-GCM encryption and is not optimized for structured data. The encrypted Parque file metadata completely loses its semantics, forcing the query engine to abandon all optimization strategies.
[0006] In summary, existing cloud-based big data confidential computing solutions cannot simultaneously protect metadata privacy while maintaining performance efficiency similar to plaintext queries. Computing solutions that rely entirely on full disk decryption within TEE are limited by memory capacity and paging overhead, resulting in a significant reduction in query performance and making it difficult to support the needs of large-scale data analysis.
[0007] As a columnar storage format, Apache Parquet is widely used in big data analysis scenarios. It has become one of the mainstream storage formats for big data computing engines such as Spark, Hive, and Presto due to its advantages such as high compression rate, efficient query, and support for columnar pruning. The internal structure of a Parquet file consists of multiple row groups, each of which is stored in columns and contains column-level statistics and optional Bloom filters. This information is written to the metadata area Footer of the Parquet file to accelerate subsequent query operations. However, the traditional Parquet format is not optimized for efficient query of encrypted data. Especially after data encryption, query acceleration mechanisms such as column pruning, block filtering, and partition pruning will fail due to the loss of metadata, resulting in a significant decrease in query performance. Summary of the invention
[0008] In order to overcome the problems existing in the above-mentioned prior art, the purpose of the present invention is to disclose a big data hierarchical encryption filtering mechanism and query optimization method, system, device and medium based on TEE. On the basis of TEE's hardware isolation, combined with the hierarchical encryption strategy, a three-layer filtering mechanism at the partition level, file level and row group level is constructed, which realizes efficient query optimization of encrypted data while avoiding metadata leakage; in order to solve the problem of partition directory semantic leakage, the partition directory name is converted into HMAC encryption, which not only hides the partition semantic information, but also supports fast matching based on ciphertext hash values, solving the I / O problem caused by the failure of partition pruning in traditional encryption schemes. O performance bottleneck problem; in addition, in response to the residual risk of file-level column name fingerprints, an event-driven stable Bloom filter is designed, which actively eliminates the column name fingerprints of deleted files through a directional attenuation mechanism; at the same time, in order to solve the problem of query optimization with finer-grained filtering, the present invention integrates order-preserving encryption (OPE) and encrypted Bloom filter (EBF) at the row group level, and realizes ciphertext range index and data fingerprint matching inside the Parquet file, so that the query engine can screen candidate row groups without decryption, and finally only decrypts the target column on demand, limiting the computing overhead within the TEE security boundary, taking into account both privacy protection and query efficiency.
[0009] In order to achieve the above object, the present invention adopts the following technical solutions:
[0010] A big data hierarchical encryption filtering mechanism and query optimization method based on TEE, specifically comprising the following steps:
[0011] Step 1: When writing Parquet data to disk, use the custom Spark Parquet data source interface to perform hierarchical encryption on the data in TEE and generate a multi-level metadata index.
[0012] Step 2: Based on the data hierarchical encryption processing and the generated multi-level metadata index in Step 1, perform data query. During the query execution process, the Spark query engine first receives the SQL query request from the user, and then rewrites the original plaintext query predicate into a ciphertext predicate condition that can be executed in the ciphertext data domain. Then, the query execution process enters the hierarchical encryption filtering stage, and performs filtering operations at the partition level, file level, row group level, and column level in sequence to reduce the ciphertext decryption calculation amount inside the TEE.
[0013] The specific method of Step 1 is as follows:
[0014] Step 1.1 Processing of partition-level metadata:
[0015] In the data writing stage, customize the Spark Parquet data source interface, and use the HMAC encryption method to generate the partition directory name. The specific rule is:
[0016] E_Dir = {HMAC(K,Key) prefix}_{HMAC(K,Value) prefix}
[0017] where K is the key, Key and Value represent the partition key and partition value respectively, and prefix represents the first 8 bits of HMAC;
[0018] Step 1.2 Processing of file-level metadata:
[0019] In the data writing stage, customize the Spark Parquet data source interface to build a stable Bloom filter (SBF) and a red-black tree index for the partition where the data is written;
[0020] Step 1.2.1 Building a stable Bloom filter (SBF). The stable Bloom filter (SBF) is an array of counters used to globally and quickly filter whether a column name exists in the partition. The counter value range is [0, Max]. The building process includes:
[0021] Step 1.2.1a Generate an encrypted fingerprint for each column name ci in each Parquet file in the partition;
[0022] Step 1.2.1b Calculate the 20-bit CRC hash value of the column fingerprint;
[0023] Step 1.2.1c Use k hash functions for the CRC hash value to calculate k positions of the CRC hash value on the calculator, and update the bit values of the counter at these positions in the stable Bloom filter (SBF);
[0024] Step 1.2.2 Build a red - black tree index and insert the 20 - bit CRC hash value calculated in Step 1.2.1b into the leaf nodes of the red - black tree. The red - black tree index is divided into internal nodes and leaf nodes. The internal nodes store hash value ranges for fast data retrieval, and the leaf nodes store the specific CRC hash values and their corresponding Parquet file sets for precise matching of the files where column names are located.
[0025] After generating the file - level index during the data writing phase, write the metadata information including the stable Bloom filter (SBF), red - black tree index, and the number of partitioned Parquet files encrypted into the metadata file partition_meta.enc of each partition. The column names included in the partitioned Parquet files and the maximum and minimum metadata information of the columns within the row groups of the partitioned Parquet files are written into the Footer part of the corresponding partitioned Parquet files.
[0026] Step 1.3 Processing of row - group - level metadata:
[0027] Step 1.3.1 For SQL range queries involving predicates such as '>', '<', 'BETWEEN' in the user's query conditions, during the data writing phase, customize the Parquet data source interface of Spark. When writing each row group of the Parquet file, extract the maximum and minimum values of the data in each column of the row group and perform encryption processing using the Order - Preserving Encryption (OPE) algorithm. The encrypted statistical information will be written into the Footer of the corresponding Parquet file, enabling users to perform fast filtering and data skipping based on the ciphertext statistical information during subsequent SQL range queries, improving the query performance of encrypted data. The specific process is as follows:
[0028] Step 1.3.1a: Calculate the maximum and minimum values of each column in the row group.
[0029] Step 1.3.1b: Use the order - preserving encryption algorithm to encrypt the above - mentioned maximum and minimum values to ensure that the order of the values after order - preserving encryption is the same as the order of the plaintext before encryption.
[0030] Step 1.3.1c: Take the encrypted maximum and minimum values as encrypted statistical information and store the generated order - preserving encryption statistical information in the Footer part of the partitioned Parquet file.
[0031] Step 1.3.2 For the SQL equivalent query involving predicates such as "=" and "IN" in the user query conditions, customize the Parquet data source interface of Spark to construct an encrypted Bloom filter for each row group inside the Parquet file and embed it into the Footer of the Parquet file for use in subsequent query equivalent filtering optimization:
[0032] Step 1.3.2a Use the row group identifier (RowGroupID) as the unique identifier for each row group in the partitioned Parquet file, calculate the HMAC hash value for each column value in the row group to ensure that the stable Bloom filter is bound to a specific row group;
[0033] Step 1.3.2b Divide the obtained hash values into k sub-segments;
[0034] Step 1.3.2c Calculate k mapping positions and apply an encryption transformation;
[0035] Use the used row group identifier (RowGroupID) as metadata and encrypt and store it together with the encrypted Bloom filter parameter metadata in the Footer of the Parquet file for matching during query;
[0036] Step 1.4 Processing of column-level metadata:
[0037] Write the column offset, column data length encryption information of each column in the Parquet file into the Footer metadata of the partitioned Parquet file.
[0038] The specific method of Step 2 is as follows:
[0039] When performing partition filtering queries, the Spark query engine performs a ciphertext transformation on the plaintext predicate; specifically, the Spark query engine calculates the HMAC value of the plaintext query partition in the TEE in the same way as in Step 1.1 to generate the corresponding ciphertext predicate value; this ciphertext value will be used to match the encrypted storage directory HMAC value; if they are equal, the match is successful, and the encrypted data in this partition directory is selected by the Spark query engine into the subsequent query path, otherwise this partition is skipped to avoid unnecessary data loading;
[0040] Step 2.2 performs the file-level encryption filtering phase. The Spark query engine extracts the set of target column names of all plaintext predicates in the query condition, repeats the column fingerprint generation and hash calculation process in Step 1.2 for each column name to obtain k positions; then, reads the counter values at these positions in the stable Bloom filter: if the counter values at all k positions are greater than 0, it indicates that the column name may exist in the partition Parquet file, and the Parquet file of the candidate partition enters the red-black tree exact matching phase; if the counter value at any position is 0, it means that the column name definitely does not exist in the partition, and all partition Parquet files within this partition are directly skipped.
[0041] In the red-black tree exact matching phase, for the Parquet files within the partition screened by the stable Bloom filter (SBF), the Spark query engine performs exact matching in the red-black tree index; the red-black tree index stores the CRC hash value and its corresponding set of partition Parquet files. The Spark query engine repeats the CRC hash value calculation in Step 1.2 for the target column, and then traverses the obtained hash value from the root node to the leaf node; if a leaf node with an equal CRC hash value is found, the set of Parquet files stored in the leaf node is obtained as the candidate partition Parquet files and enters the row group-level screening; if not found, it means that the column name does not exist in the partition, and all Parquet files within this partition are directly skipped; the file-level encryption filtering phase reduces the data scan volume through the two-level index mechanism of the stable Bloom filter and the red-black tree, and only retains the candidate partition record Parquet files containing the target column.
[0042] Step 2.3 performs the row group-level encryption filtering phase. The Spark query engine combines order-preserving encryption and the encrypted stable Bloom filter (SBF) to further screen the candidate row groups inside the partition Parquet files, which is divided into range queries and equality queries according to different types of user query predicates.
[0043] Step 2.3.1 For range queries: SQL range queries in the user query condition involving predicates such as '>', '<', 'BETWEEN'.
[0044] The Spark query engine performs a ciphertext transformation on the plaintext predicates; transmits the query range boundary values in the plaintext predicates to the inside of the TEE, encrypts the boundary values using the same key K and the order-preserving encryption algorithm as in step 1.3, and the generated ciphertext predicates contain the encrypted boundary values. The ciphertext boundary values generated by order-preserving encryption are in the same numerical order as the corresponding plaintext; compares the ciphertext maximum and minimum values of each column within the row groups stored inside the partitioned Parquet files; determines and filters out the candidate row groups that have an intersection with the query encrypted boundary value range and the ciphertext maximum and minimum values of each column, and includes them in the subsequent decryption and exact matching process; for the row groups that have no intersection with the query encrypted boundary value range and the ciphertext maximum and minimum values of each column, directly skip them;
[0045] Step 2.3.2 For equality queries: SQL equality queries in the user query conditions involving predicates such as '=' and 'IN';
[0046] The Spark query engine performs a ciphertext transformation on the plaintext predicates, transmits the query column names and query column values in the plaintext predicates to the inside of the TEE, calculates the query hash in the same way as step 1.3, and generates k encrypted positions; subsequently, checks the stable Bloom filter (SBF) of each row group in the candidate partitioned Parquet files. If the sites at all k positions are 1, then this row group may contain the query value and enters the decryption process; if any position site is 0, then the row group does not contain the query value and is directly skipped;
[0047] Step 2.4, after filtering at the partition, file, and row group levels, only the target columns that meet the query conditions need to be decrypted; reads the Footer of the Parquet file to obtain the offsets and column data ciphertext lengths of each query's target column within the row group; decrypts the target column data using the key K inside the TEE and performs the final query calculation.
[0048] A big data hierarchical encryption filtering mechanism and query optimization system based on TEE, including an encryption metadata generation module, a partition-level encryption index module, a file-level stable Bloom filter (SBF) and red-black tree, a row group-level encryption filtering module, and a column-level on-demand decryption mechanism; among them:
[0049] The encryption metadata generation module is used for step 1, and by synchronously executing during the process of writing data into the partitioned Parquet files, it realizes that the encryption metadata is embedded into the Footer metadata structure of the partitioned Parquet files during the data writing stage;
[0050] The partition-level encryption index module is used for step 2.1, and through the hash prefix and OPE hybrid encryption mechanism, it generates an encrypted partition directory name and performs ciphertext matching during query, realizing efficient partition pruning and data distribution privacy protection;
[0051] File-level Stable Bloom Filter (SBF) and Red-Black Tree, for Step 2.2. By combining the stable Bloom filter and the Red-Black tree index, generate fingerprints for column names and perform two-level filtering to quickly locate candidate partition Parquet files containing the target column and reduce false positive errors;
[0052] Row-group level encryption filtering module, for Step 2.3. By analyzing the SQL query predicate, determine whether it contains an equality query or a range query; if the query predicate contains a range query, extract the encrypted statistical values of each column within the row group from the Footer of the partition Parquet file and compare them with the order-preserving encrypted query boundary values to determine whether there is an intersection between the ciphertext value range and the query boundary range; if the query predicate contains an equality query, determine whether there may be a matching value by checking whether the corresponding k positions in the encrypted Bloom filter (SBF) are hit, so as to filter out the row groups that match the query conditions in the ciphertext environment;
[0053] Column-level on-demand decryption mechanism, for Step 2.4. By reading the offset and length of the target column from the Footer of the partition Parquet file and decrypting using the key K within the TEE, achieve decrypting only the target column data and completing the final query calculation.
[0054] A big data hierarchical encryption filtering mechanism and query optimization device based on TEE, including:
[0055] A memory for storing computer programs;
[0056] A processor for, when executing the computer program, implementing the big data hierarchical encryption filtering mechanism and query optimization method of any one of Steps 1 to 2 of the TEE.
[0057] A computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, can implement the big data hierarchical encryption filtering mechanism and query optimization based on the big data hierarchical encryption filtering mechanism and query optimization method of the TEE.
[0058] Compared with the prior art, the advantages of the present invention are as follows:
[0059] 1. By combining TEE with a multi-granularity encryption strategy, the present invention constructs an efficient screening mechanism at multiple data granularities of the partition level, file level, row-group level, and column level, avoiding the problem of query optimization failure caused by traditional encryption schemes, while reducing the decryption calculation overhead inside the TEE and improving the query performance.
[0060] 2. The present invention performs HMAC encryption conversion on the partition directory names, which not only hides the partition semantic information but also supports fast matching based on the ciphertext hash value, solving the security risks brought by metadata exposure in the transparent encryption scheme.
[0061] 3. In view of the security risks brought by the file-level column name data residue, the present invention designs an event-driven stable Bloom filter to implement column existence verification, and at the same time ensures that the column name fingerprints of the deleted Parquet files can be eliminated in time.
[0062] 4. The present invention integrates order-preserving encryption and encrypted Bloom filter, making optimization strategies such as predicate pushdown, range query, and equality query still available in the ciphertext state, avoiding the problem of excessive computational overhead caused by full-disk decryption in the traditional TEE query scheme, and at the same time realizing efficient range query and equality query optimization, making the query performance more stable.
[0063] 5. Compared with the existing cloud big data privacy protection scheme based on TEE, the present invention breaks through the limitations of the traditional method in query optimization. Existing schemes usually destroy the optimization ability of the query engine due to data encryption, resulting in the failure of key mechanisms such as partition pruning and predicate pushdown, thus triggering the performance bottleneck of full-disk scanning. The present invention realizes lightweight and secure maintenance of metadata through an innovative encrypted metadata index and hierarchical filtering mechanism: the partition-level HMAC index only needs to store one-way hash values, and the file-level stable Bloom filter actively compresses the fingerprint space through a directional attenuation mechanism; while ensuring data privacy, the query optimization ability is restored, and the query performance close to plaintext calculation is achieved.
[0064] In summary, the present invention uses encrypted metadata to ensure that even if the cloud service provider or external attacker obtains the metadata, they cannot infer the true distribution of the data without leaking the metadata, completely avoiding the metadata leakage risk of the transparent encryption scheme; reducing the decryption computational overhead inside the TEE, thereby improving the query efficiency; in the query execution stage, a hierarchical pre-filtering mechanism is also introduced to gradually screen the data set at the partition level, file level, and row group level based on the encrypted metadata, and only decrypt the column-level data required in the final result as needed, thereby greatly reducing the decryption computation amount inside the TEE. This strategy avoids the EPC memory pressure caused by full-disk decryption in the traditional TEE scheme, enables the Parquet data query to be efficiently executed with low memory occupancy, and finally achieves the balance between privacy protection and query optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 is the overall structure diagram of the present invention.
[0066] Figure 2 is the flowchart of the multi-layer filtering query optimization for big data confidential computing of the present invention.
[0067] Figure 3 This is the construction process diagram of the stable Bloom filter in the embodiments of the present invention.
[0068] Figure 4 This is the structural design diagram of the red - black tree in the embodiments of the present invention. Detailed implementation manners
[0069] The present invention will be further described in detail below with reference to the accompanying drawings.
[0070] See Figure 2 , a big - data hierarchical encryption and filtering mechanism and query optimization method based on TEE, specifically including the following steps:
[0071] Step 1, in the data writing stage, the custom Spark Parquet data source interface internally realizes the secure storage of data and the preparation for query optimization through the following hierarchical encryption and index generation steps:
[0072] Step 1.1 Processing of partition - level metadata:
[0073] Before writing data into the Parquet file within the Hadoop partition, the custom Spark Parquet data source interface constructs the partition directory name in the HMAC encryption mode for this Hadoop partition. The specific rules are as follows:
[0074] E_Dir = {HMAC(K, Key) prefix}_{HMAC(K, Value) prefix}
[0075] Where K is the encryption key, and HMAC(K, Key) prefix represents taking the first 8 - bit hash prefix after encrypting the partition key name through the HMAC algorithm, and HMAC(K, Value) prefix represents taking the first 8 - bit hash prefix after encrypting the partition key value through the HMAC algorithm. These encrypted partition directory names will be stored in the system for partition matching during subsequent queries.
[0076] Step 1.2 Processing of file - level metadata
[0077] In the stage of writing data into the partition Parquet file, the custom SparkParquet data source interface first constructs the index structure. For column - name filtering at the partition Parquet file level, a stable Bloom filter (SBF) needs to be constructed for each partition where the file is located. This filter is an array of counters, and the value range of each counter is [0, Max], where Max is the preset maximum count value.
[0078] Step 1.2.1 SeeFigure 3 , the process of constructing a Stable Bloom Filter (SBF) is as follows:
[0079] ① Generate a unique fingerprint for each column name in the partitioned Parquet file. The fingerprint generation mainly calculates the hash value of the column name using the SHA3-256 algorithm, and then processes this hash value using the encryption key K through the HMAC algorithm to obtain the encrypted fingerprint of the column name. This process can be expressed as:
[0080] FP(c i ) = HMAC(K, SHA3-256(c i ))
[0081] where c i represents the column name, and FP represents the fingerprint of the column name.
[0082] ② After generating the fingerprint of the partitioned Parquet file column name, calculate its 20-bit hash value through the CRC32 algorithm, denoted as CRC, with the aim of reducing the storage space of the red-black tree:
[0083] fp_crc = CRC32(FP(c i )) & 0xFFFFF
[0084] Then, the construction method of the Hash function in the Stable Bloom Filter (SBF) is constructed using two basic hash functions H1 and H2 to generate k mutually independent hash functions. Each hash function h i is defined as:
[0085] h i (fp_crc) = (H1(fp_crc) + i · H2(fp_crc)) mod m
[0086] where the value range of i is [1, k], and m is the size of the counter array of the Stable Bloom Filter (SBF).
[0087] ③ Then, calculate k positions based on these k hash functions and update the counter values of these k positions in the Stable Bloom Filter. The update rule is:
[0088] C[hi] = min(Max, C[hi] + 1)
[0089] where i is the position index obtained by the hash function calculation, and the counter upper limit is set to Max to prevent counter overflow.
[0090] Step 1.2.2 Index Construction In addition to constructing a stable Bloom filter for global column name existence screening within the partition, a red-black tree also needs to be constructed as the second-level exact index for precisely matching the set of partition Parquet files where the column name exists within the current partition. A red-black tree is a self-balancing binary search tree with good search, insertion, and deletion performance, and its time complexity is O(log n). A red-black tree consists of internal nodes and leaf nodes. See Figure 4 , and the structure design of the red-black tree is as follows:
[0091] Internal nodes store the hash value range interval [min_hash, max_hash] to guide the search process; leaf nodes store the specific CRC hash value, its corresponding set of partition Parquet files, and the position mapping.
[0092] During the process of inserting a column name in the construction of the red-black tree index, start from the root node first. According to the comparison result between the CRC hash value and the node interval, decide whether to move to the left subtree or the right subtree until the appropriate leaf node position is found. If the same CRC hash value already exists at this position, add the partition Parquet file name to the corresponding set of partition Parquet files and record the hash position list. If the same CRC hash value does not exist, create a new leaf node, set its hash value to the CRC hash value, the file set to the set containing the current file name, and insert it into the red-black tree. To ensure the balance of the red-black tree, after inserting a new node, the red-black tree will perform color adjustment and rotation operations according to its own property rules to ensure the height balance of the tree, thereby ensuring the efficiency of query operations.
[0093] For the deletion of partition Parquet files, the system locates the set of partition Parquet file positions fp_crc_pos through the leaf nodes of the red-black tree and performs local attenuation on the corresponding counters in the stable Bloom filter. To avoid directly clearing the counters and affecting query accuracy, an exponential probability attenuation is introduced:
[0094] P = 1 - e^(-C[hp] / Max)
[0095] where C[hp] is the bit value of the current bit of the counter and Max is the counter length. If the random number t < P, then C[hp] = max(0, C[hp] + 1). Subsequently, remove the association of the partition Parquet file name from the red-black tree. If the set of partition Parquet files is empty, delete the leaf node and rebalance the tree. This method effectively solves the defect that traditional Bloom filters cannot delete elements and ensures efficiency and accuracy.
[0096] Step 1.3 Processing of Row Group-Level Metadata
[0097] Step 1.3.1 For SQL range queries involving predicates such as '>', '<', and 'BETWEEN' in the user query conditions, during the data writing phase, customize the Parquet data source interface of Spark. When writing each Parquet row group, calculate the maximum and minimum values of the data in each column of the row group, and perform encryption processing using the Order-Preserving Encryption (OPE) algorithm. Generate the encrypted maximum and minimum statistical values enc_min and enc_max to ensure that the order of the plaintext numerical values in the query conditions after encryption is the same as before encryption, and store the generated order-preserving encryption statistical information in the Footer part of the partitioned Parquet file;
[0098] Step 1.3.2 For SQL equality queries involving predicates such as '=' and 'IN' in the user query conditions, customize the Parquet data source interface of Spark to construct an encrypted Bloom filter for each row group inside the Parquet file and embed it in the Footer of the partitioned Parquet file for use in subsequent equality filtering optimization; the specific steps include:
[0099] ① Assign a unique row group identifier RowGroupID to each row group;
[0100] ② For each column value val in the row group, combine the column identifier col and RowGroup, and calculate the hash value using the key K: h = HMAC(K, col||val||BlockID)
[0101] ③ Evenly divide the generated fixed-length binary hash value h into k sub-segments, and each sub-segment contains a continuous bit sequence to form k mapping positions;
[0102] ③ Apply an encryption transformation to each mapping position: enc_pos = AES_Encrypt(K, pos) ⊕ BlockID
[0103] ④ In the Bloom filter bit array, set the bit positions corresponding to all the calculated encrypted positions to 1, and store the finally generated encrypted Bloom filter as metadata in the Footer.
[0104] Step 1.4 Column-level metadata generation
[0105] Generate a storage location index for each column, recording the offset Offset of the column within the row group j and the data length Length j。These index information is used to accurately locate the ciphertext data of the target column during query, avoiding loading irrelevant columns or the entire row group. Use the key K to encrypt the storage location index of the column, generating the encrypted index values E(Offset j ) and E(Length j ).
[0106] Step 2, during the Spark query execution process, the Spark query engine first receives the SQL query statements submitted by the user. These statements may contain various query conditions, such as data table selection, column filtering, filtering conditions in the WHERE clause, JOIN operations, etc. The optimization engine needs to perform lexical analysis and syntactic analysis on these SQL statements, and calculate the plaintext predicates including operators such as "=", ">", "<", field names, and corresponding values into ciphertext predicates through HMAC or order-preserving encryption according to their predicate types belonging to equality or range query types. This predicate structure facilitates subsequent query matching.
[0107] Then the Spark query engine enters the hierarchical encryption filtering stage, and sequentially performs partition-level, file-level, row group-level, and column-level filtering operations to reduce the ciphertext decryption calculation amount inside the TEE:
[0108] Step 2.1 Partition-level encryption filtering stage, the Spark query engine performs partition filtering through the HMAC directory index; the Spark query engine calculates the HMAC value of the ciphertext predicate in the same steps as 1.1 for the plaintext predicate according to the query conditions inside the TEE, obtaining the encrypted partition directory name of the query partition, and matching it in the ciphertext directory list; for example, for the user query "SELECT id, name FROM table WHERE age BETWEEN 30 AND 50 AND gender = '0' AND dt = '2025-03-01';", where dt is the partition column, then rewrite the partition filtering predicate dt = '2025-03-01' as the equivalent ciphertext predicate condition dt = HMAC(K, '2025-03-01') that can be executed in the ciphertext data domain. Here K is the encryption key. If the Hash encrypted partition directory name of the query partition obtained after calculation is equal to the Hadoop encrypted directory name stored on the system, it is considered a successful match, and the encrypted data of this partition directory is selected by Spark into the subsequent query path; otherwise, skip this partition to avoid unnecessary data loading;
[0109] Step 2.2 File-level Encryption Filtering Phase: The Spark query engine uses a stable Bloom filter and a red-black tree to determine whether the candidate partition Parquet file contains the query target column. This method adopts a two-level index structure: the first level uses a stable Bloom filter for fast filtering, and the second level uses a red-black tree for exact matching. This combination makes full use of the efficient filtering ability of the Bloom filter and the exact search ability of the red-black tree, effectively improving the query performance of encrypted partition Parquet files in a big data environment.
[0110] When the user submits a query, the Spark query engine first extracts the set of all column names involved in the plaintext query predicate. For example, for "SELECT id, name FROM table WHERE age BETWEEN 30 AND 50 AND gender = '0' AND dt = '2025-03-01';", the extracted set of column names is: [id, name, age, gender], and the same fingerprint generation process as in Step 1.2 is performed for each column name. For each column name fingerprint, the Spark query engine calculates its corresponding k hash positions and checks the counter values at these k positions in the stable Bloom filter. If the counter values at all positions are greater than 0, the column name may exist in the partition and needs to enter the second stage for exact matching; if the counter value at one position is 0, it can be determined that the column name does not exist in the partition, and the subsequent processing of this column name can be directly skipped. The characteristics of the Bloom filter guarantee no false negatives, that is, if the filter indicates that a column name does not exist, then the column name does not exist in any partition Parquet file of this partition.
[0111] Step 2.3 For the column names that are preliminarily screened by the stable Bloom filter and exist in the partition, the Spark query engine further performs an exact matching query in the red-black tree to find the set of partition Parquet files where they are located. Calculate the CRC hash value of the query column name, and then start from the root node, compare the CRC hash value of the column name with the hash interval of the internal node, and decide whether to move to the left subtree or the right subtree until the leaf node is found. If the hash value of this leaf node is the same as the CRC hash value of the column name, obtain the set of partition Parquet files corresponding to this leaf node. These partition Parquet files are the candidate partition Parquet files that may contain the target column name; if there is no matching hash value in the red-black tree, it indicates that the column name does not exist in the partition, and the Spark query engine can directly skip the subsequent processing of this column name. The exact matching mechanism of the red-black tree can eliminate the possible false positives of the stable Bloom filter and further improve the query accuracy.
[0112] For the set of candidate partition Parquet files obtained through exact matching in the red - black tree, the Spark query engine finally needs to perform metadata verification to ensure the accuracy of the query results. The Spark query engine decrypts the partition - level metadata file partition_meta.enc and verifies whether the target column names actually exist in the candidate partition Parquet files. If the target column is included in the candidate partition Parquet file, the partition Parquet file is added to the final result set; if the target column is not included in the candidate partition Parquet file, the partition Parquet file is excluded from the result set. Excluding the partition Parquet file from the result set is equivalent to directly skipping the partition Parquet file, reducing the amount of data scanned; metadata verification can completely eliminate the false positives that may occur in the first two stages and ensure the accuracy of the query results.
[0113] Step 2.3 In the row - group - level encryption and filtering stage, the Spark query engine combines order - preserving encryption and encrypted Bloom filters to optimize range queries and equality queries respectively during query execution:
[0114] Step 2.3.1 For range queries, the Spark query engine transmits the upper and lower bounds of the plaintext predicates in the query condition to the inside of the TEE. Inside the TEE, the query engine uses the same key key and order - preserving encryption algorithm as in Step 1.3 to encrypt the boundary values, generating the encrypted boundary sets enc_lower = OPE(K, lower) and enc_upper = OPE(K, upper), where OPE(K, x) represents the ciphertext generated by applying order - preserving encryption to the boundary value x under the key K. Ensure that the original size relationship is still maintained after encryption. For example, for "SELECT id, name FROM table WHERE age BETWEEN 30 AND 50 AND gender = '0' AND dt = '2025 - 03 - 01';", the plaintext predicate "age BETWEEN 30 AND 50" is extracted and rewritten as the ciphertext - executable predicate: age ≥ OPE(K, 30) AND age ≤ OPE(K, 50).
[0115] Subsequently, the query engine reads the encrypted metadata statistics of each row group stored in the Footer during the data writing phase, namely the encrypted minimum value enc_min and the maximum value enc_max, and makes the following judgments: If enc_lower > enc_max, it means that all values in this row group are less than the query lower bound, and it is directly skipped; if enc_upper < enc_min, it means that all values in this row group are greater than the query upper bound, and it is directly skipped; if enc_lower ≤ enc_max and enc_upper ≥ enc_min, it means that the encrypted range of this row group may overlap with the query range, and it is filtered as a candidate row group.
[0116] For example, the query ciphertext boundaries exemplified here are OPE(K, 30) and OPE(K, 50); assume that the row group in the Footer contains the order-preserving encryption statistics of column age, the minimum value is enc_min(age), and the maximum value is enc_max(age). Then the judgment conditions are as follows: If enc_max(age) ≥ OPE(K, 30) and enc_min(age) ≤ OPE(K, 50). This row group may hit the query condition and needs to be included in the candidate row group set. Otherwise, this row group cannot hit the query and is skipped.
[0117] In step 2.3.2 for equality queries, when the user initiates a query, the Spark query engine passes the column name col and the query value val in the query condition into the TEE internally. For example, "SELECT id, name FROM table WHERE age BETWEEN 30 AND 50 AND gender = '0' AND dt = '2025-03-01';", then the plaintext predicate "gender = 0" is extracted and rewritten as a ciphertext executable predicate: gender = HMAC(K, gender||0||BlockID). Inside the TEE, similar to step 1.3, the query engine performs the following steps for each row group to be checked:
[0118] Calculate the hash-based message authentication code query_mac = HMAC(K, col||val||BlockID) using the key k;
[0119] Evenly divide the fixed-length binary string of query_mac into k sub-segments, and each sub-segment contains a continuous bit sequence to form k original positions {pos_1, pos_2,..., pos_k}.
[0120] Apply the same encryption transformation to each position pos_i to obtain the encrypted position:
[0121] enc_pos_i = AES_Encrypt(key, pos_i) ⊕ BlockID
[0122] The Spark query engine checks the bit values of these k positions enc_pos_1,..., enc_pos_k in the encrypted Bloom filter (SBF): If all bits are 1, the row group may contain the query value and is included in the subsequent decryption process; if at least one bit is 0, the row group definitely does not contain the query value and is directly skipped without decryption.
[0123] Step 2.4, after partition, file, and row group level filtering, only the target columns that meet the query conditions need to be decrypted; the goal of column-level filtering is to accurately locate the set of target columns {C1, C2,..., C m}, to avoid loading and decrypting irrelevant data columns, thereby reducing waste of computing and storage resources. During the query execution phase, the Spark query engine first reads the specific storage locations of each target column C in the Footer stored during the data writing phase in Step 1.4 j within the candidate row group, including the offset and row group length. Based on this information, the query engine directly reads the ciphertext data of the target column from the disk without loading the entire row group or the complete content of the row group. After reading the ciphertext data of the target column; then, the Spark query engine decrypts it using the key K inside the TEE and performs the final query calculation.
[0124] See Figure 1 , a big data hierarchical encryption filtering mechanism and query optimization system based on TEE, including an encrypted metadata generation module, a partition-level encryption index module, a file-level stable Bloom filter, a row group-level encryption filtering module, and a column-level on-demand decryption mechanism; among them:
[0125] The encrypted metadata generation module is used for Step 1. When writing data to a partitioned Parquet file through a custom Spark Parquet data source interface, this module processes the data based on a hierarchical filtering algorithm to generate multi-level encrypted metadata information. Including: the encrypted partition directory name, the column name index at the file level, the encrypted statistical information at the row group level, and the offset description at the column level. These metadata will be written into the metadata file partition_meta.enc in the partition directory and the Footer area respectively, providing an efficient filtering basis for subsequent query operations.
[0126] The partition-level encryption index module is used in Step 2.1. Through the Spark query engine within the TEE, using the key K and the HMAC algorithm, the encrypted directory name is calculated according to the partition filtering conditions in the query and matched with the stored list of encrypted text directories, achieving partition-level screening. The partitions with successful matches enter the subsequent query, and those with failures are skipped.
[0127] The file-level stable Bloom filter is used in Step 2.2. Through the Spark query engine within the TEE, using the stable Bloom filter and the red-black tree index, the fingerprint of the query column name is calculated, achieving the screening of candidate partition Parquet files. The Parquet files containing the target columns enter the next step, otherwise they are skipped, reducing the scanning volume.
[0128] The row group-level encryption filtering module is used in Step 2.3. Through the Spark query engine within the TEE, combining order-preserving encryption and the encrypted Bloom filter, row group screening is achieved. The row groups that meet the conditions are retained, and those with no intersection or mismatch are skipped.
[0129] The column-level on-demand decryption mechanism is used in Step 2.4. Through the Spark query engine within the TEE, the column offset and length information in the encrypted Footer are identified and read, achieving precise positioning and on-demand decryption.
[0130] The effects of the present invention can be further verified through the following experiments:
[0131] I. Experimental Conditions
[0132] In this experiment, a test cluster consisting of 5 physical servers was built, with 1 being the Spark Driver node and 4 being the Spark Executor nodes. In terms of hardware configuration, each server is equipped with an Intel Xeon Silver4410Y processor (main frequency 2.0GHz, 12 cores), 32GB of memory, and a hybrid storage solution of 512GB SSD and 1TB HDD, and supports Intel SGX2 technology to enhance security. The network environment uses Gigabit Ethernet to ensure the efficiency and stability of data transmission between nodes.
[0133] II. Experimental Settings
[0134] The experiment uses the TPC-H standard benchmark dataset (scales of 50GB, 100GB, and 200GB respectively), simulating the transaction scenario between suppliers and purchasers, including 8 tables and 22 analytical queries. The data generates metadata and indexes through the hierarchical encryption filtering engine of the present invention. Three query schemes are tested and compared: plaintext query, ciphertext query, and the hierarchical encryption filtering scheme proposed by the present invention. The index construction efficiency, query execution time, and data scanning volume are mainly evaluated, and Q6 and Q17 are selected as representative queries.
[0135] III. Experimental Process
[0136] 1. Metadata Construction: On a 50GB dataset, construct a partition-level HMAC directory, a file-level Bloom filter, a row-group-level encrypted Bloom filter, and an order-preserving encrypted statistical value and column-level metadata respectively, and record the construction time consumed at each level.
[0137] 2. Query Testing: Execute Q6 and Q17 queries on 50GB, 100GB, and 200GB datasets, and measure the execution time of the three schemes. Among them, Q6 is a single-table range query, and Q17 is a complex query scenario involving table joins, equality filtering, range filtering, and aggregate subqueries.
[0138] IV. Experimental Analysis
[0139] The hierarchical encryption and filtering scheme of the present invention shows significant advantages in performance optimization, and the specific analysis is as follows:
[0140] 1. Metadata Construction Efficiency:
[0141] Table 1 Metadata Construction Time (50GB Dataset)
[0142]
[0143]
[0144] It can be observed from the table that the partition-level index construction is the fastest, and the HMAC index only takes 15.5 seconds. This is mainly because the number of partitions is relatively small, and the computational complexity of the HMAC hash operation is relatively low.
[0145] In contrast, the file-level index needs to construct a stable filter and a red-black tree for data files. Although the Bloom filter itself has high computational efficiency, a large number of Parquet files still need to be processed. At the same time, adding the time required for balancing the red-black tree insertion nodes, it reaches 68.3 seconds. The row-group-level index has a longer construction time due to finer operation granularity. Among them, the encrypted Bloom filter and the OPE statistical information index require 92.2 seconds. The column-level index only needs to record the Parquet file offset and has a shorter construction time. The total construction time of the complete hierarchical encryption index system is 220.1 seconds. Although the construction time seems relatively long, considering that the index construction belongs to a one-time offline process and can support a large number of subsequent query operations, overall, the construction overhead is completely acceptable.
[0146] 2. Query Performance Improvement:
[0147] Table 2 shows the execution time of Q6 query. Under the scales of 50GB, 100GB, and 200GB, the proposed solution of the present invention is 10.16s, 820ms, and 1550ms respectively, which is 35.6%-37.4% higher than that of the encrypted query (118ms, 363ms, and 724ms). The performance of Q17 query (Table 3) is improved by 39.8%-40.7%.
[0148] Table 2 TPC-H Q6 Query Execution Time (Unit: Millisecond)
[0149]
[0150] Table 3 Comparison of TPC-H Q17 Query Execution Time (Unit: Second)
[0151]
[0152] The experimental results show that the hierarchical encryption and filtering solution proposed in the present invention achieves significant performance optimization. The average performance is improved by 36%, and the average improvement of Q17 query is 40%, which is far better than the encrypted query. This achievement fully verifies that the multi-layer index mechanism can effectively reduce the data scanning volume during the query process, thereby significantly reducing the query latency. It should be noted that although there is still a certain performance loss for the encrypted query compared with the plaintext query, the present invention effectively narrows the gap between security and performance while ensuring data security, achieving an optimal balance between the two.
Claims
1. A big data hierarchical encryption and filtering mechanism and query optimization method based on TEE, characterized in that Specifically, it includes the following steps: Step 1, during the process of writing Parquet data to disk, using a custom Spark Parquet data source interface, perform hierarchical encryption processing on the data in the TEE and generate multi-level metadata indexes; Step 2, based on the data hierarchical encryption processing and the generated multi-level metadata indexes in Step 1, perform data query. During the query execution process, the Spark query engine first receives the user's SQL query request, and then rewrites the original plaintext query predicate into a ciphertext predicate condition that can be executed in the ciphertext data domain; then, the query execution process enters the hierarchical encryption filtering stage, and sequentially performs partition-level, file-level, row group-level, and column-level filtering operations to reduce the ciphertext decryption calculation amount inside the TEE.
2. The big data hierarchical encryption and filtering mechanism and query optimization method based on TEE according to claim 1, characterized in that The specific method of Step 1 is as follows: Step 1.1 Processing of partition-level metadata: During the data writing stage, customize the Spark Parquet data source interface and generate a partition directory name using the HMAC encryption method. The specific rule is: E_Dir = {HMAC(K, Key) prefix}_{HMAC(K, Value) prefix} Among them, K is the key, Key and Value represent the partition key and partition value respectively, and prefix represents the first 8 bits of the HMAC; Step 1.2 Processing of file-level metadata: During the data writing stage, customize the Spark Parquet data source interface to build a stable Bloom filter (SBF) and a red-black tree index for the partition where the data is written; Step 1.2.1 Building a stable Bloom filter (SBF). The stable Bloom filter (SBF) is an array of counters used to globally and quickly filter whether a column name exists in the partition. The counter value range is [0, Max]. The building process includes: Step 1.2.1a Generate an encrypted fingerprint for each column name ci of each Parquet file in the partition; Step 1.2.1b Calculate the 20-bit CRC hash value of the column fingerprint; Step 1.2.1c Use k hash functions for the CRC hash value to calculate k positions of the CRC hash value on the counter, and update the bit values of the counter at these positions in the stable Bloom filter (SBF); Step 1.2.2 Build a red-black tree index. Insert the 20-bit CRC hash value calculated in Step 1.2.1b into the leaf node of the red-black tree; the red-black tree index is divided into internal nodes and leaf nodes. The internal nodes store the hash value range for fast data retrieval, and the leaf nodes store the specific CRC hash value and its corresponding Parquet file set for accurate matching of the file where the column name is located; After the file-level index is generated during the data writing stage, write the metadata information including the stable Bloom filter (SBF), red-black tree index, and the number of partition Parquet files into the metadata file partition_meta.enc of each partition in an encrypted manner. The column names included in the partition Parquet file and the maximum and minimum metadata information of the columns in the row group of the partition Parquet file are written into the Footer part of the corresponding partition Parquet file; Step 1.3 Processing of row group-level metadata: Step 1.3.1 For the SQL range query involving predicates such as '>', '<', 'BETWEEN' in the user query conditions, during the data writing stage, customize the Parquet data source interface of Spark. When writing each row group of the Parquet file, extract the maximum and minimum values of each column in the row group, and use the Order-Preserving Encryption (OPE) algorithm for encryption processing. The encrypted statistical information will be written into the Footer of the corresponding Parquet file, so that when the user performs a range query in SQL later, it can perform fast filtering and data skipping based on the ciphertext statistical information, improving the query performance of encrypted data. The specific process is as follows: Step 1.3.1a: Calculate the maximum and minimum values of each column in the row group; Step 1.3.1b: Use the order-preserving encryption algorithm to encrypt the above maximum and minimum values to ensure that the order of the values after order-preserving encryption is the same as the order of the plaintext before encryption; Step 1.3.1c: Take the encrypted maximum and minimum values as encrypted statistical information, and store the generated order-preserving encryption statistical information in the Footer part of the partitioned Parquet file; Step 1.3.2 For the SQL equality query involving predicates such as "=", "IN" in the user query conditions, customize the Parquet data source interface of Spark to build an encrypted Bloom filter for each row group inside the Parquet file and embed it into the Footer of the Parquet file for use in equality filtering optimization during subsequent queries: Step 1.3.2a Use the row group identifier (RowGroupID) as the unique identifier for each row group in the partitioned Parquet file, and calculate the HMAC hash value for each column value in the row group to ensure that the stable Bloom filter is bound to a specific row group; Step 1.3.2b Divide the obtained hash values into k sub-segments; Step 1.3.2c Calculate k mapping positions and apply encryption transformation; Use the used row group identifier (RowGroupID) as metadata and encrypt and store it together with the metadata of the encrypted Bloom filter parameters in the Footer of the Parquet file for matching during queries; Step 1.4 Processing of column-level metadata: Write the column offset, column data length encryption information of each column in the Parquet file into the Footer metadata of the partitioned Parquet file.
3. The big data hierarchical encryption filtering mechanism and query optimization method based on TEE according to claim 1, characterized in that The specific method of the said Step 2 is: Step 2.1 When performing a partition filtering query, the Spark query engine performs a ciphertext transformation on the plaintext predicate. Specifically, the Spark query engine calculates the HMAC value of the plaintext query partition in the TEE in the same way as in Step 1.1 to generate the corresponding ciphertext predicate value. This ciphertext value will be used to match the HMAC value of the encrypted storage directory; If they are equal, the match is successful, and the encrypted data in the partition directory is selected by the Spark query engine for the subsequent query path; otherwise, the partition is skipped to avoid unnecessary data loading; In step 2.2, during the file-level encryption filtering phase, the Spark query engine extracts the set of target column names of all plaintext predicates in the query conditions, and repeats the column fingerprint generation and hash calculation processes in step 1.2 for each column name to obtain k positions; then, reads the counter values at these positions in the stable Bloom filter: if the counter values at all k positions are greater than 0, the column name may exist in the partition Parquet file, and the Parquet file of the candidate partition enters the red-black tree exact matching phase; if the counter value at any position is 0, the column name definitely does not exist in the partition, and all partition Parquet files within the partition are directly skipped; In the red-black tree exact matching phase, for the Parquet files within the partition filtered by the stable Bloom filter (SBF), the Spark query engine performs exact matching in the red-black tree index; the red-black tree index stores the CRC hash value and its corresponding set of partition Parquet files, and the Spark query engine repeats the CRC hash value calculation in step 1.2 for the target column, and then traverses the obtained hash value from the root node to the leaf node; if a leaf node with an equal CRC hash value is found, the set of Parquet files stored in the leaf node is obtained as the candidate partition Parquet files and enters the row group-level filtering; if not found, the column name does not exist in the partition, and all Parquet files within the partition are directly skipped; the file-level encryption filtering phase reduces the data scan volume through the two-level index mechanism of the stable Bloom filter and the red-black tree, and only retains the candidate partition directory Parquet files containing the target column; In step 2.3, during the row group-level encryption filtering phase, the Spark query engine combines order-preserving encryption and the encrypted stable Bloom filter (SBF) to further filter the candidate row groups inside the partition Parquet files, which are divided into range queries and equality queries according to the different types of user query predicates; Step 2.3.1 For range queries: SQL range queries in the user query conditions involving predicates such as '>', '<', 'BETWEEN'; The Spark query engine performs ciphertext transformation on the plaintext predicate; transmits the query range boundary values in the plaintext predicate to the inside of the TEE, and encrypts the boundary values using the same key K and order-preserving encryption algorithm as in step 1.3 to generate a ciphertext predicate containing the encrypted boundary values, and the ciphertext boundary values generated by order-preserving encryption are in the same numerical order as the corresponding plaintext; Compare the ciphertext maximum and minimum values of each column within the row groups stored in the partition Parquet files; Judge and filter out the candidate row groups that intersect with the query encryption boundary value range and the ciphertext maximum and minimum values of each column, and include them in the subsequent decryption and exact matching processes; for the row groups that have no intersection with the query encryption boundary value range and the ciphertext maximum and minimum values of each column, they are directly skipped; Step 2.3.2 For Equivalent Query: SQL equivalent query where the user's query conditions involve predicates such as '=', 'IN'; The Spark query engine performs a ciphertext transformation on the plaintext predicates, transmits the query column names and query column values in the plaintext predicates inside the TEE, calculates the query hash in the same way as Step 1.3, and generates k encrypted positions; subsequently, checks the stable Bloom filter (SBF) of each row group of the candidate partition Parquet file. If the bits at all k positions are 1, then this row group may contain the query value and enters the decryption process; if any position bit is 0, then the row group does not contain the query value and is directly skipped; Step 2.4, after filtering at the partition, file, and row group levels, only the target columns that meet the query conditions need to be decrypted; Read the Footer of the Parquet file to obtain the offsets and column data ciphertext lengths of each query's target column within the row group; Use the key K to decrypt the target column data inside the TEE and perform the final query calculation.
4. A big data hierarchical encryption filtering mechanism and query optimization system based on TEE, characterized in that Including an encrypted metadata generation module, a partition-level encrypted index module, a file-level stable Bloom filter (SBF) and red-black tree, a row group-level encrypted filtering module, and a column-level on-demand decryption mechanism; among them: The encrypted metadata generation module is used in Step 1. By synchronously executing during the process of writing data into the partition Parquet file, the encrypted metadata is embedded into the Footer metadata structure of the partition Parquet file at the data writing stage; The partition-level encrypted index module is used in Step 2.
1. Through the hash prefix and OPE hybrid encryption mechanism, an encrypted partition directory name is generated and ciphertext matching is performed during the query, to achieve efficient partition pruning and data distribution privacy protection; The file-level stable Bloom filter (SBF) and red-black tree are used in Step 2.
2. By combining the stable Bloom filter and the red-black tree index, fingerprints are generated for the column names and two-level filtering is performed, to quickly locate the candidate partition Parquet files containing the target columns and reduce false positive errors; The row group-level encrypted filtering module is used in Step 2.
3. By analyzing the SQL query predicates, it is judged whether it contains an equivalent query or a range query; if the query predicate contains a range query, the encrypted statistical values of each column within the row group are extracted from the Footer of the partition Parquet file and compared with the order-preserving encrypted query boundary values to determine whether there is an intersection between the ciphertext value interval and the query boundary range; if the query predicate contains an equivalent query, it is judged whether there may be a matching value by checking whether the corresponding k positions in the encrypted Bloom filter (SBF) are hit, so as to filter out the row groups that match the query conditions in the ciphertext environment; The column-level on-demand decryption mechanism is used in Step 2.
4. By reading the Footer of the partition Parquet file to obtain the offsets and lengths of the target columns and decrypting them with the key K inside the TEE, only the target column data is decrypted and the final query calculation is completed.
5. A big data hierarchical encryption and filtering mechanism and query optimization device based on TEE, characterized in that, Including: A memory for storing computer programs; A processor for executing the computer program according to the big data hierarchical encryption filtering mechanism and query optimization method of the TEE in any one of steps 1 to 2.
6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement the big data hierarchical encryption filtering mechanism and query optimization based on the big data hierarchical encryption filtering mechanism and query optimization method of the TEE.
Citation Information
Cited By
Data query method and device, electronic equipment and storage medium
CN120929480A
Encrypted database system based on trusted execution environment
CN120951359A