A private information retrieval method based on a trusted execution environment
Patent Information
- Application Number
- CN202310926387.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-26
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-07-26
AI Technical Summary
现有的私有信息检索方法虽然在理论上都是安全的,但都引入了较大的计算量和计算复杂度,从而不具备现实可行性
[0035]与现有技术相比,本发明的有益效果为:本发明提供一种基于可信执行环境TEE的私有信息检索协议,该方法在保护用户请求内容的同时,还能够保护用户对于数据库的访问模式,对可能恶意的外部数据库屏蔽用户的真实访问模式;该协议分为常规状态(normal)和重排状态(shuffle)两个状态。具体具有如下优点:
Smart Images

Figure CN117056964B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data privacy protection, and in particular to a method for retrieving private information based on a Trusted Execution Environment (TEE). Background Technology
[0002] In traditional information retrieval protocols, users send query requests to servers, which then return corresponding data. However, with the advent of the big data era, the value of user personal data is no longer limited to the protection of personal privacy; its value in production activities is becoming increasingly prominent. For example, e-commerce platforms can learn about user preferences by saving users' browsing history, creating user profiles, and thus accurately recommending relevant products to users. Through this, platforms not only increase sales volume but can also provide search ranking services to merchants for further profit. All of this is driven by user personal data. The issue of user data privacy protection urgently needs to be addressed.
[0003] In traditional information retrieval protocols, users send query requests to servers, and the servers return corresponding information based on the requests. In this scenario, the server may collect and store user-related information without authorization, infringing on user privacy. Private Information Retrieval (PIR) is an important means of protecting user data privacy because it allows servers to return the information a user needs without knowing the specific information the user is querying. While existing PIR methods are theoretically secure, they all introduce significant computational load and complexity, making them impractical.
[0004] Unintentional sorting can reorder a sequence by element size without revealing the size information between elements. Batcher's bitonic sort is an important method for implementing unintentional sorting; it accesses and swaps data in a fixed, predefined order. Because its access pattern is independent of the final order of the data being sorted, it can be used as a method for unintentional sorting. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a private information retrieval method based on a trusted execution environment.
[0006] In a first aspect, the present invention provides a private information retrieval method based on a trusted execution environment, comprising: switching between a normal state and a rearranged state based on the results of data access;
[0007] The normal state includes: setting up a cache with a two-layer hash structure within the trusted execution environment, obfuscating the arriving data based on the cache, and shielding the external database from repeated access to arbitrary data;
[0008] The rearrangement state includes: randomly generating a new permutation, applying it to the original database through unintentional sorting, and making the new permutation unrelated to the original permutation; and writing the data changes recorded in the cache back to the database.
[0009] Furthermore, the normal state also includes:
[0010] Waiting for encrypted access requests;
[0011] When the first batch of access requests arrives at the trusted execution environment, each request in the first batch of access requests is decrypted and mapped into a cache area with a two-level hash table structure for searching, answering and marking the hit requests;
[0012] Copy the first batch of access requests to the second batch of access requests, and iterate through the second batch of access requests; for each request in the second batch of access requests, generate a randomly generated request, and map the randomly generated request to the cache for searching; repeat the generation until the randomly generated request corresponding to a certain request in the second batch of access requests does not exist in the cache.
[0013] For each request in the second batch of access requests, check the marker at the corresponding position of the first batch of access requests. If the marker exists, replace the request with a randomly generated request.
[0014] Send all requests in the second batch of access requests to a database outside the trusted execution environment for querying and obtain the results; write the results into the second batch of access requests;
[0015] Iterate through all requests in the second batch of access requests, check the marker at the corresponding position of the first batch of access requests, and if the marker exists, write the request to the corresponding position of the first batch of access requests in an unintentional manner.
[0016] Encrypt and return the first batch of access requests to the client that sent the data request;
[0017] Map all data from the second batch of access requests and write it into the cache, then update the cache; if the first-level hash table overflows, write it into the second-level hash table; if the second-level hash table overflows, the protocol enters the rearrangement state.
[0018] Furthermore, the two-level hash table structure cache includes:
[0019] Divide the buffer that can hold n data into two equal parts, denoted as the first-level hash table and the second-level hash table; the hash function corresponding to the first-level hash table is denoted as the first hash function, and the hash function corresponding to the second-level hash table is denoted as the second hash function; further divide the first-level hash table and the second-level hash table into b buckets, each of which can hold k data, then n = 2*b*k.
[0020] Furthermore, the process of searching in the cache includes:
[0021] When mapping the data to be searched to a first-level hash table or a second-level hash table, the hash value of the data to be searched is first calculated using the hash function corresponding to the hash table of that level, and then modulo the number of buckets b to locate a specific bucket in that hash table. Subsequently, all data items in the bucket are traversed to search for whether the data has been matched.
[0022] Furthermore, iterate through all requests in the second batch of access requests. If a request is marked, then subtly write that request to the corresponding location in the first batch of access requests, including:
[0023] Simultaneously, retrieve the corresponding data bits from the first batch of access requests and the same data bits from the corresponding requests in the second batch of access requests. Check the flag of the request. If it exists, write the corresponding data bits of the first batch of access requests back to the first batch of access requests; otherwise, write the corresponding data bits of the second batch of access requests back to the first batch of access requests.
[0024] Furthermore, the rearranged states include:
[0025] Based on the amount of data i in the database, set the total number of sorting rounds to 1. Where i is a positive integer;
[0026] When the current sorting round is 1, generate a random hash key; read the first batch of data from the external database in batches according to the size of the trusted execution environment and in a predefined order using the bitonic sorting algorithm; calculate the new hash value of the first batch of data read in using the random hash key; sort the read data in a predefined address access order according to the new hash value and write it back to the external database.
[0027] Re-read the second batch of data for the current round, and repeat the above steps until all data has been processed, at which point the current round ends.
[0028] Begin the next round, repeating the above steps until the preset sorting round is reached. Then the rearrangement state ends, and the system switches to the normal state.
[0029] The rearrangement state also includes:
[0030] The read-in external data is mapped to the cache for searching. If the data is found, the read-in external data is updated with the data in the cache.
[0031] Furthermore, reading data into an external database and / or writing data back to an external database includes:
[0032] Set up a shared cache area outside the trusted execution environment and the external space, and create read and write threads in the trusted execution environment and the external space respectively.
[0033] Secondly, the present invention provides an electronic device, including a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the aforementioned private information retrieval method based on a trusted execution environment.
[0034] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the above-described private information retrieval method based on a trusted execution environment.
[0035] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a private information retrieval protocol based on a Trusted Execution Environment (TEE). This method protects not only the content requested by the user but also the user's access pattern to the database, shielding the user's true access pattern from potentially malicious external databases. The protocol has two states: a normal state and a shuffle state. Specifically, it has the following advantages:
[0036] In the normal state, the existence of the cache and the efficient request replacement algorithm design enable high protocol access performance while ensuring access mode security.
[0037] In the shuffle state, in addition to the confidentiality of the new database arrangement, the use of buffers and parallel techniques avoids the TEE reentrancy overhead that may exist under single-threaded reordering conditions, speeds up the reordering process, and reduces the dead time of service interruption in the shuffle state.
[0038] In addition, the present invention can also easily increase throughput by expanding the number of running machines. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 This is the overall architecture diagram of private information retrieval based on TEE;
[0041] Figure 2 This is a diagram of the cache area;
[0042] Figure 3 This is a diagram illustrating the normal state;
[0043] Figure 4 This is a diagram illustrating the rearranged state;
[0044] Figure 5 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0046] It should be noted that, where there is no conflict, the features in the following embodiments and implementations can be combined with each other. All unintentional sorting problems involved in this invention are solved using bitonic sorting.
[0047] The machines and hardware involved in this invention, such as Figure 1As shown, the server-side application of this invention is hardware-divided into two modules: a data storage service and a Trusted Execution Environment (TEE). The data storage service can store at least N data items, each a key-value pair, enabling O(1) data access. The TEE is required to allocate a cache area, which can store at most n data items, where n is much smaller than N. This method uses a symmetric key to encrypt user-submitted requests, thus protecting the request content before it reaches the server's Trusted Execution Environment. After the request reaches the server's Trusted Execution Environment, a private cache with a special structure allocated in secure memory is used to obfuscate the actual access request, thereby shielding the database from the client's actual data access patterns. Before and after cache overflow, the method can be divided into a normal state and a rearranged state. In the normal state, an efficient and secure replacement algorithm is used to transform repeated access to the same data into random access to previously unaccessed data. In the rearranged state, the database is rearranged and then returned to the normal state, achieving the goal of securely reusing the limited cache area.
[0048] Unlike conventional databases, in the method proposed in this invention, requests from clients are not directly submitted to the data storage service. Instead, they are first processed by an instance of this method running in the TEE (Transaction Execution Environment). This instance constructs a new request based on the current cycle state and the content characteristics of the request, or performs a reordering before processing in the normal state, thus ensuring the security of the request access pattern.
[0049] The private information retrieval method proposed in this invention is periodic, with each period divided into a normal state and a shuffle state before and after cache overflow. The request processing procedures for these two states are described in detail below with reference to the accompanying drawings.
[0050] 1) The server receives the first batch of access requests (Batch1) and forwards them to the trusted execution environment for processing.
[0051] 2.1) As Figure 3 As shown, if the protocol is in the normal state, the system first uses the data in the cache to answer part of the request, and replaces the part of the request in the cache with the request that has not been accessed since entering the current normal state. Then, it submits the request to the data storage service to retrieve the data, and finally writes the newly retrieved data into the cache for future use.
[0052] Furthermore, such as Figure 2As shown, the buffer that can hold n data items is divided into two equal parts, denoted as the first-level hash table and the second-level hash table. The hash function corresponding to the first-level hash table is denoted as the first hash function, and the hash function corresponding to the second-level hash table is denoted as the second hash function. The first-level hash table and the second-level hash table are further divided into b buckets, and each bucket can hold k data items, so n = 2*b*k.
[0053] The specific replacement and processing methods are as follows:
[0054] 2.1.1) First, iterate through the first batch of access requests (Batch1), mapping each request sequentially to a two-level hash table in the cache for searching. Specifically: First, use the first hash function Hc1 of the first-level hash table to calculate the corresponding hash position of the current request, and then iterate through and search for all data at that hash position. Then, use the second hash function Hc2 of the second-level hash table to calculate the corresponding hash position of the current request, and then iterate through and search for all data at that hash position. For cases where a request hits either the first or second-level hash table, directly use the hit data to respond and mark the request as hit. For cases where a request misses, use the original request data to respond, ensuring the unintentionality of memory access patterns.
[0055] 2.1.2) Copy the first batch of access requests (Batch1) to create a second batch of access requests (Batch2). Iterate through the second batch of access requests (Batch2). For each request in the second batch of access requests (Batch2): randomly generate a Random request and map it to the two-level hash tables in the cache for searching. Specifically: first, use the first hash function Hc1 of the first-level hash table to calculate the corresponding hash position of the current request, and iterate through all data in that hash position; then, use the second hash function Hc2 of the second-level hash table to calculate the corresponding hash position of the current request, and iterate through all data in that hash position. Repeat generating Random requests until the above search does not hit. If the currently iterated request was marked as hit in step 2.1.1, replace the request with Random requests unintentionally. Specifically: retrieve the same data bits of the current request and the Random request, check the hit mark, and if it exists, write the corresponding data bits of the Random request back to the second batch of access requests (Batch2); otherwise, write the corresponding data bits of the current request back.
[0056] 2.1.3) Forward the processed second batch of access requests (Batch2) to the external database to query the results;
[0057] 2.1.4) Iterate through the second batch of access requests Batch2. For each request in the second batch of access requests Batch2: if it was marked as hit in step 2.1.1, write it into the corresponding position of the first batch of access requests Batch1 in an unintentional manner. Specifically, retrieve the data bits of the corresponding request in the first batch of access requests Batch1 and the same data bits of the corresponding request in the second batch of access requests Batch2. Check the hit mark. If there is one, write the corresponding data bits of the first batch of access requests Batch1 back to the first batch of access requests Batch1. Otherwise, write the corresponding data bits of the second batch of access requests Batch2 back to the second batch of access requests Batch2.
[0058] It should be noted that by simultaneously retrieving the data bits of the corresponding requests in the first batch of access requests and the same data bits of the corresponding requests in the second batch of access requests, and then putting the data bits back into the first batch of access requests according to the flag, it is unclear from the perspective of the memory observer whether the data bits put back into the first batch of access requests are the original (i.e., the data bits of the corresponding requests in the first batch of access requests) or obtained from the outside (i.e., the same data bits of the corresponding requests in the second batch of access requests). This achieves the unintentional replacement of requests that have not been accessed, without exposing which data has been accessed or not.
[0059] 2.1.5) Encrypt all requests and data in the first batch of access requests (Batch 1) and return them to the client.
[0060] 2.1.6) Traverse the second batch of access requests Batch2 and write each request in Batch2 into the cache: First, try to write to the first-level hash table. If the first-level hash table overflows, try to write to the second-level hash table. If the second-level hash table also overflows, the normal state terminates and the protocol enters the shuffle state.
[0061] 2.2) As Figure 4 As shown, if the protocol is in shuffle state, a hash function Hs is first randomly generated to create a new permutation of the original database. Bitonic sort is then applied to rearrange the database, reorganizing it according to the newly generated permutation. After the bitonic sort algorithm finishes executing, the protocol switches back to the normal state. The specific processing method is as follows:
[0062] 2.2.1) Based on the amount of data i in the database, set the total number of sorting rounds to... Where i is a positive integer;
[0063] 2.2.2) Read the first batch of data from the external database in batches according to the size of the trusted execution environment and in a predefined order using the bitonic sorting algorithm;
[0064] 2.2.3) Calculate the new hash value of the first batch of data read in using the hash function Hs;
[0065] 2.2.4) Randomly sort the read-in data according to the new hash value in a predefined address access order and write it back to the external database;
[0066] 2.2.5) Re-read the second batch of data for the current round, and repeat the above steps until all data has been processed, at which point the current round ends;
[0067] 2.2.6) Start the next round and repeat the above steps until the preset sorting round is reached. Then the shuffle state ends and switches to the normal state.
[0068] The process of reading data into an external database and / or writing data back to an external database includes:
[0069] A shared cache area between the Trusted Execution Environment (TEE) and the external space is set up outside the TEE, and read / write threads are created separately outside the TEE and in the external space. This avoids the time overhead of repeated re-entry into the TEE that may occur when reading external data under single-threaded conditions.
[0070] The rearrangement state also includes:
[0071] 2.2.7) The read external data is mapped to the cache for searching. Specifically: First, the first hash function Hc1 of the first-level hash table is used to calculate the corresponding hash position of the current data, and all data at that hash position are searched. Then, the second hash function Hc2 of the second-level hash table is used to calculate the corresponding hash position of the current data, and all data at that hash position are searched. If the data hits in either the first or second-level hash table, the hit data is used to update the original data; otherwise, the original data itself is used to update the data, ensuring the unintentionality of memory access patterns.
[0072] In summary, this invention provides a private information retrieval protocol based on a Trusted Execution Environment (TEE). This protocol sets up a cache area in a trusted execution environment, avoiding repeated access to certain data by caching accessed data and replacing duplicate requests. This protects the security of user request content while ensuring the security of access patterns and achieving efficient data access. When the cache overflows, this invention rearranges the data to reuse the limited cache area, and writes the data in the cache back to the original database during the rearrangement process, ensuring data consistency.
[0073] The access protocol proposed in this invention is periodic, and a protocol cycle can be divided into a normal state and a shuffle state before and after a cache overflow. By efficiently utilizing existing data in the cache to answer and replace repeated access requests, the protocol in the normal state can ensure the security of the request access pattern and achieve high access performance. In the shuffle state, this invention uses a cache and parallel technology to avoid the overhead of repeated re-entry into the TEE, reducing the time that the protocol will stop service due to data rearrangement.
[0074] Accordingly, this application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; and, when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the aforementioned private information retrieval method based on a trusted execution environment. Figure 5 The diagram shown illustrates a hardware structure of any device with data processing capabilities for a private information retrieval method based on a trusted execution environment, as provided in an embodiment of the present invention. (Except for...) Figure 5 In addition to the processor, memory, and network interface shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.
[0075] Accordingly, this application also provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, implement the aforementioned method for retrieving private information based on a trusted execution environment. The computer-readable storage medium can be an internal storage unit of any data-processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units of any data-processing device and external storage devices. The computer-readable storage medium is used to store the computer program and other programs and data required by the data-processing device, and can also be used to temporarily store data that has been output or will be output.
[0076] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only.
[0077] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A method for retrieving private information based on a trusted execution environment, characterized in that, include: Switches between normal and rearranged states based on the results of data access; The normal state includes: setting up a cache with a two-layer hash structure in the trusted execution environment, obfuscating the arriving data based on the cache, and shielding the database from repeated access to arbitrary data; The rearrangement states include: randomly generating a new permutation, applying it to the original database through unintentional sorting, making it impossible for the new permutation to establish a connection with the original permutation; and returning to the normal state after writing the data changes recorded in the cache back to the database. The normal state also includes: Waiting for encrypted access requests; When the first batch of access requests arrives at the trusted execution environment, each request in the first batch of access requests is decrypted and mapped into a cache area with a two-level hash table structure for searching, answering and marking the hit requests; Copy the first batch of access requests to the second batch of access requests, and iterate through the second batch of access requests; for each request in the second batch of access requests, generate a randomly generated request, and map the randomly generated request to the cache for searching; repeat the generation until the randomly generated request corresponding to a certain request in the second batch of access requests does not exist in the cache. For each request in the second batch of access requests, check the marker at the corresponding position of the first batch of access requests. If the marker exists, replace the request with a randomly generated request. Send all requests in the second batch of access requests to a database outside the trusted execution environment for querying and obtain the results; write the results into the second batch of access requests; Iterate through all requests in the second batch of access requests, check the marker at the corresponding position of the first batch of access requests, and if the marker exists, write the request into the corresponding position of the first batch of access requests in an unintentional manner. Encrypt and return the first batch of access requests to the client that sent the data request; Map all data from the second batch of access requests and write it into the cache, then update the cache. If the first-level hash table overflows, write it into the second-level hash table. If the second-level hash table overflows, the protocol enters a reordering state.
2. The private information retrieval method based on a trusted execution environment according to claim 1, characterized in that, The two-level hash table structure cache includes: Divide the buffer that can hold n data into two equal parts, denoted as the first-level hash table and the second-level hash table; the hash function corresponding to the first-level hash table is denoted as the first hash function, and the hash function corresponding to the second-level hash table is denoted as the second hash function; further divide the first-level hash table and the second-level hash table into b buckets, each of which can hold k data, then n = 2*b*k.
3. The private information retrieval method based on a trusted execution environment according to claim 2, characterized in that, The process of searching in the cache includes: When mapping the data to be searched to a first-level hash table or a second-level hash table, the hash value of the data to be searched is first calculated using the hash function corresponding to the hash table of that level, and then modulo the number of buckets b to locate a specific bucket in that hash table. Subsequently, all data items in the bucket are traversed to search for whether the data has been matched.
4. The private information retrieval method based on a trusted execution environment according to claim 1, characterized in that, Iterate through all requests in the second batch of access requests. If a request is marked, subtly write that request to the corresponding location in the first batch of access requests, including: Simultaneously, retrieve the corresponding data bits from the first batch of access requests and the same data bits from the corresponding requests in the second batch of access requests. Check the flag of the corresponding request in the second batch of access requests. If it exists, write back the corresponding data bits of the first batch of access requests to the first batch of access requests. Otherwise, write back the corresponding data bits of the second batch of access requests.
5. The private information retrieval method based on a trusted execution environment according to claim 1, characterized in that, Rearrangement states include: Based on the amount of data i in the database, set the total number of sorting rounds to I = 1 + 2 + … + , where i is a positive integer; When the current sorting round is 1, generate a random hash key; read the first batch of data from the external database in batches according to the size of the trusted execution environment and in a predefined order using the bitonic sorting algorithm; calculate the new hash value of the first batch of data read in using the random hash key; sort the read data in a predefined address access order according to the new hash value and write it to the database. Re-read the second batch of data for the current round, and repeat the above steps until all data has been processed, at which point the current round ends. Start the next round and repeat the above steps until the preset sorting round is reached. Then the rearrangement state ends and switches to the normal state. The rearrangement state also includes: The read data is mapped to the cache for searching. If the data is matched, the read data is updated using the data in the cache.
6. The private information retrieval method based on a trusted execution environment according to claim 5, characterized in that, Reading data into an external database and / or writing data back to an external database includes: Set up a cache area shared by the trusted execution environment and the external space outside the trusted execution environment, and create read and write threads in the external space and the external space respectively.
7. An electronic device comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the private information retrieval method based on a trusted execution environment as described in any one of claims 1-6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the private information retrieval method based on a trusted execution environment as described in any one of claims 1-6.