Symmetrical searchable encryption index construction method with enhanced security

Through the blocking strategy of multiple mapping and Path-ORAM technology, the information leakage and overhead of symmetric searchable encrypted indexes during the query process is solved, and an efficient and privacy-protected single-keyword query index structure is realized.

CN120337265APending Publication Date: 2025-07-18FUJIAN NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510508936.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing symmetric searchable cryptographic index construction method has information leakage and large storage and communication overhead problems during the query process, especially in single-keyword query of large data sets.

Method used

Multi-mapping and Path-ORAM technology are used to fill and block the document identifier through chunking policy, generate Position table, Stash table and Residue table, and hide access mode and capacity mode through query trap gate and eviction policies during the query process, reducing storage and communication overhead.

Benefits of technology

Effectively hide information leakage during the query process, reduce storage and communication overhead, ensure that each query returns the correct document identifier, avoid false negative phenomena, and optimize storage utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337265A_ABST
    Figure CN120337265A_ABST
Patent Text Reader

Abstract

The invention provides a security-enhanced symmetric searchable encrypted index construction method, which comprises the following steps of: scanning a plaintext document set by a user, extracting keywords contained in each document, and generating an inverted index; according to the occurrence frequency of each keyword in the document set, a user firstly adopts a partitioning strategy to carry out filling and partitioning processing on a corresponding document identifier, then loads a block corresponding to each keyword into a Path-ORAM tree, and generates a Position table, a Sash table and a Resistance table; the user encrypts the Path-ORAM tree and uploads the Path-ORAM tree to a cloud server, and other table items are stored locally by the user; and the user queries by sending the query trap door # imgabs0 # corresponding to the keyword w to obtain the document identifier corresponding to the keyword w. According to the method, the query index capable of hiding information leakage in symmetric searchable encryption can be designed by utilizing multiple mapping and a Path-ORAM technology, and the storage and communication overhead is reduced while common leakage is hidden by modifying a mapping and expelling strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information security, and particularly to a method for constructing a symmetric searchable encryption index with enhanced security. Background Art

[0002] Existing methods for index construction in symmetric searchable encryption include: PRT-EMM, VH-EMM, TWORAM, Eurus, ZPH-OBI.

[0003] PRT-EMM uses a pseudo-random transformation function to shuffle the length of the document index tuple corresponding to the keyword, hiding the capacity pattern leakage; VH-EMM uses the cuckoo hashing technique to obscure the number and storage location of the document identifiers corresponding to the keyword. PRT-EMM and VH-EMM technologies hide the capacity pattern by obscuring the number of document identifiers corresponding to the keyword. However, both schemes are implemented in a static manner, that is, once the index is generated, it will not be updated anymore. Therefore, the search pattern will still be leaked during the query process.

[0004] Path-ORAM is implemented based on a binary tree to achieve oblivious reading of stored data. Conventional Path-ORAM includes three components: a Path-ORAM binary tree, a Position mapping table, and a Stash buffer. Among them, the height of the ORAM binary tree is L. Each node in the tree is called a bucket, and each bucket contains Z blocks, and the size of each block is B (in bytes), which is used to store user privacy information; the Position mapping table is used to store the physical location information of each block; the Stash buffer is used to temporarily store the read or overflowed blocks. TWORAM stores a single keyword and document pair in a block in the Path-ORAM tree. Therefore, a single keyword query corresponds to D w (indicating the number of documents corresponding to the keyword w) times of ORAM tree readings, resulting in a large storage overhead.

[0005] Eurus stores all the identifiers corresponding to a single keyword in a block and fills the number of document identifiers corresponding to all keywords to be the same (worst filling). Although each query only requires one ORAM tree reading, the worst filling results in a large storage overhead and communication overhead.

[0006] To hide the capacity pattern, ZPH-OBI still uses the worst filling and divides every u keyword and document identifier pairs into a block. Therefore, a single keyword query also corresponds to D w / u ORAM tree path readings, still resulting in a large storage and communication overhead.

[0007] In view of this, to construct a single-keyword query index for large datasets and construct a query index with high search efficiency and no information leakage, the present invention proposes a method for constructing a security-enhanced symmetric searchable encryption index. Summary of the Invention

[0008] The object of the present invention is to propose a method for constructing a security-enhanced symmetric searchable encryption index, which uses multiple mapping and Path-ORAM technology to design a query index that can hide information leakage in symmetric searchable encryption, and reduces storage and communication overhead while hiding common leaks by modifying the mapping and eviction strategies.

[0009] To achieve the above object, the technical solution of the present invention is: a method for constructing a security-enhanced symmetric searchable encryption index, specifically including the following steps:

[0010] S1. The user scans the plaintext document set, extracts the keywords contained in each document, and generates an inverted index;

[0011] S2. According to the frequency of each keyword appearing in the document set, the user first adopts a chunking strategy to fill and chunk the corresponding document identifiers, and then loads the chunks corresponding to each keyword into the Path-ORAM tree, and generates a Position table, a Stash table, and a Residue table;

[0012] S3. The user encrypts the Path-ORAM tree and uploads it to the cloud server, and stores the remaining table entries locally;

[0013] S4. The user sends the query trapdoor τ w corresponding to the keyword w for querying to obtain the document identifiers corresponding to the keyword w.

[0014] Preferably, the step of adopting a chunking strategy to fill and chunk the corresponding document identifiers specifically is:

[0015] Each node in the Path-ORAM tree represents a bucket, each bucket contains Z blocks, and each block can fixedly store MAX / Z document identifiers; a bucket in the Path-ORAM tree can store all the document identifiers corresponding to a single keyword, and MAX represents the maximum value of the number of document identifiers corresponding to a single keyword; the document identifiers corresponding to each keyword are chunked according to a fixed size of MAX / Z, and if the number of document identifiers corresponding to a certain keyword is less than MAX / Z, the number of document identifiers is filled to the fixed size.

[0016] Preferably, the Position table is used to store the mapping relationship from each keyword to the leaf node, and the document identifier blocks associated with the keyword w are all stored in random blocks on the path P(l w ) in the Path-ORAM tree; the mapping is implemented through a pseudo-random function:

[0017]

[0018] where l w is the leaf node, G is the pseudo-random function, K1 is the key, r is the random number, and l w is stored in Position[w].

[0019] Preferably, the Path-ORAM tree is used to store the document identifier blocks corresponding to each keyword, and the document identifier blocks associated with the keyword w are randomly stored in the nodes on the path P(l w ); the plaintext format of the content stored in the block in the Path-ORAM tree is:

[0020]

[0021] link w is a structure representing the association relationship between the keyword w and its corresponding document identifier. For the keyword w, ids represents a fixed number (MAX / Z) of document identifiers stored in the corresponding block, represents the number of times the user accesses link w , and each block in the Path-ORAM tree stores an encrypted link.

[0022] Preferably, the Residue table is used to record the number of remaining free blocks in each node in the Path-ORAM tree, which is used to determine whether the subsequent insertion operation on the node can be executed; and R(P(l w )) is used to represent the number of remaining free blocks on the path P(l w ).

[0023] Preferably, the Stash table is used to temporarily store the pairs of keywords and document identifiers read or overflowed, and then insert the elements in the Stash into the ORAM tree according to the rules; among them, if link x is read from the node numbered seq on the path P(l w ), the storage format is (seq||link x ||l w ), if link x cannot be inserted into the ORAM tree due to overflow, the storage format is (-1||link x||l x ),l x Indicates the leaf node number corresponding to the keyword x.

[0024] Preferably, the user sends a query trapdoor τ corresponding to the keyword w w for querying to obtain the document identifier corresponding to the keyword w, specifically as follows:

[0025] For the search operation of the keyword w, the user reads the node corresponding to the number in the query trapdoor τ w in the Path-ORAM tree to the local, decrypts each one and matches it with the keyword w, so as to obtain the correct document identifier under the condition of hiding the user's access intention; after several reads are completed, the user evicts the link stored in the local Stash table back to the Path-ORAM tree;

[0026] a) Generate the query trapdoor τ w : By looking up the Position table, obtain the leaf node number l corresponding to the keyword w w and generate the path P(l w ), then traverse the elements in the Stash table (seq||link x ||l w ): If the path P(l w ) contains the node number seq, then delete this node number from P(l w ) to generate P'(l w ); finally, let the query trapdoor τ w = P'(l w );

[0027] b) Query: The user sends the query trapdoor τ w = P'(l w ) to the cloud server, then the server returns and deletes all the nodes corresponding to the query trapdoor, and the format of the returned nodes is (seq||block); the user stores the received nodes in the Stash table and clears the corresponding positions in the local Residue table;

[0028] c) Local operations after query: The user decrypts each returned block one by one to generate elements in the form of where x is a placeholder representing the current keyword being processed, and updates Subsequently, store (seq||link x ||l w ) in the Stash table; then the user traverses the Stash table and compares one by one. If w = x, then read the corresponding ids;

[0029] d) Update: After several queries, the update operation is triggered, and the user evicts all (seq || link x || l w ) in the Stash table; First, the user creates a corresponding free block for path P(l w ) locally. If x ≠ w, for all links x corresponding to keyword x, let l x = Position[x], encrypt the link x and randomly store it in the block on path P(l x ) ∩ P(l w ); If x = w, then update Position[w] such that the number of free blocks R(P(l' w ) ∩ P(l w )) on the path P(l' w ) ∩ P(l w ) is greater than the number of links w corresponding to keyword w. Then encrypt all links w and randomly store them in this path, and update the Residue table.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] Compared with VH-EMM, the present invention places the inverted index in the Path-ORAM tree and, through the strategy of querying and updating simultaneously, hides the access pattern and search pattern leaked by the user during the query process; ensures that all document identifiers corresponding to the keyword are returned for each query, without generating false negative phenomena and without causing query information loss; does not require worst-case padding, and divides the document identifiers corresponding to each keyword into blocks according to the occurrence frequency of different keywords.

[0032] Compared with Eurus and ZPH-OBI, the present invention does not require worst-case padding. According to the occurrence frequency of different keywords, the document identifiers corresponding to each keyword are divided into blocks of the same size, and the blocks corresponding to different keywords share the nodes in the tree, hiding the capacity pattern leaked by the user during the query process while reducing the storage overhead; the query for a single keyword only requires one access to the Path-ORAM tree and only returns one path in the tree, which can reduce the communication overhead required for the query; free blocks are introduced into the Path-ORAM tree as the basis for subsequent rewriting. And by adopting deferred rewriting, the read blocks are temporarily stored in the local Stash table and then uniformly evicted later, which can greatly reduce the subsequent query communication overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a schematic diagram of the index structure of the present invention;

[0034] Figure 2 Schematic diagram of the query and update process of the present invention Specific implementation manners

[0035] The following combines the attached Figure 1-2 , and specifically describes the technical solution of the present invention.

[0036] The present invention proposes a method for constructing a security-enhanced symmetric searchable encryption index. By introducing a Path-ORAM for recording free blocks to load the blocks corresponding to each keyword, it hides the access pattern, search pattern, and capacity pattern leakage, and specifically includes the following steps:

[0037] S1. The user scans the plaintext document set, extracts the keywords contained in each document, and generates an inverted index;

[0038] S2. According to the frequency of each keyword appearing in the document set, the user first uses a chunking strategy to fill and chunk the corresponding document identifiers, and then loads the chunks corresponding to each keyword into the Path-ORAM tree, and generates a Position table, a Stash table, and a Residue table;

[0039] S3. The user encrypts the Path-ORAM tree and uploads it to the cloud server, and stores the remaining table entries locally;

[0040] S4. After the index construction is completed, the user sends a query trapdoor τ w corresponding to the keyword w for querying to obtain the document identifier corresponding to the keyword w.

[0041] In this embodiment, the step of using the chunking strategy to fill and chunk the corresponding document identifiers specifically is:

[0042] Each node in the Path-ORAM tree represents a bucket, each bucket contains Z blocks, and each block can fixedly store MAX / Z document identifiers; a bucket in the Path-ORAM tree can store all the document identifiers corresponding to a single keyword, and MAX represents the maximum value of the number of document identifiers corresponding to a single keyword; the document identifiers corresponding to each keyword are chunked according to a fixed size of MAX / Z. If the number of document identifiers corresponding to a certain keyword is less than MAX / Z, the number of document identifiers is filled to the fixed size. The document identifiers of each keyword are chunked, and the insufficient part is filled to this fixed size.

[0043] In this embodiment, the Position table is used to store the mapping relationship between each keyword and its corresponding leaf node. All document identifier blocks associated with the keyword w are stored in random nodes on the path P(l w ) in the Path-ORAM tree; the mapping is implemented through a pseudo-random function:

[0044]

[0045] where l w is a leaf node, G is a pseudo-random function, K1 is a key, r is a random number, and l w is stored in Position[w].

[0046] In this embodiment, the Path-ORAM tree is used to store the document identifier blocks corresponding to each keyword. All document identifier blocks associated with the keyword w are randomly stored in the nodes on the path P(l w ); the plaintext format of the content stored in the block in the Path-ORAM tree is:

[0047]

[0048] link w is a structure representing the association relationship between the keyword w and its corresponding document identifier. For the keyword w, ids represents a fixed number (MAX / Z) of document identifiers stored in the corresponding block, represents the number of times the user accesses link w , and each block in the Path-ORAM tree stores an encrypted link.

[0049] In this embodiment, the Residue table is used to record the number of remaining free blocks in each node in the Path-ORAM tree, which is used to determine whether the subsequent insertion operation on the node can be executed; and R(P(l w )) is used to represent the number of remaining free blocks on the path P(l w ).

[0050] In this embodiment, the Stash table is used to temporarily store the keyword and document identifier pairs read or overflowed, and then insert the elements in the Stash into the ORAM tree according to the rules; among them, if link x is read from the node numbered seq on the path P(l w ), the storage format is (seq||link x ||l w ), and if link x cannot be inserted into the ORAM tree due to overflow, the storage format is (-1||linkx ||l x ),l x Indicates the leaf node number corresponding to the keyword x.

[0051] In this embodiment, the user sends a query trapdoor τ corresponding to the keyword w w to perform a query and obtain the document identifier corresponding to the keyword w, specifically as follows:

[0052] For the search operation of the keyword w, the user reads the node corresponding to the query trapdoor τ in the Path-ORAM tree w to the local, decrypts each one and matches it with the keyword w, so as to obtain the correct document identifier under the condition of hiding the user's access intention; after several reads, the user evicts the link stored in the local Stash table back to the Path-ORAM tree;

[0053] a) Generate the query trapdoor τ w : By looking up the Position table, obtain the leaf node number l corresponding to the keyword w w and generate the path P(l w ), and then traverse the elements (seq||link x ||l w ) in the Stash table: If the path P(l w ) contains the node number seq, then delete this node number from P(l w ) to generate P'(l w ); finally, let the query trapdoor τ w = P'(l w );

[0054] b) Query: The user sends the query trapdoor τ w = P'(l w ) to the cloud server, and then the server returns and deletes all the nodes corresponding to the query trapdoor, and the format of the returned nodes is (seq||block); the user stores the received nodes in the Stash table and clears the corresponding position in the local Residue table;

[0055] c) Local operations after query: The user decrypts each returned block to generate elements in the form of where x is a placeholder representing the currently processed keyword, and updates Subsequently, store (seq||link x ||l w ) in the Stash table; then the user traverses the Stash table and compares one by one. If w = x, then read the corresponding ids;

[0056] d) Update: After several queries, the update operation is triggered and the user evicts all (seq||link x ||l w ); First, the user locally sets the path P(l w ) Create the corresponding free block. If x≠w, for all links corresponding to keyword x x , let l x =Position[x], encryptedlink x And store it randomly in path P(l x )∩P(l w ) in the block on the left; if x=w, update Position[w] so that P(l' w )∩P(l w )The number of free blocks in the path R(P(l′ w )∩P(l w )) is greater than the link corresponding to keyword w w The number of links, then w Encrypt and store it randomly in this path, and update the Residue table.

[0057] In summary, the prior art uses the worst filling to hide the query capacity mode, without considering the frequency of occurrence (importance) of each keyword in the actual scenario, and some technologies cannot hide the search mode, that is, the same query will correspond to the same response content. In order to solve the above problems, the present invention fills each keyword according to its importance and divides it into blocks of the same size, hiding the capacity mode without using the worst filling; using Path-ORAM to implement indexing to enhance privacy, and at the same time, mapping all blocks corresponding to a single keyword to a path in the ORAM tree; introducing the number of free blocks for each node of the ORAM tree to record the number of free blocks in the node, and only allowing it to be loaded when the number of free blocks in a certain path is greater than the number of blocks corresponding to the keyword. By observing the frequency of occurrence of different keywords, the present invention divides the document identifier corresponding to each keyword into blocks and loads it into a path in the ORAM tree, which can not only reduce storage and query overhead, but also effectively hide capacity mode and search mode leakage.

[0058] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.

Claims

1. A method for constructing a security-enhanced symmetric searchable encryption index, characterized in that, Specifically, it includes the following steps: S1. The user extracts the keywords contained in each document by scanning the plaintext document set and generates an inverted index; S2. According to the frequency of each keyword appearing in the document set, the user first adopts a chunking strategy to fill and chunk the document identifiers corresponding to the keywords. Subsequently, the chunks corresponding to each keyword are loaded into the Path-ORAM tree, and a Position table, a Stash table, and a Residue table are generated; S3. The user encrypts the Path-ORAM tree and uploads it to the cloud server, and stores the remaining table entries locally; S4. The user sends the query trapdoor τ corresponding to the keyword w w to perform a query and obtain the document identifier corresponding to the keyword w.

2. The method for constructing a security-enhanced symmetric searchable encryption index according to claim 1, characterized in that The adoption of the chunking strategy to fill and chunk the document identifiers corresponding to the keywords specifically is: Each node in the Path-ORAM tree represents a bucket, and each bucket contains Z blocks; Let each block fixedly store MAX / Z document identifiers; where a node in the Path-ORAM tree can store all the document identifiers corresponding to a single keyword, MAX represents the maximum value of the number of document identifiers corresponding to a single keyword; the document identifiers corresponding to each keyword are chunked by a fixed size of MAX / Z. If the number of document identifiers corresponding to a certain keyword is less than MAX / Z, the number of document identifiers is filled to the fixed size.

3. The method for constructing a security-enhanced symmetric searchable encryption index according to claim 2, characterized in that, The Position table is used to store the mapping relationship from each keyword to the leaf node, and the document identifier blocks associated with the keyword w are all stored in the leaf node l of the Path-ORAM tree w on the corresponding path P(l w ); the mapping is implemented through a pseudo-random function: where G is a pseudo-random function, L1 is the key, r is a random number, and l w is stored in Position[w].

4. The method for constructing a security-enhanced symmetric searchable encryption index according to claim 3, characterized in that The Path-ORAM tree is used to store the document identifier blocks corresponding to each keyword, and the document identifier blocks associated with the keyword w are randomly stored in the nodes on the path P(l w );The plaintext format of the content stored in the block in the Path-ORAM tree is as follows: link w is a structure representing the association relationship between the keyword w and the corresponding document identifier. For the keyword w, ids represents a fixed number (MAX / Z) of document identifiers stored in the corresponding block. represents the number of times the user accesses the link w Each block in the Path-ORAM tree stores an encrypted link.

5. The method for constructing a security-enhanced symmetric searchable encryption index according to claim 3, characterized in that, The Residue table is used to record the number of remaining free blocks in each node of the Path-ORAM tree, which is used to determine whether the subsequent insertion operation on the node can be executed; and R(P(l w )) is used to represent the number of remaining free blocks on the path P(l w ).

6. The method for constructing a security-enhanced symmetric searchable encryption index according to claim 5, characterized in that, The Stash table is used to temporarily store the keyword-document identifier pairs read or overflowed, and then insert the elements in the Stash into the ORAM tree according to the rules; among them, if link x is read from the node numbered seq in the path P(l w ), the storage format is (seq||link x ||l w ), if link x cannot be inserted into the ORAM tree due to overflow, the storage format is (-1||link x ||l x ), l x represents the leaf node number corresponding to the keyword x.

7. The method for constructing a security-enhanced symmetric searchable encryption index according to claim 6, wherein The user queries by sending a query trapdoor τ corresponding to the keyword w w to obtain the document identifier corresponding to the keyword w, which is specifically as follows: For the search operation for keyword w, the user reads the query trapdoor τ from the corresponding node with the number in the Path-ORAM tree to the local, decrypts them one by one and matches them with the keyword w, so as to obtain the correct document identifier under the condition of hiding the user's access intention; after several reads are completed, the user evicts the links stored in the local Stash table back to the Path-ORAM tree; w After several reads are completed, the user evicts the links stored in the local Stash table back to the Path-ORAM tree; a) Generate the query trapdoor τ w : By looking up the Position table, obtain the leaf node number l corresponding to the keyword w w and generate the path P(l w ). Then, traverse the elements (seq||link x ||l w ) in the Stash table: If the path P(l w ) contains the node number seq, then delete this node number from P(l w ) to generate P'(l w ); Finally, let the query trapdoor τ w = P'(l w ); b) Query: The user sends a query trapdoor τ to the cloud server w = P'(l w ), after which the server returns and deletes all nodes corresponding to the query trapdoor, and the format of the returned nodes is (seq||block); the user stores the received nodes in the Stash table and clears the corresponding positions in the local Residue table; c) Local operations after query: The user decrypts each returned block one by one to generate elements in the form of , where x is a placeholder representing the keyword being processed currently, and update Subsequently, store (seq || link x || l w ) into the Stash table; then the user traverses the Stash table to make comparisons one by one. If w = x, read the corresponding ids; d) Update: After several queries, the update operation is triggered and the user evicts all (seq||link x ||l w ); First, the user locally sets the path P(l w ) Create the corresponding free block. If x≠w, for all links corresponding to keyword x x , let l x =Position[x], encryptedlink x And the encrypted link x Randomly store on path P(l x )∩P(l w ) in the block on the left; if x=w, update Position[w] so that P(l' w )∩P(l w )The number of free blocks in the path R(P(l' w )∩P(l w )) is greater than the link corresponding to keyword w w The number of links, then w Encrypt and save the encrypted link w Randomly store in this path and finally update the Residue table.