Multi-keyword private information retrieval method and system based on inverted index

By using inverted indexes and fully homomorphic encryption, the database is divided into two parts: keyword-index set and index-data value. A hash bucket structure is constructed using Gödel encoding and probabilistic batch encoding, which solves the problem of high communication and computational complexity in existing technologies and achieves flexibility and security for multi-keyword queries.

CN121029976APending Publication Date: 2025-11-28UNIV OF JINAN

Patent Information

Application Number
CN202511562852.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing private information retrieval technologies suffer from problems such as high communication and computational overhead when performing multi-keyword queries, inability to effectively support non-primary key queries, and violation of privacy protection principles.

Method used

The database is split into keyword-index set and index-data value parts using an inverted index structure. Hash buckets are constructed using Gödel encoding and probabilistic batch encoding, and ciphertext calculation is performed using fully homomorphic encryption to enable multi-keyword queries.

Benefits of technology

It implements a symmetric PIR in a single-server architecture, supports multi-keyword queries without primary keys, protects the privacy of both the client and the server, reduces communication and computational complexity, and improves query flexibility and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029976A_ABST
    Figure CN121029976A_ABST
Patent Text Reader

Abstract

The invention provides a multi-keyword private information retrieval method and system based on inverted indexes, and belongs to the field of information security and privacy computing. The method comprises the following steps that: a server splits an original database into a reverse index part of a keyword-index set and a key value pair part of an index-data value; mapping the index set into an integer by utilizing Godel coding; processing the data by adopting probability batch coding and binary random linear coding and publishing a hash function; the client maps a multi-keyword query into a bucket by using a hash function to generate a query vector, and the query vector is sent to the server after being subjected to fully homomorphic encryption; the server performs homomorphic calculation on a matching result in a ciphertext state and returns the result; and the client obtains a matching index set after decryption verification, and then initiates a second round of query to obtain final data. According to the method, symmetric privacy protection of multi-keyword non-primary key query in a single-server environment is realized, leakage of query content and database information is effectively prevented, and private information retrieval security is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of information security and privacy computing, and particularly relates to a multi-keyword private information retrieval method and system based on inverted index. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute the prior art.

[0003] Under the tide of high-speed development of Internet technology and full digitalization of personal and institutional data, data has become a core production factor, but the problem of its privacy protection has become increasingly serious. The traditional database query mode has significant security risks: when a user or an institution performs an information retrieval operation on a network platform, its query content (such as medical diagnosis keywords, financial transaction demands, etc.) is extremely easy to be monitored, collected and stored by search engines, data service providers or other third parties. Such information exposure not only directly leads to privacy leakage, but also may derive a series of potential security risks such as data abuse. Especially in the fields of medical health, financial transactions, government services, etc., which are highly sensitive to data, once the query content is leaked, it may cause serious consequences.

[0004] Private information retrieval (PIR) technology is a key solution that has emerged in response to the above needs, and its core goal is to enable the client to securely obtain the required information from a remote database without revealing any query intent to the server, thereby balancing between efficient information acquisition and privacy protection. However, existing private information retrieval usually requires the client to know the exact index position of the target data in advance, which is often impractical in actual query scenarios. Secondly, keyword PIR supports keyword-based queries, but usually requires the keyword to exist as the primary key of the database, and it is difficult to effectively support joint queries of multiple keywords. In addition, existing schemes when processing multi-keyword queries either require the client to perform multiple independent queries and locally calculate the intersection, resulting in significant increase in communication and computation overhead, or require the server to expose additional database information, which violates the principle of symmetric privacy protection. SUMMARY

[0005] To overcome the deficiencies of the above existing technologies, the present application provides a multi-keyword private information retrieval method and system based on inverted index, aiming to realize symmetric PIR under a single server architecture, support non-primary key multi-keyword queries, and protect the privacy of both the client query and the server data.

[0006] To achieve the above purpose, one or more embodiments of the present application provide the following technical solutions: The application provides a multi-keyword private information retrieval method based on an inverted index. A multi-keyword private information retrieval method based on an inverted index comprises the following steps: Obtaining a server database and a client query set; Dividing the database into an inverted index structure database and a key-value pair structure database by using a server; encoding the inverted index structure database by using a Godel code to obtain an encoded inverted index structure database; Processing data of the key-value pair structure database and the encoded inverted index structure database by using a probability batch encoding and a binary random linear encoding, constructing a hash bucket structure and publishing a hash function; Mapping elements in the client query set to buckets by using the hash function to generate a query vector, performing full-homomorphic encryption on the query vector and sending the query vector to the server; Homomorphically calculating a matching result in a ciphertext state by the server and returning the matching result to the client; restoring a matching index set after decryption and verification by the client, and obtaining final query data by initiating a second round of query.

[0007] As a further technical solution, the server database is:

[0008] wherein, is an index of the i th data item, is a related keyword set of the i th data item is a data value of the i th data item. , is an index of the i th data item, is a related keyword set of the i th data item The client query set is:

[0009] wherein, m is the number of multi-keywords held by the client.

[0010] As a further technical solution, the step of encoding the inverted index structure database by using the Godel code to obtain the encoded inverted index structure database comprises the following steps: Mapping all indexes in the inverted index structure database to different prime numbers, and sequentially establishing mapping from the smallest prime number; For each mapped index, calculating a Godel code value, combining the Godel code value with a keyword in the inverted index structure database, and obtaining the encoded inverted index structure database.

[0011] ​As a further technical solution, the probability batch encoding and binary random linear encoding are used to process the key-value pair structure database and the coded inverted index structure database, construct a hash bucket structure and publish a hash function, including: The probability batch encoding is performed on the key-value pair structure database and the coded inverted index structure database respectively, a hash bucket structure is constructed, and a first group of hash functions and corresponding bucket number parameters are published; A predefined unit element is inserted into the coded inverted index structure database, and binary random linear encoding is performed on the data in each bucket using a second group of hash functions to generate a first encoding result and a second encoding result.

[0012] As a further technical solution, the client uses a hash function to map elements in a client query set to a bucket to generate a query vector, and sends the homomorphically encrypted query vector to the server, including: The client uses the first group of hash functions to map elements in the client query set to corresponding buckets using a cuckoo hashing strategy, and generates a first query vector for each bucket using the second group of hash functions; the client encrypts the first query vector using a homomorphic encryption algorithm to obtain a first ciphertext vector, and sends the first ciphertext vector to the server.

[0013] As a further technical solution, the server homomorphically computes the matching result in a ciphertext state and returns it to the client; the client decrypts and verifies the matching index set after restoration, including: The server performs homomorphic computation on the first encoding result, the second encoding result and the first ciphertext vector to obtain a first intermediate result and a second intermediate result, and returns the first intermediate result and the second intermediate result to the client; The client decrypts the second intermediate result and verifies its validity; if the verification is passed, the first intermediate result is decrypted and the multi-keyword corresponding index set is restored through the Gödel encoding; otherwise, the protocol is terminated.

[0014] As a further technical solution, the final query data is obtained by initiating a second round of query, including: The client uses the index set and the first group of hash functions to generate a second query vector for the key-value pair structure database using a cuckoo hashing strategy; the client encrypts the second query vector using a homomorphic encryption algorithm to obtain a second ciphertext vector, and sends the second ciphertext vector to the server; the server calculates the inner product of the data in each bucket of the key-value pair structure database and the second ciphertext vector to obtain a result vector, and returns the result vector to the client; the client decrypts the result vector to obtain the final query data.

[0015] A second aspect of the present invention provides a multi-keyword private information retrieval system based on an inverted index.

[0016] A multi-keyword private information retrieval system based on an inverted index includes: The data acquisition module is configured to acquire the server database and client query sets. The server segmentation and encoding module is configured to: use the server to segment the server database into an inverted index structure database and a key-value pair structure database; and use Gödel encoding to encode the inverted index structure database to obtain the encoded inverted index structure database. The data processing module is configured to: perform data processing on the key-value pair structure database and the encoded inverted index structure database using probabilistic batch coding and binary random linear coding, construct a hash bucket structure and publish the hash function; The query vector generation and encryption module is configured as follows: the client uses a hash function to map the elements in the client query set to buckets to generate a query vector, and then sends the query vector to the server after performing fully homomorphic encryption. The query data retrieval module is configured as follows: the server performs homomorphic calculation of the matching result in encrypted state and returns it to the client; after the client decrypts and verifies the result, it restores the matching index set and retrieves the final query data by initiating a second round of query.

[0017] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of a multi-keyword private information retrieval method based on an inverted index as described in the first aspect of the present invention.

[0018] A fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of a multi-keyword private information retrieval method based on an inverted index as described in the first aspect of the present invention.

[0019] The above one or more technical solutions have the following beneficial effects: (1) This invention enables symmetric private information retrieval supporting multiple keywords and non-primary key conditions under a single-server computing security model. It not only protects the client's query keyword set from being known by the server, but also prevents the client from obtaining any additional database information besides the query results from the server response, achieving two-way privacy protection and greatly expanding the practicality and security of PIR technology in complex query scenarios. By introducing an inverted index structure and splitting the database into two parts—keyword-index set and index-data value—the query keywords do not need to be unique identifiers of data items. Users can freely combine any attribute tags of data items for queries, which is more suitable for actual application scenarios with multi-attribute tag data such as medical records, literature retrieval, and e-commerce products, significantly improving the flexibility of queries and the applicability of the solution.

[0020] (2) This invention uses Gödel encoding to encode the index set into an integer. By utilizing the multiplicative property of integers, the intersection operation of the set is transformed into a multiplication operation under the homomorphic encryption domain, enabling the server to efficiently calculate the intersection of multiple keyword corresponding index sets on the ciphertext, thus providing a mathematical basis for the core multi-keyword matching. By processing the database through probabilistic batch encoding technology and combining it with the client's cuckoo hash strategy, multiple keyword queries are efficiently mapped to fixed buckets and batch processed, significantly reducing the number of communication rounds and computational complexity, and achieving sublinear communication overhead. At the same time, the introduction of binary random linear encoding ensures that each query can be correctly mapped and verified, and its randomness guarantees the high probability full rank characteristic of the encoding matrix, providing a guarantee for the server to perform correct linear operations on the ciphertext and effectively verifying the legality of the client's query keywords. Based on the homomorphic computation capability of fully homomorphic encryption, the entire query matching process is carried out in the ciphertext state, ensuring that the server cannot obtain any plaintext information, and achieving a strong security core cryptographic guarantee.

[0021] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0022] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0023] Figure 1 This is a flowchart of the method in the first embodiment.

[0024] Figure 2 This is a schematic diagram illustrating the experimental results of the number of database entries and total online time in the first embodiment.

[0025] Figure 3 This is a schematic diagram illustrating the experimental results of different round numbers and communication volumes in the first embodiment.

[0026] Figure 4 This is a system structure diagram of the second embodiment. Detailed Implementation

[0027] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0028] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0029] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0030] Example 1 This embodiment discloses a multi-keyword private information retrieval method based on an inverted index; like Figure 1 As shown, a multi-keyword private information retrieval method based on an inverted index includes: Step S1: Obtain the server database and client query set; Step S2: The database is divided into an inverted index structure database and a key-value pair structure database using the server; the inverted index structure database is encoded using Gödel encoding to obtain the encoded inverted index structure database. Step S3: Use probabilistic batch coding and binary random linear coding to process the key-value pair structure database and the encoded inverted index structure database, construct a hash bucket structure and publish the hash function; Step S4: The client uses a hash function to map the elements in the client query set to buckets to generate a query vector, and then sends the query vector to the server after performing fully homomorphic encryption. In step S5, the server homomorphically calculates the matching result in encrypted state and returns it to the client; after the client decrypts and verifies, it restores the matching index set and obtains the final query data by initiating a second round of query.

[0031] Specifically, it also includes the following: Step S1: Obtain the server database and client query set.

[0032] In the process of retrieving the server database and the client query set, the server database... for:

[0033] in, It is the first Index of data items It is the first A set of keywords related to the data item. It is the first The data value of the data item .

[0034] Client query collection for:

[0035] Where m represents the number of keywords held by the client.

[0036] Step S2: The database is divided into an inverted index structure database and a key-value pair structure database using the server; the inverted index structure database is encoded using Gödel encoding to obtain the encoded inverted index structure database.

[0037] Step S2: The database is divided into an inverted index structure database and a key-value pair structure database using the server; Gödel encoding is used to map each index set in the inverted index structure database to integer values, resulting in an encoded inverted index structure database; probabilistic batch encoding is performed on the key-value pair structure database and the encoded inverted index structure database respectively to construct a hash bucket structure, and the first set of hash functions and the corresponding bucket number parameters are published; a predefined unit element is inserted into the encoded inverted index structure database, and a second set of hash functions is used to perform binary random linear encoding on the data in each bucket to generate a first encoding result and a second encoding result.

[0038] Specifically, step S2 also includes the following: Step S21: Use the server to divide the database into an inverted index structure database and a key-value pair structure database.

[0039] Use the server to store the server database It is decomposed into an inverted index structure database and a key-value pair structure database, where the inverted index structure database for:

[0040] in, It is the number of non-repeating keywords. It is the first in the inverted index structure database One keyword, Keywords The corresponding inverted index set.

[0041] Key-value pair structure database for: 。

[0042] Step S22: Use Gödel encoding to map each index set in the inverted index structure database to integer values ​​to obtain the encoded inverted index structure database.

[0043] The server first indexes all Mapping to different prime numbers, to facilitate calculations, mappings are established sequentially from the smallest prime number. The purpose is to implement set operations using encoding under multiplication. Then, for... Each index in Calculate the Gödel code value:

[0044] in, Gödel encoded values; Let j be the j-th prime number; if ,but ,otherwise 。

[0045] Compare Gödel codes with inverted index structure database By combining keywords, an encoded inverted index structure database is obtained. :

[0046] in, Keywords in an inverted index database.

[0047] Step S3: Use probabilistic batch coding and binary random linear coding to process the key-value pair structure database and the encoded inverted index structure database, construct a hash bucket structure and publish the hash function.

[0048] In step S3, probabilistic batch coding is an encoding method that supports batch queries. Its principle is that on the server side, multiple independent hash functions are used with a common hash strategy to map and copy all data items into multiple buckets. The client uses the same hash function with a cuckoo hash strategy to map multiple keywords to be queried into different buckets, and then executes individual queries in all sub-buckets to complete the overall batch query. Binary random linear coding is expressed as follows: if each row of a matrix generates a binary short-band vector based on randomness, such a matrix has a high probability of being full rank. This, combined with solving linear equation systems, can be used to encode databases to obtain data structures that support keyword queries. Specifically, in this embodiment, it includes: Step S31: Perform probabilistic batch encoding on the key-value pair structure database and the encoded inverted index structure database respectively, construct a hash bucket structure, and publish the first set of hash functions and corresponding bucket number parameters.

[0049] For the encoded inverted index structure database Establish For each bucket, select three independent hash functions to form the first group of hash functions. , Adjusting the output of a hash function through modulo operations is... For each keyword Calculate the values ​​of the three hash functions Then Copy and save them to the first one. Inside the bucket.

[0050] For key-value pair structured databases Establish Each bucket, and the output of the hash function is adjusted through a modulo operation. Index for each data item calculate Then Copy and save them to the first one. The data is stored in the bucket. Finally, the server publishes the hash function used in the first set. And the number of buckets and .

[0051] Step S32: Insert predefined unit elements into the encoded inverted index structure database, and use the second set of hash functions to perform binary random linear encoding on the data in each bucket to generate the first encoding result and the second encoding result.

[0052] The server calculates the unit element based on Gödel encoding. ,in This is the agreed-upon empty keyword, and 1 is the Gödel code value for the entire set. This is used in the encoded inverted index structure database. Insert a unit element into each of the generated sub-buckets .

[0053] Then for Each bucket, respectively Perform binary random linear encoding, where, To be Input to hash function The obtained function value. Based on the bucketing results of step S31, the following is adopted: To represent the encoded inverted index structure database The keywords and their corresponding Gödel codes. This indicates which bucket it is in. This represents the position inside the bucket, where That is the maximum capacity of the bucket.

[0054] The server selects a hash function. Output Used for partitioning, where It's the size of the block; the server then selects an independent hash function. The output is Used to generate random binary values; finally, an independent hash function is selected. Used for calculation The authentication value.

[0055] Specifically, regarding Gödel encoding values Taking linear encoding as an example, the server for For each bucket, first calculate This involves dividing the data into blocks. Then, for each block, each element within that block... Cascade integer calculations , Generate row vectors, with the starting positions of the vectors determined by... To generate randomly, where A parameter less than 1. Within each block, the determinant of the row vectors is obtained. Each block... The values ​​and determinants together form a system of linear equations, which are then solved using Gaussian elimination. Finally, each bucket outputs... The solution to a system of linear equations is expressed as: , .right Similarly, the linear encoding method outputs the solution to the linear equation system using... express.

[0056] Will be and The two sets of codes are denoted as the first encoding result and the second encoding result, respectively. Wherein:

[0057]

[0058] in, This is the first encoding result; This is the second encoding result.

[0059] In step S4, the client uses a hash function to map the elements in the client query set to buckets to generate a query vector, and then sends the query vector to the server after performing fully homomorphic encryption.

[0060] Specifically, in step S4, the first query interaction is performed using the client, including: Step S41: The client uses the first set of hash functions and adopts the cuckoo hash strategy to map the elements in the client query set to the corresponding buckets, and uses the second set of hash functions to generate a first query vector for each bucket. The client uses the first set of hash functions Calculate client query set Each keyword is assigned three hash values, and then the Cuckoo Hash Strategy is used to distribute all keywords into separate buckets. For any keywords not assigned to a bucket, an empty keyword is assigned. .

[0061] use This represents the keywords assigned to each bucket, where .for ,calculate To determine the location of the blocks, a vector consisting of 0s and 1s is generated. ,in The position is 1 and the other positions are 0. Then calculate. To generate vectors Based on vectors sum vector Obtain the first query vector for all buckets:

[0062] in, This is the first query vector.

[0063] In step S42, the client encrypts the first query vector using a fully homomorphic encryption algorithm to obtain a first ciphertext vector, and sends the first ciphertext vector to the server.

[0064] The client uses fully homomorphic encryption to query the vector; in this embodiment, BFV homomorphic encryption is employed, and the plaintext field is represented as follows. It is a polynomial ring, where T is the modulus of the coefficients and N is the degree of the polynomial. The ciphertext field is... It is also a polynomial ring, where Q is the modulus of the coefficients. The BFV homomorphic encryption includes: a key generation algorithm. Homomorphic encryption algorithm Homomorphic decryption algorithm .

[0065] Specifically, in this embodiment, the public key is first obtained by running a key generation algorithm on the client side. and private key Among them, the key generation algorithm Used to output the public key and private key ,in From the ring Randomly selected elements from the data. It is a polynomial whose degree coefficients are randomly selected from a small discrete distribution. It is an error polynomial randomly sampled from a discrete Gaussian distribution.

[0066] Use the query vector in this example as plaintext input. Homomorphic encryption algorithm Based on input plaintext and public key Output ciphertext ,in Random sampling from a discrete Gaussian distribution .

[0067] In this embodiment, the public key will be input. and the first query vector Running the homomorphic encryption algorithm yields a two-dimensional first ciphertext vector. ,in Then Send to the server.

[0068] In step S5, the server homomorphically calculates the matching result in encrypted state and returns it to the client; after the client decrypts and verifies, it restores the matching index set and obtains the final query data by initiating a second round of query.

[0069] Step S51: The server performs homomorphic computation with the first encoding result, the second encoding result and the first ciphertext vector to obtain the first intermediate result and the second intermediate result, and returns the first intermediate result and the second intermediate result to the client.

[0070] The server will assign the first encoded result to each bucket. Second coding result With the first ciphertext vector To perform homomorphic computation, taking the i-th bucket as an example, first input... calculate The data respectively compared with the data in the first encoding result inner product A total of The inner product result ,in It is an array of vectors. Then, the second inner product is calculated. >. For the second encoding result Perform the same calculation steps to obtain the query vector and Calculation results .

[0071] Finally, calculate the iteration from the first bucket to the... 1 bucket, obtain the first intermediate result Second intermediate result The first and second intermediate results are returned to the client.

[0072] Step S52: The client decrypts the second intermediate result and verifies its validity; if the verification is successful, the first intermediate result is decrypted and the index set corresponding to the multiple keywords is restored by reverse engineering using Gödel encoding; otherwise, the protocol is terminated.

[0073] The client first uses a homomorphic decryption algorithm Decrypting the second intermediate result The entire content, including the plaintext output of the homomorphic decryption algorithm. in That is, first use the private key The intermediate value is obtained by linearly combining the ciphertext components. Then, the original plaintext is restored through scaling and rounding.

[0074] In this embodiment, a homomorphic decryption algorithm is used. Decryption, as shown below:

[0075] in, To The decryption result represents the hash value of the keyword; Using the above algorithm, From 1 to get A plain text, in comparison and The protocol is terminated if all keys are equal, indicating that the client's input keywords exceeded the server's range. If all keys are equal, the first intermediate result is decrypted. .in, To The decoding result represents the product of multiple corresponding Gödel codes. An empty set is initialized. Then Calculate from 1 to n If the value is 0, then use Gödel encoding to reverse-engineer the index set. .

[0076] Step S53, based on the index set obtained in step S20, performs a second query interaction using the client, specifically including: Step S531: The client uses the index set and the first set of hash functions to generate a second query vector for the key-value pair structure database using the Cuckoo Hash strategy.

[0077] The client uses the first set of hash functions Calculate the index set All index values ​​are used in the database, and the Cuckoo Hash algorithm is employed to distribute these index values ​​to the key-value pair structure database. Within each bucket, a vector of 0s and 1s is generated, with the index set to 1 at the specified position and 0 at the rest, resulting in a total of [number missing]. A vector, represented as the second query vector. .

[0078] Step S532: The client uses a fully homomorphic encryption algorithm to encrypt the second query vector to obtain the second ciphertext vector, and sends the second ciphertext vector to the server. Client inputs public key Second query vector Running the homomorphic encryption algorithm will generate the second query vector. Encryption yields a two-dimensional second ciphertext vector. ,in Then the second ciphertext vector Send to the server.

[0079] Step S533: The server calculates the inner product of each bucket of data in the key-value pair structure database with the second ciphertext vector to obtain the result vector, and returns the result vector to the client. Let key-value pair structure database The bucket in the middle is Calculated by server Each bucket in the middle and dot product of median vectors The result vector is obtained. .

[0080] Step S534: Decrypt the result vector using the client to obtain the final query data.

[0081] The client will From 1 to Decrypt to obtain query data It then outputs the decryption result, thus obtaining the final query data. .

[0082] Furthermore, the technical solution of this invention will be described in detail in conjunction with specific application scenarios. The performance test of this invention was conducted on a virtual database, where the content of the database was randomly generated. The experimental results are as follows... Figure 2 and Figure 3 As shown.

[0083] First, assume the server is a product provider holding a product database list. Clients want to search for products of interest within this list but don't want their query to be known by the server. Therefore, a multi-keyword private information retrieval scheme is needed. Multiple keywords allow for flexible client queries, while private information retrieval ensures the client's query remains undetected by the server. Specifically, the system first gathers input data from both parties: the server provides the product list, and the client provides multiple keywords for the query. Then, the product list is divided into an inverted index table and a table mapping indexes to values, and hash-bucketed. Gödel codes and binary random linear codes are then calculated on the inverted index tables of each bucket. After the server's pre-calculation, its parameters are published. The client uses the calculated query vector, encrypts it using a homomorphic encryption scheme, and sends it to the server. The server receives the vector, performs a homomorphic inner product calculation, and returns the result to the client. After verification, the client decrypts the vector, restores the index values, generates a second query vector, and the server performs the same homomorphic inner product calculation, returning the final result. The client then decrypts the vector to obtain the final result. Figure 2 and Figure 3 In the test experiment, the runtime and communication overhead increased with the increase of parameters (database size, number of keywords), and when the number of database entries reached... When the number of keywords on the client reaches 8, the total online time is close to 400 seconds and the total communication volume is close to 8000KB.

[0084] Example 2 This embodiment discloses a multi-keyword private information retrieval system based on an inverted index; like Figure 4 As shown, a multi-keyword private information retrieval system based on an inverted index includes: The data acquisition module is configured to acquire the server database and client query sets. The server segmentation and encoding module is configured to: use the server to segment the server database into an inverted index structure database and a key-value pair structure database; and use Gödel encoding to encode the inverted index structure database to obtain the encoded inverted index structure database. The data processing module is configured to: perform data processing on the key-value pair structure database and the encoded inverted index structure database using probabilistic batch coding and binary random linear coding, construct a hash bucket structure and publish the hash function; The query vector generation and encryption module is configured as follows: the client uses a hash function to map the elements in the client query set to buckets to generate a query vector, and then sends the query vector to the server after performing fully homomorphic encryption. The query data retrieval module is configured as follows: the server performs homomorphic calculation of the matching result in encrypted state and returns it to the client; after the client decrypts and verifies the result, it restores the matching index set and retrieves the final query data by initiating a second round of query.

[0085] Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.

[0086] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a multi-keyword private information retrieval method based on an inverted index as described in Example 1.

[0087] Example 4 The purpose of this embodiment is to provide an electronic device.

[0088] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of a multi-keyword private information retrieval method based on an inverted index as described in Embodiment 1.

[0089] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0090] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0091] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for retrieving private information using multiple keywords based on an inverted index, characterized in that, include: Retrieve the server database and client query set; The server database is divided into an inverted index structure database and a key-value pair structure database using the server. The inverted index structure database is encoded using Gödel coding to obtain the encoded inverted index structure database; The key-value pair structure database and the encoded inverted index structure database are processed using probabilistic batch coding and binary random linear coding to construct a hash bucket structure and publish the hash function. The client uses a hash function to map the elements in the client query set to buckets to generate a query vector, and then sends the query vector to the server after performing fully homomorphic encryption. The server performs homomorphic computation of the matching result in encrypted form and returns it to the client. After the client decrypts and verifies the data, it restores the matching index set and then initiates a second round of queries to obtain the final query data.

2. The method for retrieving private information using multiple keywords based on an inverted index as described in claim 1, characterized in that, The server database is: in, It is the first Index of data items It is the first A set of keywords related to the data item. It is the first The data value of the data item; The client query set is: Where m represents the number of keywords held by the client.

3. The method for retrieving private information using multiple keywords based on an inverted index as described in claim 1, characterized in that, The process of encoding the inverted index structure database using Gödel coding to obtain the encoded inverted index structure database includes: Map all indexes in the inverted index structure database to different prime numbers, starting with the smallest prime number. For each mapped index, a Gödel code value is calculated, and the Gödel code value is combined with the keywords in the inverted index structure database to obtain the encoded inverted index structure database.

4. The method for retrieving private information using multiple keywords based on an inverted index as described in claim 1, characterized in that, The key-value pair structure database and the encoded inverted index structure database are processed using probabilistic batch coding and binary random linear coding to construct a hash bucket structure and publish the hash function, including: The key-value pair structure database and the encoded inverted index structure database are subjected to probabilistic batch encoding to construct a hash bucket structure, and the first set of hash functions and corresponding bucket number parameters are published. Insert predefined unit elements into the encoded inverted index structure database, and use the second set of hash functions to perform binary random linear encoding on the data in each bucket to generate the first encoding result and the second encoding result.

5. The method for retrieving private information using multiple keywords based on an inverted index as described in claim 1, characterized in that, The client uses a hash function to map elements in the client's query set to buckets to generate a query vector. This query vector is then fully homomorphically encrypted and sent to the server, including: The client uses the first set of hash functions and the Cuckoo Hash strategy to map the elements in the client query set to the corresponding buckets. It then uses the second set of hash functions to generate a first query vector for each bucket. The client uses a fully homomorphic encryption algorithm to encrypt the first query vector to obtain a first ciphertext vector, and sends the first ciphertext vector to the server.

6. The method for retrieving private information using multiple keywords based on an inverted index as described in claim 1, characterized in that, The server performs homomorphic computation of the matching result in encrypted form and returns it to the client. After decryption and verification by the client, the matching index set is restored, including: The server performs homomorphic computation using the first encoding result, the second encoding result, and the first ciphertext vector to obtain the first intermediate result and the second intermediate result, and then returns the first intermediate result and the second intermediate result to the client. The client decrypts the second intermediate result and verifies its validity. If the verification passes, the first intermediate result is decrypted and the index set corresponding to the multiple keywords is restored by reverse engineering using Gödel encoding; otherwise, the protocol is terminated.

7. The method for retrieving private information using multiple keywords based on an inverted index as described in claim 1, characterized in that, The final query data is obtained by initiating a second round of queries, including: The client uses the index set and the first set of hash functions to generate a second query vector for the key-value pair structure database using a cuckoo hashing strategy. The client encrypts the second query vector using a fully homomorphic encryption algorithm to obtain a second ciphertext vector, and sends the second ciphertext vector to the server. The server calculates the inner product of each bucket of data in the key-value pair structure database with the second ciphertext vector to obtain a result vector, and returns the result vector to the client. The client decrypts the result vector to obtain the final query data.

8. A multi-keyword private information retrieval system based on an inverted index, characterized in that, include: The data acquisition module is configured to acquire the server database and client query sets. The server segmentation and encoding module is configured to: use the server to segment the server database into an inverted index structure database and a key-value pair structure database; and use Gödel encoding to encode the inverted index structure database to obtain the encoded inverted index structure database. The data processing module is configured to: perform data processing on the key-value pair structure database and the encoded inverted index structure database using probabilistic batch coding and binary random linear coding, construct a hash bucket structure and publish the hash function; The query vector generation and encryption module is configured as follows: the client uses a hash function to map the elements in the client query set to buckets to generate a query vector, and then sends the query vector to the server after performing fully homomorphic encryption. The query data retrieval module is configured as follows: the server performs homomorphic calculation of the matching result in encrypted state and returns it to the client; after the client decrypts and verifies the result, it restores the matching index set and retrieves the final query data by initiating a second round of query.

9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the multi-keyword private information retrieval method based on an inverted index as described in any one of claims 1-7.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the multi-keyword private information retrieval method based on inverted index as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Homomorphic keyword privacy information retrieval method and system based on polynomial link

    CN119918088A

  • Batch privacy information retrieval method and device

    WO2024239204A1

Cited By

  • Precision-controllable lattice-based homomorphic encryption inner product data similarity retrieval method and system

    CN121501866A