A method and a client for implementing a constructed confusion set

By adopting an interchangeable encrypting/decrypting algorithm and a method of constructing an obfuscation set between the client and the server, the problem of difficulty in realizing efficient data retrieval while protecting data privacy is solved, and effective protection of client privacy and efficient data retrieval are achieved.

CN115801233BActive Publication Date: 2025-05-27ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211229255.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-09
Publication Date
2025-05-27
Estimated Expiration
2042-10-09

AI Technical Summary

Technical Problem

Existing privacy computing technologies are difficult to achieve efficient data retrieval and query while protecting data privacy. Especially in the interaction between the client and the server, how to ensure that the privacy of the client is not leaked is a challenge.

Method used

By using an interchangeable encrypted/decrypted algorithm between the client and the server, the client receives the encrypted query base sent by the server and obtains sensitive fields encrypted by the server through interaction. The client retrieves the identification set of matching records in the query base and constructs an obfuscation set to protect privacy.

Benefits of technology

It realizes efficient data retrieval and query without revealing client privacy, ensuring the server-side privacy protection of the database and avoiding privacy leakage caused by unreasonable confusion set construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115801233B_ABST
    Figure CN115801233B_ABST
Patent Text Reader

Abstract

One or more embodiments of this specification provide a method and a client for implementing a construction of a confusion set. In the method, the client sends a sensitive field encrypted by itself to the server, and obtains the same sensitive field encrypted by the server through interaction with the server; the client retrieves in a query base according to the sensitive field encrypted by the server to obtain a first identification set of matching records; the client selects at least one point in the query base except for the ID field and the fields of interest, uses the selected at least one point as a retrieval condition, constructs a retrieval statement according to this retrieval condition and the original fields of interest, and executes the retrieval on the query base to obtain a second identification set of matching records; the client constructs a confusion set according to the first identification set and the second identification set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification belong to the field of privacy computing technology, and in particular, relate to a method and client for constructing a confusion set. Background Art

[0002] Privacy-preserving computing is a collection of technologies that enable data analysis and computation without leaking the data itself, making the data available but invisible. Privacy-preserving computing can transform and unlock the value of data while fully protecting data and privacy.

[0003] Currently, mainstream technologies for achieving privacy-preserving computing fall into three main areas: the first is cryptography-based privacy computing, represented by Secure Multi-Party Computation (SMPC); the second is the integration of AI and privacy-preserving technologies, represented by Federated Learning (FL); and the third is trusted hardware-based confidential computing (CC), represented by Trusted Execution Environment (TEE). Differential Privacy (DP) also protects the computational results, not the computational process itself. Federated Learning, Secure Multi-Party Computation, and Confidential Computing protect the computational process and its intermediate results.

[0004] The first category of multi-party secure computation includes four basic technologies: garbled circuits (GC), secret sharing, oblivious transfer, and homomorphic encryption (HE). Homomorphic encryption is a special encryption algorithm that performs calculations directly on ciphertext, producing the same result as the decrypted plaintext. It includes semi-homomorphic encryption (PHE) and fully homomorphic encryption (FHE).

[0005] Secure multi-party computation, with its solid theoretical foundation, provides privacy protection for secret input data, ensuring the security of the privacy-preserving computation process. Currently, there are two main technical approaches to secure multi-party computation: general-purpose secure multi-party computation and problem-specific secure multi-party computation. The former can solve a wide range of computational problems, but this "one-size-fits-all" approach often involves a large system and high overhead. The latter, designed specifically for specific problems, uses dedicated protocols such as Private Set Intersection (PSI) and Privacy Information Retrieval (PIR). While these protocols can often achieve computational results at a lower cost than general-purpose secure multi-party computation protocols, they require careful design by domain experts tailored to the application scenario, are generally not applicable to general scenarios, and are expensive to design.

[0006] Private set intersection allows two parties to obtain the intersection of their data without revealing any additional information. Additional information refers to any information beyond the intersection of the two parties' data. Private set intersection is very useful in real-world scenarios, such as data alignment in vertical federated learning and friend discovery through address books in social apps.

[0007] Privacy information retrieval is a method by which clients retrieve information from a database. During the retrieval process, the querying party hides the target identifier, and the data service provider provides matching query results without being able to identify the specific query object. Summary of the Invention

[0008] The purpose of this specification is to provide a method and client for constructing a confusion set, including:

[0009] A method for constructing a confusion set, wherein a client receives a query base sent by a server, wherein the query base is obtained by encrypting the query base from a database; the encryption / decryption performed by the client and the server on the same target uses an encryption / decryption algorithm with an interchangeable sequence;

[0010] The client sends the sensitive field encrypted by itself to the server, and obtains the same sensitive field encrypted by the server through interaction with the server;

[0011] The client searches the query base according to the sensitive field encrypted by the server to obtain a first identification set of matching records;

[0012] The client selects at least one point in the query base other than the ID field and the field of interest, uses the selected at least one point as a search condition, constructs a search statement based on the search condition and the original field of interest, and performs the search on the query base to obtain a second set of identifiers of matching records;

[0013] The client constructs a confusion set according to the first identification set and the second identification set.

[0014] A client for constructing a confusion set, wherein the encryption / decryption performed by the client and the server on the same target uses an encryption / decryption algorithm with an interchangeable order, and:

[0015] The client is configured with a query base, which is obtained by encrypting the database by the server;

[0016] The client sends a sensitive field encrypted by itself to the server, and obtains the same sensitive field encrypted by the server through interaction with the server; searches the query base based on the sensitive field encrypted by the server to obtain a first identification set of matching records; selects at least one point in the query base other than the ID field and the field of interest, uses the selected at least one point as a search condition, constructs a search statement based on the search condition and the original field of interest, and executes the search on the query base to obtain a second identification set of matching records; and constructs a confusion set based on the first identification set and the second identification set.

[0017] A client for constructing a confusion set, comprising:

[0018] processor,

[0019] The memory stores a program, wherein when the processor executes the program, the above method is performed.

[0020] A storage medium is used to store a program, wherein when the program is executed, the client executes the above method.

[0021] In the above embodiment, by constructing a search statement and performing a search on the query base, a second identification set is obtained, and then a confusion set is constructed, which can make the confusion set more reasonable, thereby avoiding client privacy leakage due to unreasonable confusion set construction. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0023] Figure 1 This is a flow chart of an embodiment;

[0024] Figure 2 This is a flow chart of an embodiment;

[0025] Figure 3It is a flowchart of an embodiment. DETAILED DESCRIPTION

[0026] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. All other embodiments derived by those skilled in the art based on the embodiments in this specification without creative effort shall fall within the scope of protection of this specification.

[0027] As mentioned previously, PIR is a method for clients to retrieve information from a database. The PIR scheme, proposed by Chor B et al. in 1995, aims to protect user query privacy. The primary goal of the PIR scheme is to ensure that queries submitted by users to the database server are completed without leaking the user's private information. Specifically, during the retrieval process, the server remains unaware of the user's specific query information and retrieved data items.

[0028] Application scenarios of privacy information retrieval include:

[0029] i. If a patient wants to search for medications for their disease through the medical system, using the disease name as the query condition will cause the medical system to know that the patient may have this disease, thereby leaking the patient's privacy. This can be avoided by using private information query.

[0030] ii. During the domain name and trademark application process, users must first submit their domain name or trademark information to the relevant database to check whether it already exists. However, some users may not want the service provider to know their application name so that they can register it first.

[0031] iii. In the securities market, a user wants to query stock information, but cannot disclose the stocks he is interested in to the service provider, which would affect the stock price and his own preferences.

[0032] A simple implementation involves the database sending all data to the client, but this fails to protect the database's security, meaning that the server's privacy cannot be guaranteed. PIRs that guarantee both client and database privacy are called symmetric PIRs (SPIRs). PIRs that guarantee the privacy of either the client or the database are called asymmetrical PIRs (APIRs). Based on the number of database replicas, PIRs can be categorized as multi-replica PIRs and single-replica PIRs. Multi-replica PIR protocols require that multiple database replicas cannot collude, a requirement difficult to achieve in real-world scenarios. Therefore, single-replica PIRs are more commonly considered. Single-replica PIRs can only achieve computational security (CPIRs). Most PIR schemes assume that the client knows the exact bit in the database they wish to retrieve (a single bit). However, in real-world scenarios, clients often search based on keywords (without knowing the specific location of the keyword in the database) and wish to retrieve strings (multi-bit strings). In summary, a practical PIR typically needs to simultaneously meet multiple requirements, such as symmetry, single copy, keyword search, and string return, while also striking a balance between computational and communication efficiency. Cryptographic techniques such as homomorphic encryption, oblivious transfer (OT), and one-way trapdoor functions can partially or completely meet these requirements.

[0033] This specification provides an embodiment of a method for implementing private information retrieval.

[0034] In this embodiment, the server may encrypt the database in advance to obtain a query base, and then send the query base to the client.

[0035] Generally, the server has a local database that can be queried by the client. For example, the server's local database is as follows:

[0036]

[0037]

[0038] Table 1. Databases on the server

[0039] In the example of Table 1 above, there are four fields: ID, Name, Age, and Native_place. For example, there are 10 records with id_0, ..., id_9, and each row is a record. id_0, ..., id_9 are identifiers for each row.

[0040] To allow clients to perform searches without exposing the server's privacy, the server can encrypt the database to obtain a query basis. This encryption can use RSA (a widely used asymmetric encryption algorithm proposed by Ronald Rivest, Adi Shamir, and Leonard Adleman in 1977) or ECC (Elliptical Curve Cryptography). Specifically, the server can encrypt the data using the RSA private key / ECC private key α. This means that each field (i.e., the data in each cell) other than the ID column is encrypted using the RSA private key / ECC private key α.

[0041] When using the ECC encryption and decryption algorithm, specifically, the server can generate and properly store a secret value α, which is also the ECC private key. In addition, the server can use a hash function to convert the value of the name field into a point on the elliptic curve, which can be expressed as Hash(C) or H(C).

[0042] Due to the operational properties of scalar multiplication on elliptic curves, it is easy to calculate Q = kP for a point P on an elliptic curve and an integer k, and the result Q is also a point on the elliptic curve; conversely, if a point pair P and Q on the elliptic curve is known, it is difficult to find the value of k that makes the equation Q = kP valid.

[0043] Here, α·H(C) is easily calculated using scalar multiplication on the elliptic curve, but knowing the result of α·H(C) and H(C) makes it difficult to deduce the value of α. If the value of α is difficult to determine, knowing the result of α·H(C) also makes it difficult to determine the value of H(C).

[0044] Then, the database encrypted by the server using the secret value α is as follows:

[0045]

[0046]

[0047] Table 2. Query base encrypted by ECC private key on the server

[0048] It should be noted that the above hash function can not only convert the original input into an output of fixed length and format, but also convert the output into the x-axis coordinate of a point on an elliptic curve. For example, using an elliptic curve such as curve25519, any 256-bit data can be used as a legal x-axis coordinate on this elliptic curve. Accordingly, sha256 or sha3-256 can be used, or 256 bits can be intercepted from the results of sha384, sha512, or sha3-384, sha3-512. More broadly speaking, any hash value (not limited to hash results of 256 bits) can be modulo the order of the elliptic curve, and the product of the modulo result and the generator point multiplication (scalar multiplication) is a point on the elliptic curve.

[0049] The server can then send the query base to the client that needs to perform the search. In one approach, the server can send the query base directly to the client, such as directly to the client's device or to a proxy server of the client. In another approach, the server can publish the query base on a Uniform Resource Locator (URL), and the client can then retrieve the query base from the URL.

[0050] Correspondingly, the client may receive the query base and save the received query base locally.

[0051] Similarly, when using RSA, the server can generate and properly store a secret value α, which is the RSA private key. In addition, the server can use a hash function to convert the value of the name field into a point on the elliptic curve, which can be expressed as Hash(C) or H(C).

[0052] According to the properties of modular exponentiation, given a secret value α, for a large prime number q and base g, calculate p=g α mod q is easy; conversely, if we know p, q, and the base g, we can solve for p = g α The value of α that makes the equation true mod q is difficult to find. The base g is also called a primitive root.

[0053] Here, (H(C)) is calculated based on simulation calculation α mod q is easy, but knowing (H(C)) α It is difficult to deduce the value of α from the result of mod q and H(C) and q. If it is difficult to obtain the value of α, knowing (H(C)) α The result of mod q is also difficult to know the value of H(C). αThe expression of mod q omits mod q and is simply expressed as (H(C)) α .

[0054] Then, the database encrypted by the server using the secret value α is as follows:

[0055]

[0056]

[0057] Table 3. Query base encrypted by RSA private key on the server

[0058] The server can then send the query base to the client that needs to perform the search. Similarly, the server can send the query base directly to the client, such as directly to the client's device or to a proxy server of the client. Alternatively, the server can publish the query base on a Uniform Resource Locator (URL), and the client can then retrieve the query base from the URL.

[0059] Correspondingly, the client may receive the query base and save the received query base locally.

[0060] For example Figure 1 As shown, the interaction process between the client and the server may include the following steps:

[0061] S110: The client sends the sensitive field encrypted by itself to the server, and obtains the same sensitive field encrypted by the server through interaction with the server.

[0062] For example, the client's search condition is that the value of the Age field is 25. However, 25 is a sensitive field and is not intended to be known to the other party. To prevent the server from knowing that the client's search condition is the value of 25 in the Age field, the client can encrypt the value 25. For example, this can be done using RSA / ECC private key encryption, with the client using the same encryption algorithm as the server used to generate the query base.

[0063] Specifically, when using RSA private key encryption, the client generates a secret β and stores it properly. Then, the client can use its own private key β to encrypt 25. Specifically, it can encrypt 25 or the hash value of 25. Here, the hash encryption of 25 is used as an example to illustrate. The case of directly encrypting 25 is similar. The client and the server use the same hash algorithm. For example, the client uses the same large prime number q as the server as the modulus. The client can use β to perform RSA encryption on the hash value of 25 to obtain (H(25)) βThe sensitive field sent by the client to the server can be (H(25)) β , where (H(25)) β The ciphertext represents the value 25 of the sensitive field.

[0064] On the other hand, the client can also construct a search statement, encrypt the sensitive fields in the search statement to obtain the private fields, replace the sensitive fields with the private fields, and send the replaced private search statement to the server.

[0065] For example, the query statement constructed by the client is select Name where Age=25.

[0066] To protect privacy, the server is not allowed to obtain the query condition Age=25. For example, the privacy of 25 is protected. The result is as follows:

[0067] select Name where Age=?

[0068] Where? represents the search statement after replacement.

[0069] Specifically, the client can encrypt 25 with the RSA private key. For example, the client can use the same hash function as the server to perform hash calculation on 25, and then use β to perform RSA encryption on the hash value of 25 to obtain (H(25)) β The query statement sent by the client to the server is as follows:

[0070] select Name where Age=(H(25)) β

[0071] As mentioned above, (H(25)) β The ciphertext is the content represented by “?” in the search statement above. The server cannot know β and 25 after obtaining it.

[0072] When using ECC private key encryption, the client uses the same elliptic curve as the server, that is, it has the same elliptic curve parameters and generators. The client generates the secret β and stores it properly. Then, the client can use its own private key β to encrypt 25. Specifically, it can encrypt the hash value of 25, and the client and the server use the same hash algorithm. For example, the client can use β to perform ECC encryption on the hash value of 25 to obtain β·H(25). Then the sensitive field sent by the client to the server can be β·H(25), where β·H(25) represents the ciphertext of the value 25 of the sensitive field.

[0073] On the other hand, the client can also construct a search statement, encrypt the sensitive fields in the search statement to obtain the private fields, replace the sensitive fields with the private fields, and send the replaced private search statement to the server.

[0074] For example, the query statement constructed by the client is select Name where Age=25.

[0075] To protect privacy, the server is not allowed to obtain the query condition Age=25. For example, the privacy of 25 is protected. The result is as follows:

[0076] select Name where Age=?

[0077] Where? represents the search statement after replacement.

[0078] Specifically, the client can encrypt 25 with an ECC private key. For example, the client uses the same elliptic curve as the server, that is, it has the same elliptic curve parameters and generators. The client can replace the sensitive fields in the search statement with its own ECC private key and send the replaced private search statement to the server. For example, the client generates the secret β itself and saves it properly. In addition, the client can use the same hash function as the server to perform hash calculation on 25, and then use β to perform ECC encryption on the hash value of 25 to obtain β·H(25). Then the query statement sent by the client to the server is, for example, as follows:

[0079] select Name where Age=β·H(25)

[0080] As mentioned above, β·H(25) is the ciphertext, which is the content represented by “?” in the above search statement. The server cannot know β and 25 after obtaining it.

[0081] The client obtains the same sensitive field encrypted by the server through interaction with the server, which may include the server using its own key to encrypt the sensitive field encrypted by the client again and then sending it to the client, and the client using its own key to decrypt the sensitive field encrypted twice to obtain the sensitive field encrypted by the server. The core of this content is the need to find an encryption algorithm that can exchange the order of decryption for two consecutive encryption operations (two parties encrypt in succession). According to the cryptographic properties of ECC, the two parties agree to use the same elliptic curve, that is, have the same elliptic curve parameters and generators, each holding private keys α and β, and the encryption operation is a scalar multiplication operation using α (or β). Regardless of whether it is encrypted with α first and then β or encrypted with β first and then α, it can be decrypted in the same or different order, that is, the encryption results can be decrypted in different orders. Similarly, based on the cryptographic properties of RSA encryption, both parties agree to use the same large prime number q and primitive root g, each holding private keys α and β. The encryption operation is to exponentiate α (or β) and modulo q. Whether encrypting with α first and then β, or encrypting with β first and then α, the decryption can be done in the same or different order. In other words, the encryption results can be decrypted in different orders. Overall, the encryption and decryption performed by the client and server on the same target use a commutative encryption and decryption algorithm.

[0082] Specifically, after receiving the private search statement, the server can re-encrypt the private fields and return them to the client. Alternatively, after receiving the sensitive fields from the client that were encrypted by the client itself, the server can re-encrypt the encrypted sensitive fields using its own key and return them to the client. The client then uses its own key to decrypt the twice-encrypted sensitive fields, obtaining the sensitive fields encrypted by the server.

[0083] For example, case 1: the server can receive (H(25)) sent by the client. β .

[0084] The server can re-encrypt the encrypted sensitive field (ie, the private field) and return the re-encrypted sensitive field to the client. Specifically, the server can re-encrypt the private field (H(25)) β Use its own RSA private key α to encrypt again and get ((H(25)) β ) α .

[0085] For example, case 1': the server can receive the privacy search statement select Name where Age = (H(25)) sent by the client. β In this way, the server can obtain the privacy field (H(25)) in the privacy search statement. β .

[0086] The server can re-encrypt the private field and return the re-encrypted private field to the client.

[0087] Specifically, the server can set the privacy field (H(25)) β Use its own RSA private key α to encrypt again and get ((H(25)) β ) α The specific process is similar to the above and will not be repeated here.

[0088] For example, case 2: the server can receive β·H(25) sent by the client.

[0089] The server can re-encrypt the private field and return the re-encrypted private field to the client. Specifically, the server can re-encrypt the private field β·H(25) using its own ECC private key α to obtain α·β·H(25).

[0090] For example, in case 2', the server can receive the privacy search statement "select Name where Age = β·H(25)" sent by the client. In this way, the server can obtain the privacy field β·H(25) in the privacy search statement.

[0091] The server can re-encrypt the private field and return the re-encrypted private field to the client.

[0092] Specifically, the server can re-encrypt the privacy field β·H(25) using its own ECC private key α to obtain α·β·H(25). The specific process is similar to the above and will not be repeated here.

[0093] After the server uses its own key to re-encrypt the sensitive field (i.e., the privacy field) encrypted by the client and sends it to the client, the client can use its own key to decrypt the twice-encrypted privacy field to obtain the sensitive field encrypted by the server.

[0094] For example, corresponding to the above cases 1 and 1', the client receives ((H(25)) sent by the server β ) α , where the power operation has the following properties: ((H(25)) β ) α =(H(25)) βα =(H(25)) αβ =((H(25)) α ) β . Then, the client can use the inverse element of its own private key β Decrypt the twice-encrypted sensitive fields as follows: In this way, the client obtains the same sensitive field encrypted by the server, namely (H(25)) α .

[0095] Corresponding to cases 2 and 2' above, the client receives α·β·H(25) sent by the server, where the scalar multiplication operation has the following property: α·β·H(25)=β·α·H(25). Furthermore, the client can use the inverse element β of its own private key β -1 Decrypt the twice-encrypted sensitive fields as follows: -1 ·α·β·H(25)=β -1 ·β·α·H(25)=α·H(25). In this way, the client also obtains the same sensitive field encrypted by the server, namely α·H(25).

[0096] It should be noted that in RSA, according to Euler's theorem, pk·sk = 1 mod (p-1)·(q-1), where p and q are two large prime numbers, so pk and sk are inverses of each other. Similarly, in ECC, pk = sk*G, where G is a generator on the ECC curve, so pk and sk are also inverses of each other.

[0097] S120: The client searches the query base according to the sensitive field encrypted by the server, obtains the identifier of the matching record, and returns the identifier to the server.

[0098] After S110 is executed, the client can obtain the same sensitive field encrypted by the server.

[0099] The client can query the query base based on the sensitive field encrypted by the server. For example, after decryption, the client obtains the private field α·H(25) or (H(25)) encrypted by the server. α , the client can then query the query base based on this private field, for example, in Table 2 or Table 3, and retrieve the record with ID = d_1 containing the private field in its Age field. ID is the identifier of this record. In this way, the client can query the query base based on this private field encrypted by the server, and upon finding a matching record, locate and obtain the matching record's identifier.

[0100] The client returns the identifier of the matching record to the server, which may include two situations.

[0101] One scenario is that in S110, the client constructs a search statement, encrypts the sensitive fields in the search statement to obtain private fields, replaces the sensitive fields with the private fields, and sends the replaced private search statement to the server. In this case, the client can directly return the identifier of the matching record to the server.

[0102] Another case is that in S110, the client sends the value of the sensitive field encrypted by itself to the server. In this case, in S120, the client can construct a search statement, for example, the search statement is:

[0103] select Name where ID=id_1

[0104] Thus, in S120 , the client may send the constructed search formula to the server. The search formula includes the identifier of the matching record and indicates that the field of interest is Name, that is, the field name immediately following select.

[0105] In other words, in steps S110 and S120 , you can choose to send a search formula in one of the steps, and the search formula includes the field of interest.

[0106] S130: The server returns the value of the field of interest in the record corresponding to the identifier in the database to the client.

[0107] Continuing with the above example, after receiving the identifier sent by the client, the server can search the database for the record corresponding to the identifier, retrieve the corresponding value from the found record according to the field of interest in S110 or S120, and return the retrieved value of the field of interest to the client. For example, returning Name = B in the record corresponding to id_1 means returning B to the client.

[0108] In the above-described embodiment, by preconfiguring the query base on the client, the client can locate the identifier of the field to be queried in the query base through interaction with the server and the query base, without exposing the database plaintext. The client can then initiate a query to the server based on the identifier to obtain the value of the field of interest in the record corresponding to the identifier. Compared to traditional multi-replica PIR, this eliminates the need for the assumption that multiple replica databases cannot collude, resulting in improved practicality. Compared to traditional single-replica PIR, which can only perform bit-based searches, this embodiment does not require the specific location (bit position) of the keyword to be searched in the database, can perform string queries, and supports Structured Query Language (SQL). In this embodiment, the database remains on the server, while the query base, encrypted from the database, is configured on the client. This allows the client to locate data based on the query base to obtain the record identifier during searches. Furthermore, the encrypted nature of the query base prevents the client from accessing the database contents, ensuring the server's privacy protection of the database. In general, the form of the database and query base in this embodiment can be called "asymmetric dual copies" when a database is configured on a server and a query base is configured on a client, and can be called "asymmetric multiple copies" when query bases are configured on multiple clients.

[0109] In the above embodiment, through the SQL query statement, the client can initiate a query on the field of interest, such as the Name field to be queried in the above select Name... This exposes the client's fields of interest to a certain extent. In another way, you can query the records that meet the conditions, that is, the entire row of data that meets the conditions. This can protect the privacy of the client, but the server needs to return the entire record, which exposes the entire row of data on the server to a certain extent. For example, in S110 / S120, through the search statement such as "select*where Age=?" or "select*where ID=id_1". In this way, the result returned by the server can be the record of id_1, for example as follows:

[0110] id_1B 25shanghai

[0111] In addition, to ensure the security of the transmission process, the server can encrypt the record corresponding to the identifier in the database / the value of the field of interest in the corresponding record and return it to the client. For example, the server can use a symmetric key negotiated with the client to encrypt the record corresponding to the identifier in the database / the value of the field of interest in the corresponding record and return it to the client, or use the public key of the client's asymmetric key to encrypt the record corresponding to the identifier in the database / the value of the field of interest in the corresponding record and return it to the client, so that the client can decrypt it with its own private key, or use a digital envelope method, etc.

[0112] In the above S120, the client directly returns the matched ID to the server. Although the record corresponding to the ID or the field of interest in the record can be obtained from the server, as in S130, this will expose the client's privacy to a certain extent, that is, the server will know that the identifier the client wants to query is id_1. In order to protect the client's privacy, this can be achieved through the following embodiments:

[0113] S210: The client sends the sensitive field encrypted by itself to the server, and obtains the same sensitive field encrypted by the server through interaction with the server.

[0114] For example, the client's search condition is that the value of the Age field is 25. However, 25 is a sensitive field and is not intended to be known to the other party. To prevent the server from knowing that the client's search condition is the value of 25 in the Age field, the client can encrypt the value 25. For example, this can be done using RSA / ECC private key encryption, with the client using the same encryption algorithm as the server used to generate the query base.

[0115] Specifically, when using RSA private key encryption, the client generates a secret β and stores it properly. Then, the client can use its own private key β to encrypt 25. Specifically, it can encrypt 25 or the hash value of 25. Here, the hash encryption of 25 is used as an example to illustrate. The case of directly encrypting 25 is similar. The client and the server use the same hash algorithm. For example, the client uses the same large prime number q as the server as the modulus. The client can use β to perform RSA encryption on the hash value of 25 to obtain (H(25)) β The sensitive field sent by the client to the server can be (H(25)) β , where (H(25)) β The ciphertext represents the value 25 of the sensitive field.

[0116] When using ECC private key encryption, the client uses the same elliptic curve as the server, that is, it has the same elliptic curve parameters and generators. The client generates the secret β and stores it properly. Then, the client can use its own private key β to encrypt 25. Specifically, it can encrypt the hash value of 25, and the client and the server use the same hash algorithm. For example, the client can use β to perform ECC encryption on the hash value of 25 to obtain β·H(25). Then the sensitive field sent by the client to the server can be β·H(25), where β·H(25) represents the ciphertext of the value 25 of the sensitive field.

[0117] The client obtains the same sensitive field encrypted by the server through interaction with the server, which may include the server using its own key to encrypt the sensitive field encrypted by the client again and then sending it to the client, and the client using its own key to decrypt the sensitive field encrypted twice to obtain the sensitive field encrypted by the server. The core of this content is the need to find an encryption algorithm that can exchange the order of decryption for two consecutive encryption operations (two parties encrypt in succession). According to the cryptographic properties of ECC, the two parties agree to use the same elliptic curve, that is, have the same elliptic curve parameters and generators, each holding private keys α and β, and the encryption operation is a scalar multiplication operation using α (or β). Regardless of whether it is encrypted with α first and then β or encrypted with β first and then α, it can be decrypted in the same or different order, that is, the encryption results can be decrypted in different orders. Similarly, based on the cryptographic properties of RSA encryption, both parties agree to use the same large prime number q and primitive root g, each holding private keys α and β. The encryption operation is to exponentiate α (or β) and modulo q. Whether encrypting with α first and then β, or encrypting with β first and then α, the decryption can be done in the same or different order. In other words, the encryption results can be decrypted in different orders. Overall, the encryption and decryption performed by the client and server on the same target use a commutative encryption and decryption algorithm.

[0118] Specifically, after receiving the sensitive fields encrypted by the client, the server re-encrypts the encrypted sensitive fields with its own key and returns them to the client. The client then decrypts the twice-encrypted sensitive fields using its own key to obtain the sensitive fields encrypted by the server.

[0119] For example, case 1: the server can receive (H(25)) sent by the client. β .

[0120] The server can re-encrypt the encrypted sensitive field (ie, the private field) and return the re-encrypted sensitive field to the client. Specifically, the server can re-encrypt the private field (H(25)) βUse its own RSA private key α to encrypt again and get ((H(25)) β ) α .

[0121] For example, case 2: the server can receive β·H(25) sent by the client.

[0122] The server can re-encrypt the private field and return the re-encrypted private field to the client. Specifically, the server can re-encrypt the private field β·H(25) using its own ECC private key α to obtain α·β·H(25).

[0123] After the server uses its own key to re-encrypt the sensitive field (i.e., the privacy field) encrypted by the client and sends it to the client, the client can use its own key to decrypt the twice-encrypted privacy field to obtain the sensitive field encrypted by the server.

[0124] For example, corresponding to the above case 1, the client receives ((H(25)) sent by the server β ) α , where the power operation has the following properties: ((H(25)) β ) α =(H(25)) βα =(H(25)) αβ =((H(25)) α ) β . Then, the client can use the inverse element of its own private key β Decrypt the twice-encrypted sensitive fields as follows: In this way, the client obtains the same sensitive field encrypted by the server, namely (H(25)) α .

[0125] Corresponding to the above case 2, the client receives α·β·H(25) sent by the server, where the scalar multiplication operation has the following property: α·β·H(25)=β·α·H(25). Furthermore, the client can use the inverse element β of its own private key β -1 Decrypt the twice-encrypted sensitive fields as follows: -1 ·α·β·H(25)=β -1 ·β·α·H(25)=α·H(25). In this way, the client also obtains the same sensitive field encrypted by the server, namely α·H(25).

[0126] It should be noted that in RSA, according to Euler's theorem, pk·sk = 1 mod (p-1)·(q-1), where p and q are two large prime numbers, so pk and sk are inverses of each other. Similarly, in ECC, pk = sk·G, where G is a generator on the ECC selected curve, so pk and sk are also inverses of each other.

[0127] S220: The client searches the query base according to the sensitive field encrypted by the server to obtain an identifier of a matching record.

[0128] After S210 is executed, the client can obtain the same sensitive field encrypted by the server.

[0129] The client can query the query base based on the sensitive field encrypted by the server. For example, after decryption, the client obtains the private field α·H(25) or (H(25)) encrypted by the server. α , the client can then query the query base based on this private field, for example, in Table 2 or Table 3, and retrieve the record with ID = d_1 containing the private field in its Age field. ID is the identifier of this record. In this way, the client can query the query base based on this private field encrypted by the server, and upon finding a matching record, locate and obtain the matching record's identifier.

[0130] S230: The server returns the value of the field of interest in the record corresponding to the predetermined size identifier set including the matching identifier in the database to the client in an oblivious transmission manner.

[0131] In S220, the client does not return the matching IDs to the server, so the server cannot know which record or records the client is looking for. The client-encrypted sensitive fields sent by the client in S210 also prevent the server from knowing which record or records the client's search for sensitive fields will match; only the client knows. This protects the client's privacy. However, the search still needs to be completed, which requires the server to return the records the client is looking for.

[0132] Here, the server can use oblivious transmission.

[0133] Oblivious Transfer (OT) can be implemented based on RSA, ECC, etc., and can achieve various OTs such as 2-out-of-1, n-out-of-1, m-out-of-1, m-out-of-k (k < m < n), etc. Taking 2-out-of-1 OT as an example to illustrate its principle, the sender has two secrets, namely m1 and m2, and needs to send 2 secrets to the receiver. The receiver can only choose to decrypt one of them and cannot know the other, and at the same time, the sender also cannot know which one the receiver has chosen. Taking RSA as an example, a simple implementation process of 2-out-of-1 is as follows:

[0134] First, the sender generates two different pairs of public and private keys and publishes the two public keys, denoted as public key 1 and public key 2 respectively. Suppose the receiver hopes to know m1 but does not want the sender to know that he wants m1. The receiver generates a random number r, encrypts r with public key 1, and sends it to the sender. The sender decrypts the encrypted r with its own two private keys, decrypts it with private key 1 to get r1, and decrypts it with private key 2 to get r2. Obviously, only r1 is equal to r, and r2 is a string of meaningless numbers (also the decryption result). However, the sender does not know which public key the receiver used for encryption, so the sender also does not know which of the calculated r1 and r2 is the real r. After the sender receives m1 and m2, it symmetrically encrypts m1 with r1 and symmetrically encrypts m2 with r2, and sends the two symmetric encryption results to the receiver. The receiver has r = r1 locally, so the receiver can symmetrically decrypt the two received results with r to get m1, but cannot decrypt to get m2 because the receiver's r ≠ r2, and the receiver cannot decrypt with the correct symmetric key to get the value of m2. In this process, the sender also does not know which of m1 and m2 the receiver has calculated.

[0135] With 2-out-of-1 as the basis, the 2 pairs of public and private keys can be extended to n pairs of public and private keys, which becomes n-out-of-1 OT. The core of n-out-of-1 is that the server encrypts n records in the data table / the values of the interested fields in the corresponding records with n different keys to obtain n encrypted results, and sends the n encrypted results to the client; the client decrypts 1 encrypted result corresponding to the matching identifier among the n encrypted results sent by the server using the key corresponding to the matching identifier.

[0136] Combined with the above embodiments of this specification, assume that there are a total of n records in the data table of the server. In this way, there are also n encrypted records corresponding in the query base of the client. For convenience, the IDs of the data records are sequentially identified as id_0, id_1, id_2,... id_n-1. A simple implementation process is as follows:

[0137] S231: The server pregenerates n different public and private key pairs and publishes the public key.

[0138] Here n is equal to the number of records in the database.

[0139] The server generates n pairs of public and private key pairs (pk-sk; pk is publick key, indicating public key; sk is secretkey, indicating private key; public key can be made public, private key needs to be kept secret), for example, pk0-sk0, pk1-sk1, pk2-sk2, ..., pk n-1 -sk n-1 , and make these n public keys public, that is, public pk0, pk1, pk2, ..., pk n-1 After the server publishes these n ordered public keys, the client can obtain these n public keys.

[0140] S232: The client generates a random number r, encrypts r with the public key corresponding to the desired ID, and sends the encrypted number to the server.

[0141] Here we assume that the client wants to obtain the record with id_1, but does not want the server to know that the record the client wants to obtain is the one with id_1. In this way, the client can use pk1 to encrypt r and send it to the server. The above-mentioned order mainly refers to the correspondence between ID and public key, and such a correspondence can be known by the client. For example, in the above example, the client wants to obtain the record with id_1 but does not want the server to know that the record the client wants to obtain is the one with id_1. The client can use the pk1 corresponding to id_1 to encrypt r and send it to the server; similarly, the client wants to obtain the record with id_t but does not want the server to know that the record the client wants to obtain is the one with id_t. The client can use the pk1 corresponding to id_t to encrypt r. t Encrypt r and send it to the server.

[0142] S233: After receiving the encrypted r, the server uses n private keys to decrypt it respectively.

[0143] The server uses sk0, sk1, sk2, ..., sk n-1 Decrypt the random number r encrypted by pk1 respectively. For example, the server uses sk0 to decrypt to get r0, sk1 to decrypt to get r1, ..., sk n-1 Decryption yields r(n-1).

[0144] Obviously, only r1 is equal to r, because only the decryption with sk1 is encrypted with the corresponding pk1; and the decryption with sk0, sk2, ..., sk n-1The decrypted results r0, r2, ..., r(n-1) will never be the same as r. The server only obtains the same decrypted results; it does not know the true r or the public key used by the client to encrypt it. In other words, the server does not know the public key used by the client to encrypt r, and therefore does not know which of the n decrypted results r0, r1, r2, ..., r(n-1) is the true r.

[0145] S234: The server symmetrically encrypts each record in the database according to the serial number using the decryption result of the corresponding serial number, and sends the symmetrically encrypted result to the client.

[0146] For example, the server symmetric encrypts the record id_0 using r0, symmetric encrypts the record id_1 using r1, ..., symmetric encrypts the record id_n-1 using r(n-1), and sends these n symmetric encryption results to the client.

[0147] S235: The client uses the random number r to symmetrically decrypt the encryption result corresponding to the desired ID in the symmetrical encryption result to obtain a retrieval result.

[0148] The client uses the random number r to symmetrically decrypt the encryption result corresponding to the ID expected in the symmetric encryption result. Specifically, for example, in the above S232, the client expects to obtain the value of the record / interested field in the record corresponding to id_1, then the client uses the corresponding public key pk1 to encrypt the random number r; in S233, the server uses the corresponding private key sk1 to decrypt the decrypted result, and the obtained r1=r, and uses sk0, sk2, ..., sk that do not correspond to pk1. n-1 The decrypted results r0, r2, ..., and r(n-1) will not be the same as r. In S234, the server uses r0, r1, r2, ..., and r(n-1) to symmetrically encrypt the corresponding records (id_0, id_1, ..., and id_n-1) or the values ​​of the fields of interest within those records, respectively, and sends these n symmetrically encrypted results to the client. In S235, the client uses the random number r to symmetrically decrypt the n symmetrically encrypted results. Of the n symmetrically decrypted results, only the encryption result for id_1 was symmetrically encrypted using r. Therefore, only decrypting the encryption result for id_1 using r can yield the correct value. Thus, the client obtains the correct search result, namely, the value of the corresponding record or the field of interest within the corresponding record.

[0149] Of course, in order to reduce the amount of calculation, the client can only use the random number r to symmetric decrypt the encryption result corresponding to the desired ID among the n symmetric encryption results, that is, the client only uses r to decrypt the encryption result of id_1, thereby obtaining the corresponding record of id_1 / the value of the field of interest in the corresponding record, without using r to symmetric decrypt the symmetric encryption results of id_0, id_2, ..., id_n-1, because the client can know that these encryption results are not symmetric encrypted using r, and even if r is used for symmetric decryption, the correct result cannot be obtained.

[0150] For a clearer presentation, the client uses the random number r to symmetrically decrypt the n symmetrical encryption results, as explained below:

[0151] In S234, the server uses r0, r1, r2, ..., r(n-1) to symmetrically encrypt the values ​​of the fields of interest in the corresponding records id_0, id_1, ..., id_n-1 respectively:

[0152] Enc(id_0, r0), where r0≠r;

[0153] Enc(id_1, r1), where r1=r;

[0154] Enc(id_2, r2), where r2≠r; ...

[0155] Enc(id_n-1, r(n-1)), where r(n-1)≠r;

[0156] The Enc in the above code represents encryption. The first part of the Enc() bracket (id_0, id_1, id_2, ..., id_n-1) represents the value of the field of interest in n records / n records, and the second part (r0, r1, r2, ..., r(n-1)) represents the encryption key.

[0157] In S235, the client uses the random number r to symmetrically decrypt the encryption result in S234. Specifically, the client uses the random number r to symmetrically decrypt the following contents:

[0158] Dec(Enc(id_0, r0), r), where r0≠r;

[0159] Dec(Enc(id_1, r1), r), where r1=r;

[0160] Dec(Enc(id_2, r2), r), where r2≠r; ...

[0161] Dec(Enc(id_n-1, r(n-1), r), where r(n-1)≠r;

[0162] The above Dec represents decryption (Decrypt), the first part of Dec() represents the decryption object, which is the encryption result above, and the second part of Dec() represents the key used for decryption.

[0163] As you can see, the client can only decrypt the record with id_1 and cannot infer other records. This is because the server only uses the random number r for symmetric encryption of the record with id_1, and does not use the random number r for symmetric encryption of other IDs. The client cannot obtain r0, r2, ..., r(n-1) other than r1=r.

[0164] It should be noted that S231 may be after S230 or before S230, which is not limited here.

[0165] When the number of matching records retrieved by the client in the query base is 1, all records corresponding to the identifiers in the database including the matching identifier / values ​​of the fields of interest in the corresponding records are transmitted to the client through n-choose-1 oblivious transmission.

[0166] In the above embodiment, the server does not know which ID or IDs the client is querying. Instead, it encrypts all records in the database and returns them to the client, protecting the client's privacy. However, in S233, the server uses n private keys to decrypt the received encrypted r. This requires a large amount of asymmetric decryption calculations, consuming significant CPU and memory resources. Furthermore, transmitting the n symmetric encrypted results in S234 also consumes a significant amount of bandwidth. Especially when n is large, the server's computational workload and bandwidth usage are high.

[0167] In addition, if the possible matching result is greater than 1, for example, k (k>1), it can be achieved through n-choose-k oblivious transfer. Regarding n-choose-k, one implementation is to group each k of the n records into a set, and each set corresponds to a public-private key pair, so that there will be a total of (C represents the combination formula, the number of combinations consisting of any k items in n). Next, use The oblivious transmission method of option 1 transmits all records corresponding to the identifier including the matching identifier in the database / the values ​​of the fields of interest in the corresponding records to the client. The implementation process of the 1-choose-1 oblivious transmission method is similar to the implementation process of the n-choose-1 oblivious transmission method. Different keys are used to encrypt each k records in the data table / the value of the field of interest in the corresponding record to obtain The encrypted result is sent The client uses the key corresponding to the matching identifier to decrypt the encrypted result sent by the server. The specific implementation is similar to the above-mentioned process of S231-S235, which will not be repeated here.

[0168] It should be noted that r can also be the public key in an asymmetric key. In this way, after receiving the result encrypted by r, the client can decrypt it using its own private key. Specifically, in S234 and S235, the server asymmetrically encrypts the record id_0 using r0, the record id_1 using r1, ..., and the record id_n-1 using r(n-1). These n encrypted results are then sent to the client. After receiving these encrypted results, the client can asymmetrically decrypt the encrypted result corresponding to the desired ID using its own private key. The following procedures are similar and will not be repeated here.

[0169] Based on this, this specification provides the following implementation method that adds the construction of a confusion set:

[0170] S310: The client sends the sensitive field encrypted by itself to the server, and obtains the same sensitive field encrypted by the server through interaction with the server.

[0171] For example, the client's search condition is that the value of the Age field is 25. However, 25 is a sensitive field and is not intended to be known to the other party. To prevent the server from knowing that the client's search condition is the value of 25 in the Age field, the client can encrypt the value 25. For example, this can be done using RSA / ECC private key encryption, with the client using the same encryption algorithm as the server used to generate the query base.

[0172] Specifically, when using RSA private key encryption, the client generates a secret β and stores it properly. Then, the client can use its own private key β to encrypt 25. Specifically, it can encrypt 25 or the hash value of 25. Here, the hash encryption of 25 is used as an example to illustrate. The case of directly encrypting 25 is similar. The client and the server use the same hash algorithm. For example, the client uses the same large prime number q as the server as the modulus. The client can use β to perform RSA encryption on the hash value of 25 to obtain (H(25)) β The sensitive field sent by the client to the server can be (H(25)) β , where (H(25))β The ciphertext represents the value 25 of the sensitive field.

[0173] When using ECC private key encryption, the client uses the same elliptic curve as the server, that is, it has the same elliptic curve parameters and generators. The client generates the secret β and stores it properly. Then, the client can use its own private key β to encrypt 25. Specifically, it can encrypt the hash value of 25, and the client and the server use the same hash algorithm. For example, the client can use β to perform ECC encryption on the hash value of 25 to obtain β·H(25). Then the sensitive field sent by the client to the server can be β·H(25), where β·H(25) represents the ciphertext of the value 25 of the sensitive field.

[0174] The client obtains the same sensitive field encrypted by the server through interaction with the server, which may include the server using its own key to encrypt the sensitive field encrypted by the client again and then sending it to the client, and the client using its own key to decrypt the sensitive field encrypted twice to obtain the sensitive field encrypted by the server. The core of this content is the need to find an encryption algorithm that can exchange the order of decryption for two consecutive encryption operations (two parties encrypt in succession). According to the cryptographic properties of ECC, the two parties agree to use the same elliptic curve, that is, have the same elliptic curve parameters and generators, each holding private keys α and β, and the encryption operation is a scalar multiplication operation using α (or β). Regardless of whether it is encrypted with α first and then β or encrypted with β first and then α, it can be decrypted in the same or different order, that is, the encryption results can be decrypted in different orders. Similarly, based on the cryptographic properties of RSA encryption, both parties agree to use the same large prime number q and primitive root g, each holding private keys α and β. The encryption operation is to exponentiate α (or β) and modulo q. Whether encrypting with α first and then β, or encrypting with β first and then α, the decryption can be done in the same or different order. In other words, the encryption results can be decrypted in different orders. Overall, the encryption and decryption performed by the client and server on the same target use a commutative encryption and decryption algorithm.

[0175] Specifically, after receiving the sensitive fields encrypted by the client, the server re-encrypts the encrypted sensitive fields with its own key and returns them to the client. The client then decrypts the twice-encrypted sensitive fields using its own key to obtain the sensitive fields encrypted by the server.

[0176] For example, case 1: the server can receive (H(25)) sent by the client. β .

[0177] The server can re-encrypt the encrypted sensitive field (ie, the private field) and return the re-encrypted sensitive field to the client. Specifically, the server can re-encrypt the private field (H(25)) β Use its own RSA private key α to encrypt again and get ((H(25)) β ) α .

[0178] For example, case 2: the server can receive β·H(25) sent by the client.

[0179] The server can re-encrypt the private field and return the re-encrypted private field to the client. Specifically, the server can re-encrypt the private field β·H(25) using its own ECC private key α to obtain α·β·H(25).

[0180] After the server uses its own key to re-encrypt the sensitive field (i.e., the privacy field) encrypted by the client and sends it to the client, the client can use its own key to decrypt the twice-encrypted privacy field to obtain the sensitive field encrypted by the server.

[0181] For example, corresponding to the above case 1, the client receives ((H(25)) sent by the server β ) α , where the power operation has the following properties: ((H(25)) β ) α =(H(25)) βα =(H(25)) αβ =((H(25)) α ) β . Then, the client can use the inverse element of its own private key β Decrypt the twice-encrypted sensitive fields as follows: In this way, the client obtains the same sensitive field encrypted by the server, namely (H(25)) α .

[0182] Corresponding to the above case 2, the client receives α·β·H(25) sent by the server, where the scalar multiplication operation has the following property: α·β·H(25)=β·α·H(25). Furthermore, the client can use the inverse element β of its own private key β -1 Decrypt the twice-encrypted sensitive fields as follows: -1 ·α·β·H(25)=β -1 ·β·α·H(25)=α·H(25). In this way, the client also obtains the same sensitive field encrypted by the server, namely α·H(25).

[0183] It should be noted that in RSA, according to Euler's theorem, pk·sk = 1 mod (p - 1)·(q - 1), where p and q are two large prime numbers, so pk and sk are inverses of each other. Similarly, in ECC, pk = sk·G, where G is a generator on the curve selected by ECC, so pk and sk are also inverses of each other.

[0184] S320: The client retrieves in the query base according to the sensitive field encrypted by the server, and obtains the identifier of the matching record.

[0185] After S310 is executed, the client can obtain the same sensitive field encrypted by the server.

[0186] The client can query in the query base based on the sensitive field encrypted by the server. For example, after the client decrypts, it obtains the privacy field encrypted by the server, α·H(25) or (H(25)) α , so that the client queries in the query base based on this privacy field. For example, by querying in Table 2 or Table 3 respectively, it can obtain the record with ID = d_1 in Age that contains this privacy field, and ID is the identifier of this record. In this way, the client queries in the query base based on the privacy field encrypted by the server, and after matching the record, it can locate and obtain the identifier of the matching record.

[0187] S330: The server uses the oblivious transfer method to return the values of the interesting fields in the records corresponding to the set of identifiers of a predetermined size including the matching identifier in the database to the client.

[0188] The sensitive field encrypted by the client sent in S310 makes it impossible for the server to know which record the sensitive field searched by the client will hit, and only the client itself knows. In this way, the privacy of the client is protected. However, the retrieval still needs to be completed ultimately, which requires the server to return the record that the client wants to query to the client.

[0189] The above Figure 2 The corresponding embodiment gives the implementation method of 1-out-of-n oblivious transfer. Here, 1-out-of-m oblivious transfer can be used, where m < n. In S320, the client can also not return the obtained matching ID to the server alone, but mix and combine the obtained matching ID with some other forged IDs to construct a confusion set, and send the confusion set to the server. In this way, the server cannot accurately know which record in the confusion set the client wants to search for, and it is necessary to ensure that the client can only obtain the record it wants to search for, and cannot obtain other records. In S320, the confusion set sent by the client can be sent together with the retrieval statement, such as select Name where ID = confusion set. Or, the confusion set can also be sent in S332 below, and this is not limited here.

[0190] In conjunction with the above embodiments of this specification, assume that the server's data table contains a total of n records. Thus, the client's query base also contains n encrypted records. For convenience, the data record IDs are sequentially identified as id_0, id_1, id_2, ..., id_n-1. A simple implementation process is as follows:

[0191] S331: The server pregenerates n different public and private key pairs and publishes the public key.

[0192] The server generates n different public and private key pairs (pk-sk; pk is publick key, indicating public key; sk is secretkey, indicating private key), for example, pk0-sk0, pk1-sk1, pk2-sk2, ..., pk n-1 -sk n-1 , and make these n public keys public, that is, public pk0, pk1, pk2, ..., pk n-1 After the server publishes these n ordered public keys, the client can obtain these n public keys.

[0193] S332: The client generates a confusion set of size m including the desired ID, generates a random number r, encrypts r with the public key corresponding to the desired ID, and sends it to the server together with the confusion set.

[0194] Here, it is assumed that the client wants to obtain the record with id_1, but does not want the server to know that the record the client wants to obtain is the one with id_1. Therefore, a confusion set of size m is generated. When m=4, the confusion set is, for example: {id_1, id_2, id_3, id_4}.

[0195] For example, the four IDs and public key pairs have the following corresponding relationships:

[0196] pk1,id_1

[0197] pk2,id_2

[0198] pk3,id_3

[0199] pk4,id_4

[0200] The client can use pk1 to encrypt r and send it to the server together with the confusion set. For example:

[0201] In addition, the client can send the confusion set together with the search statement to the server. In this way, the client can use pk1 to encrypt r and send it together with the search statement containing the confusion set, for example:

[0202] select Name where ID={id_1,id_2,id_3,id_4}|Enc(r,pk1)

[0203] Among them, “|” is used to separate the preceding search statement and the following encrypted random number, the same below.

[0204] S333: After receiving the obfuscation set and the encrypted r, the server uses the corresponding m private keys to decrypt the encrypted r respectively.

[0205] The server uses sk1, sk2, sk3, and sk4 to decrypt the random number r encrypted by pk1. For example, the server decrypts r1 with sk1, r2 with sk2, r3 with sk3, and r4 with sk4.

[0206] Clearly, only r1 is equal to r, because only the decrypted value from sk1 is encrypted with the corresponding pk1. Decrypted values ​​from sk2, sk3, and sk4, which do not correspond to pk1, r2, r3, and r4 will all be identical to r. The server only obtains the same decrypted value; it has no idea of ​​the true r or the public key used by the client to encrypt it. In other words, the server does not know the public key used by the client to encrypt r, and therefore does not know which of the four decrypted values, r1, r2, r3, and r4, is the true r.

[0207] In addition, after the server receives the obfuscation set {id_1, id_2, id_3, id_4}, it can know from the obfuscation set that the data the client wants to obtain is one of the four IDs in the obfuscation set, but it is not sure which one it is, thereby protecting the client's privacy.

[0208] S334: The server symmetrically encrypts the record specified in the obfuscation set using the decryption result of the corresponding serial number, and sends the symmetrically encrypted result to the client.

[0209] For example, the server symmetric encrypts the record id_1 using r1, the record id_2 using r2, the record id_3 using r3, and the record id_4 using r4, and sends these four symmetric encryption results to the client.

[0210] S335: The client uses the random number r to symmetrically decrypt the encryption result corresponding to the desired ID in the symmetrical encryption result to obtain a retrieval result.

[0211] The client uses the random number r to symmetrically decrypt the encryption result corresponding to the desired ID in the symmetric encryption result. Specifically, for example, if the client in S332 desires to obtain the value of the field of interest in the record corresponding to id_1, the client encrypts the random number r using the corresponding public key pk1. In S333, the server decrypts the decrypted result using the corresponding private key sk1, resulting in r1 = r. Decryption results r2, r3, and r4 obtained using sk2, sk3, and sk4 that do not correspond to pk1 will not be the same as r. In S334, the server symmetrically encrypts the values ​​of the fields of interest in the records corresponding to id_1, id_2, id_3, and id_4 using r1, r2, r3, and r4, respectively, and sends the symmetric encryption results to the client. In S335, the client uses the random number r to symmetrically decrypt the four symmetric encryption results. Of the four symmetric decryption results, only the encryption result for id_1 is symmetrically encrypted using r. Therefore, only by decrypting the encryption result for id_1 using r can the correct value be obtained. Thus, the client can obtain the correct search result, namely the value of the corresponding record or the field of interest within the corresponding record.

[0212] Of course, in order to reduce the amount of calculation, the client can only use the random number r to symmetrically decrypt the encryption result corresponding to the ID expected to be obtained among the four symmetric encryption results, that is, the client only uses r to decrypt the encryption result of id_1, thereby obtaining the corresponding record of id_1 / the value of the field of interest in the corresponding record, and there is no need to use r to symmetrically decrypt the symmetric encryption results of id_2, id_3, and id_4, because the client can know that these encryption results are not symmetrically encrypted using r, and even if r is used for symmetric decryption, the correct result cannot be obtained.

[0213] The above S331 - S335 are just an exemplary implementation. In another implementation, the construction and transmission of the confusion set can be decoupled from the execution of the OT protocol, and the OT protocol is used to transmit the key. Specifically, on the one hand, the client can send a confusion set of size m to the server. Of course, the client knows which one in the confusion set of size m is the identifier of the result it really wants to obtain. On the other hand, the server can generate m symmetric keys. Through 1 - out - of - m OT, the client can obtain a specified symmetric key, that is, the client obtains the symmetric key corresponding to the identifier of the result it really wants to obtain. In this way, the server can encrypt the records corresponding to the m identifiers in the client's confusion set with the corresponding symmetric keys and send them to the client. Thus, the client decrypts the result it really wants to obtain with the correct symmetric key to get the result. Among them, the server can pre - generate m symmetric keys, so that the key preparation work can be completed in batches before the OT interaction, without occupying the time of the OT protocol execution.

[0214] When the number of matching records retrieved by the client in the query base is 1 as described above, through 1 - out - of - m oblivious transfer, the server transmits the records corresponding to the m identifiers specified in the confusion set / the values of the fields of interest in the corresponding records to the client. In addition, the number of possible matching results may be greater than 1, for example, it is k (k > 1), then it can be achieved through k - out - of - m oblivious transfer. The core of k - out - of - m oblivious transfer is that the client constructs a confusion set of size m and sends it to the server, 1 < m < n. One of the m confusion sets contains the identifiers of the k matching records. For example, m = 4, k = 2, and the matching identifiers are id_1 and id_3. Then the constructed confusion set is, for example, {{id_1, id_3}, {id_2 and id_4}, {id_3 and id_4}, {id_2}}. Obviously, the first one is the subset composed of the matching identifiers. Furthermore, the server generates m symmetric keys, and through the 1 - out - of - m OT protocol, the client obtains the specified symmetric key; the server encrypts the m subsets in the confusion set with m different symmetric keys to obtain m encrypted result subsets, and sends these m encrypted result subsets to the client; the client decrypts the subset composed of the k matching identifiers with the obtained symmetric key to obtain the correct decryption result. The specific implementation is similar to the process of decoupling the construction and transmission of the confusion set from the execution of the OT protocol and using the OT protocol to transmit the key as described above, and will not be elaborated here.

[0215] The examples of S310 - S330 above provide a scheme for constructing a confusion set to protect client privacy. In some cases, the confusion set may not be constructed reasonably, and it is still easy for the server to guess the ID that the client really wants to query.

[0216] For example, in the example in Table 1 above, it can be noted that the Age and Native_place in the two rows id_4 and id_7 are the same, that is, the two rows cannot be distinguished by Age and Native_place. In the above S332, the client constructs the confusion set {id_1, id_4}, and the constructed search statement is:

[0217] select Name where ID={id_1,id_4}

[0218] After receiving the search statement containing the confusion set, the server can know that the field of interest in the search statement is Name, not Age and Native_place. This means that the sensitive fields in S310 and S320 that the client previously identified should be Age and / or Native_place. However, the element id_4 in the confusion set cannot be distinguished from the row id_7 based on Age and / or Native_place. In other words, if the elements in the confusion set are located using the sensitive fields Age and / or Native_place, it is unreasonable to have only id_4 and no id_7. In this search statement, if there is id_4 in the confusion set, there should also be id_7. In this way, the server can infer that in the existing confusion set {id_1, id_4}, id_4 is a fake element, and id_1 is the actual ID to be queried, which to a certain extent causes the client's privacy to be leaked.

[0219] The above content is an example of client privacy leakage caused by unreasonable confusion set construction. In order to avoid unreasonable confusion set construction, this specification also provides an embodiment of constructing a confusion set.

[0220] The method of constructing the confusion set is as follows Figure 3 The following may be included:

[0221] S410: The client sends the sensitive field encrypted by itself to the server, and obtains the same sensitive field encrypted by the server through interaction with the server.

[0222] S420: The client searches the query base according to the sensitive fields encrypted by the server to obtain a first identification set of matching records.

[0223] S410 and S420 are similar to the aforementioned S310 and S320, and will not be repeated here. Assume that the sensitive field sent by the client to the server is (H(25)) β Or β·H(25), then the same sensitive field encrypted by the server received by the client through interaction with the server is ((H(25)) β ) αOr α·β·H(25). The client can use its own private key β to decrypt the twice encrypted sensitive fields, for example (H(25)) α Or α·H(25). In this way, the client queries the sensitive field (H(25)) encrypted by the server in the local query base. α Or α·H(25) search, the first identifier of the matching record is id_1. The following mainly uses Table 2 as an example to illustrate.

[0224] S430: The client selects at least one point in the query base except the ID field and the field of interest, uses the selected at least one point as a search condition, constructs a search statement based on the search condition and the original field of interest, and performs a search on the query base to obtain a second identifier set of matching records.

[0225] Assuming that the confusion set size is n, that is, the confusion set contains n elements, the embodiments of this specification can generate a confusion set consisting of n elements, for example, each element is an identifier, and the n identifiers include the identifier of the matching record, that is, the aforementioned first identifier set.

[0226] In S430, still using Table 2 or Table 3 as an example, the client selects at least one point in the query base other than the ID field and the field of interest, for example, one point, specifically, the point α·H(34). A point in the query base can be defined as a field value determined by a row and a column.

[0227] Furthermore, the client can use the selected at least one point as a search condition and construct a search statement based on the search condition and the original field of interest. The original field of interest is the field of interest that the client originally wanted to query, for example, the original intention was to query Name. For example, if the point α·H(34) is selected and the corresponding field name is Age, the constructed search statement can be as follows:

[0228] select Name where Age=α·H(34)

[0229] The client can execute the search statement on the query base and obtain the second identifiers of the matching records as id_4 and id_7 as the second identifier set {id_4, id_7}.

[0230] In addition, you can further select points in the query base other than the ID field and the field of interest, use the selected points as search conditions, construct a search statement based on the search conditions and the original field of interest, and perform the search on the query base. The identifiers of the matching records are added to the second identifier set. For example, the constructed search statement is as follows:

[0231] select Name where Native_place=α·H(anhui)

[0232] The identifiers of the matching records are id_0 and id_2.

[0233] Add id_0 and id_2 to the second identifier set, and obtain the second identifier set: {{id_4, id_7}, {id_0, id_2}}.

[0234] S440: The client constructs a confusion set according to the first identification set and the second identification set.

[0235] The client can combine the first identifier id_1 and the second identifiers id_4 and id_7 into a confusion set, that is, the confusion set is {id_1, id_4, id_7}. This results in a confusion set of n = 3. Further steps S331 to S335 can be performed to make the confusion set more reasonable, thereby preventing client privacy leaks caused by improper confusion set construction. Finally, the server can return the value of the field of interest to the client through 3-choose-1 oblivious transmission.

[0236] The order of the elements in the obfuscation set is not limited. For example, it can be {id_4, id_1, id_7}, {id_4, id_7, id_1}, {id_7, id_1, id_4}, {id_7, id_4, id_1}, {id_1, id_7, id_4}, etc. It can also be random, as long as the actual ID to be transmitted can be specified in the OT later. The following examples are similar and will not be further explained.

[0237] Alternatively, n = 2. For example, if the obfuscation set is {id_1, {id_4, id_7}}, the second element in the obfuscation set is the set {id_4, id_7}. Finally, the server can obliviously transmit the value of the field of interest to the client using a 2-choose-1 strategy. Alternatively, it can be {id_1, {id_4, id_7}, {id_0, id_2}}, where n = 3, or 3-choose-1.

[0238] In addition, if the sensitive field in S410 and S420 is, for example, α·H(25), then the first identification set obtained is {id_1}. For example, the second identification set is the result obtained by respectively searching the constructed search formulas select Name where Age=α·H(46) and selectName where Age=α·H(56), and the second identification set is {id_3, id_9}. Then, the client constructs a confusion set based on the first identification set and the second identification set, which can be formed by flattening the elements in the first identification set and the second identification set. In this way, the confusion set is, for example, {id_1, id_3, id_9}. Furthermore, n-choose-1 oblivious transmission can be used to return the value of the field of real interest to the client, where n=3, and the number of matching records in the confusion set is 1, that is, the value of the Name field in id_0 can be returned to the client through 3-choose-1 oblivious transmission. In addition, when the element in the first identification set is k, n-choose-k oblivious transmission can be used.

[0239] In addition to flattening the elements in the first identification set and the second identification set to form a confusion set, the elements in the first identification set can also be flattened and formed into a confusion set together with the second identification set. For example, as mentioned above, the first identification set is {id_0, id_6}. For example, the second identification set is still {id_3, id_9}, then the confusion set can be {id_0, id_6, {id_3, id_9}}. Furthermore, n-choose-k oblivious transmission can be used to return the value of the field of real interest to the client, where n=3, and the number of matching records in the confusion set is 2, that is, the value of the Name field in id_0 and id_6 can be returned to the client through 3-choose-2 oblivious transmission. When the element in the first identification set is 1, n-choose-k is n-choose-1.

[0240] Similarly, the elements in the second identification set can be flattened and together with the first identification set, form a confusion set. For example, in the above, the first identification set is {id_0, id_6}. For example, the second identification set is the result of the search formula select Name where Age = α·H(46) and select Name where Age = α·H(56) respectively, and the second identification set is {id_3, id_9}. Then the confusion set can be {{id_0, id_6}, id_3, id_9}. Furthermore, n-choose-k oblivious transmission can be used to return the value of the field of real interest to the client, where n = 3, the number of matching records in the confusion set is 1, and the identifications of these two matching records as a set are called an element in the confusion set, that is, the value of the Name field in {id_0, id_6} can be returned to the client through 3-choose-1 oblivious transmission.

[0241] However, the IDs of at least two rows that cannot be distinguished, corresponding to at least one point selected except the ID field and the field of interest, cannot be split into different sets. For example, the first identification set is {id_0, id_6}, and the second identification set is {{id_4, id_7}, id_3, id_9}. Among them, as mentioned above, except for the ID field and the field of interest Name, the remaining fields Age and Native_place cannot distinguish id_4 and id_7, so id_4 and id_7 need to be put into the same set, that is, {id_4, id_7}, and cannot be split. In other words, the identifiers of the rows that cannot be distinguished except for the ID field and the field of interest in the confusion set are placed in the same minimum set, which also meets the requirements of other implementation methods, thereby avoiding the irrationality of the confusion set construction.

[0242] In addition, the elements in the first identification set and the second identification set can be mixed to form a confusion set. For example, as mentioned above, the first identification set is {id_0, id_6}, and the second identification set is still {id_3, id_9}, then the confusion set can be {{id_0, id_6}, {id_0, id_3}, id_9}. id_3 and id_9 can be distinguished in addition to the ID field and the field of interest, so the identifiers of the corresponding rows can also be split, that is, not placed in the same minimum set. Furthermore, n-choose-k oblivious transmission can be used to return the value of the field of real interest to the client, where n=3, the number of matching records in the confusion set is 2, and the identifiers of these two matching records as a set are called an element in the confusion set, that is, the value of the Name field in {id_0, id_6} can be returned to the client through 3-choose-1 oblivious transmission.

[0243] The above example shows a case with one conditional field. The same applies to two or more conditional fields. Specifically, in S430, still using Table 2 as an example, the client selects two points in the query base other than the ID field and the field of interest, specifically, α·H(34) and α·H(shanghai).

[0244] Then, the client can use the two selected points as search conditions and construct a search statement based on the search conditions and the original fields of interest. The original fields of interest are the fields of interest that the client originally wanted to query, for example, the original intention was to query Name. For example, if the two points α·H(34) and α·H(shanghai) are selected, and the corresponding field names are Age and Native_place respectively, the constructed search statement can be as follows:

[0245] select Name where Age=α·H(34)or Native_place=α·H(shanghai)

[0246] The client can execute the search statement on the query base and obtain matching second identifiers id_4, id_7, id_1, id_5 as the second identifier set {id_4, id_7, id_1, id_5}.

[0247] In the above search statement, in the search condition "where," Age = α·H(34) is a predicate in SQL, and Native_place = α·H(shanghai) is another predicate in SQL. The connective between the two predicates is "or" here, but it can also be "and" or something like that. In other words, the connective can be a random connective. More broadly, when there are more than two predicates, the connective between two adjacent predicates is a random connective.

[0248] The following describes a client for constructing a confusion set in an embodiment of this specification. The encryption / decryption performed by the client and the server on the same target uses an encryption / decryption algorithm with an interchangeable order, and:

[0249] The client is configured with a query base, which is obtained by encrypting the database by the server;

[0250] The client sends a sensitive field encrypted by itself to the server, and obtains the same sensitive field encrypted by the server through interaction with the server; searches the query base based on the sensitive field encrypted by the server to obtain a first identification set of matching records; selects at least one point in the query base other than the ID field and the field of interest, uses the selected at least one point as a search condition, constructs a search statement based on the search condition and the original field of interest, and executes the search on the query base to obtain a second identification set of matching records; and constructs a confusion set based on the first identification set and the second identification set.

[0251] The following describes a client for constructing a confusion set in an embodiment of this specification, including:

[0252] processor,

[0253] A memory stores a program, wherein when the processor executes the program, the above Figure 3 The method in .

[0254] The following describes a storage medium in one embodiment of this specification, which is used to store a program, wherein the program, when executed, causes the client to execute the above-mentioned Figure 3 The method in .

[0255] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using "logic compiler" software. This is similar to the software compiler used during program development. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages ​​and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0256] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing various functions included therein can also be considered as structures within the hardware component. Or even, the means for implementing various functions can be considered as both a software module implementing the method and a structure within the hardware component.

[0257] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a server system. Of course, this specification does not exclude that with the future development of computer technology, the computer that implements the functions of the above embodiments may be, for example, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0258] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flow charts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps and does not represent the only execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements includes not only those elements, but also includes other elements that are not clearly listed, or also includes elements inherent to such process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements. For example, if the words first, second, etc. are used to represent the name, they do not represent any particular order.

[0259] For the convenience of description, the above devices are described in terms of functions divided into various modules. Of course, when implementing one or more of the present specifications, the functions of each module can be implemented in the same or multiple software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0260] This specification is described with reference to the flowcharts and / or block diagrams of the methods, apparatus (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0261] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0262] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0263] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0264] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0265] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0266] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0267] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In distributed computing environments, program modules may be located in local and remote computer storage media, including storage devices.

[0268] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between the various embodiments can be referenced across them. Each embodiment focuses on the differences from the other embodiments. In particular, since the system embodiments are generally similar to the method embodiments, their description is relatively simple. For relevant parts, reference can be made to the description of the method embodiments. Throughout this specification, reference to the terms "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, those skilled in the art may combine and integrate the different embodiments or examples, and features of different embodiments or examples, described in this specification, without conflict.

[0269] The foregoing is merely an example of one or more embodiments of this specification and is not intended to limit the one or more embodiments of this specification. It will be apparent to those skilled in the art that various modifications and variations may be made to one or more embodiments of this specification. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this specification shall be included within the scope of the claims.

Claims

1. A method for implementing the construction of a confusion set, where the client receives a query base sent by the server, and the query base is obtained after encrypting a database; the client and the server use an encryption / decryption algorithm whose order can be exchanged for the same target for encryption / decryption. The client sends the sensitive fields encrypted by itself to the server and obtains the same sensitive fields encrypted by the server through interaction with the server. The client retrieves in the query base according to the sensitive fields encrypted by the server and obtains a first set of identifiers of matching records. The client selects at least one point in the query base except for the ID field and the fields of interest, uses the selected at least one point as a retrieval condition, constructs a retrieval statement according to this retrieval condition and the original fields of interest, and executes the retrieval on the query base to obtain a second set of identifiers of matching records. The client constructs a confusion set according to the first set of identifiers and the second set of identifiers.

2. The method according to claim 1, wherein the client constructs a confusion set according to the first set of identifiers and the second set of identifiers. It includes: The client flattens the elements in the first set of identifiers and the second set of identifiers to form a confusion set.

3. The method according to claim 1, wherein the client constructs a confusion set according to the first set of identifiers and the second set of identifiers. It includes: The client flattens the elements in the first set of identifiers and jointly forms a confusion set with the second set of identifiers.

4. The method according to claim 1, wherein the client constructs a confusion set according to the first set of identifiers and the second set of identifiers. It includes: The client flattens the elements in the second set of identifiers and jointly forms a confusion set with the first set of identifiers.

5. The method according to claim 1, wherein the client constructs a confusion set according to the first set of identifiers and the second set of identifiers. It includes: The client mixes the elements in the first set of identifiers and the second set of identifiers to form a confusion set.

6. The method according to claim 1, when the retrieval condition includes at least two predicates, the connective between two adjacent predicates is a randomly selected connective.

7. The method according to claim 1, the order of the elements in the confusion set is a random order.

8. The method according to any one of claims 1-7, the identifiers of the indistinguishable rows in the confusion set except for the ID field and the fields of interest are placed in the same minimum set.

9. A client for implementing the construction of a confusion set, the client and the server use an encryption / decryption algorithm whose order can be exchanged for the same target for encryption / decryption, and: The client is configured with a query base, and the query base is obtained after the server encrypts a database. The client sends the sensitive fields encrypted by itself to the server, and obtains the same sensitive fields encrypted by the server through interaction with the server; retrieves in the query base according to the sensitive fields encrypted by the server to obtain the first identifier set of the matching records; selects at least one point in the query base except the ID field and the fields of interest, uses the selected at least one point as the retrieval condition, constructs a retrieval statement according to the retrieval condition and the original fields of interest, and executes the retrieval on the query base to obtain the second identifier set of the matching records; constructs a confusion set according to the first identifier set and the second identifier set.

10. A client for implementing the construction of a confusion set, comprising: a processor, a memory storing a program, wherein when the processor executes the program, it executes the method according to any one of claims 1-8 above.

11. A storage medium for storing a program, wherein when the program is executed, it causes the client to execute the method according to any one of claims 1-8 above.

Citation Information

Patent Citations

  • Searchable encryption method for hiding search mode and access mode in e-commerce platform

    CN112270006A

  • Data query method and device based on multi-party security computing

    CN113886887A