Method and apparatus for private intersection
By employing secondary encryption between the data source and the querying party, and data bucketing, the problem of time-consuming privacy intersection operations in existing technologies is solved, achieving efficient data intersection querying and privacy protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
- Filing Date
- 2023-03-06
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, when enterprises perform data intersection calculations, the large amount of data makes privacy-preserving intersection operations time-consuming and lacking in timeliness.
By performing secondary encryption between the data source and the querying party, the source identifier is distributed to multiple buckets in one encrypted form using a data bucketing method. Each privacy intersection only needs to process the data in one bucket, and the data intersection query is performed by combining an encryption algorithm that satisfies the commutative law.
This significantly reduces the amount of data processing, improves the timeliness of privacy-preserving intersection operations, and ensures the data privacy protection of the data source.
Smart Images

Figure CN116150806B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of data privacy technology, and more specifically, to a method and apparatus for privacy intersection. Background Technology
[0002] Data privacy protection refers to measures taken to protect sensitive corporate data. Each company, as a data owner, possesses data distinct from other companies, and each company's private data must be protected from leakage to other companies. However, business interactions between different companies are inevitable. In business interactions involving data, it is necessary to first determine the intersection of the data held by both parties. At this point, a privacy intersection operation is required. Privacy intersection finds the intersection of the datasets held by multiple parties while protecting their privacy. Based on the privacy intersection, subsequent business operations can then proceed. Summary of the Invention
[0003] In view of the above, embodiments of this specification provide a method and apparatus for privacy-preserving intersection. Through the technical solutions provided in these embodiments, secondary encryption processing performed during the interaction between the data source and the query party ensures the data privacy of the data source. Furthermore, the source identifier of the data source is assigned to multiple buckets, and each privacy-preserving intersection only requires processing the data in one bucket, greatly reducing the amount of data processing and thus improving the timeliness of the privacy-preserving intersection operation.
[0004] According to one aspect of the embodiments of this specification, a method for privacy-preserving intersection is provided, wherein the method is performed by a data source party, and the source identifiers corresponding to each piece of data owned by the data source party are used to generate corresponding primary ciphertexts of the source identifiers through the data source party's data source key and a commutative encryption algorithm. The generated primary ciphertexts of the source identifiers are distributed into a specified number of buckets according to a data bucketing method, each bucket corresponding to a bucket number. The method includes: obtaining the primary ciphertext of the query identifier corresponding to the query identifier sent by the query party and the query bucket number obtained according to the hash value of the query identifier corresponding to the query identifier and the data bucketing method, wherein the primary ciphertext of the query identifier... The querying party uses its querying key to encrypt the identifier to be queried once using the encryption algorithm; uses the data source key to encrypt the ciphertext of the identifier to be queried a second time using the encryption algorithm to obtain the corresponding ciphertext of the identifier to be queried; obtains the ciphertext of the source identifier in the corresponding bucket according to the bucket number to be queried; and sends the ciphertext of the identifier to be queried and the obtained ciphertext of the source identifier to the querying party, so that the querying party uses the querying key to encrypt the ciphertext of the source identifier a second time using the encryption algorithm to obtain the corresponding ciphertext of the source identifier, and determines the intersection query result based on the ciphertext of the identifier to be queried and the ciphertext of the source identifier.
[0005] According to another aspect of the embodiments of this specification, a method for privacy intersection is also provided, wherein the method is executed by a querying party, and the source identifiers corresponding to each piece of data owned by the data source party are used to generate corresponding ciphertexts of the source identifiers using the data source key of the data source party and an encryption algorithm that satisfies the commutative law. The generated ciphertexts of the source identifiers are distributed into a specified number of buckets according to a data bucketing method, with each bucket corresponding to a bucket number. The method includes: encrypting the identifier to be queried using the querying party's key according to the encryption algorithm to obtain the corresponding ciphertext of the identifier to be queried; sending the ciphertext of the identifier to be queried, and the hash value of the identifier to be queried or the bucket number to be queried, to the data source party, so that the data source party can perform the query based on the identifier to be queried. The process involves obtaining a secondary ciphertext of the corresponding queried identifier from a primary ciphertext and a primary ciphertext of the source identifier from the corresponding bucket based on the bucket number to be queried. The bucket number is obtained according to the hash value of the queried identifier and the data bucketing method. The secondary ciphertext of the queried identifier is obtained by the data source party using the data source key and the encryption algorithm to encrypt the primary ciphertext of the queried identifier a second time. The process also involves receiving the secondary ciphertext of the queried identifier and the primary ciphertext of the source identifier from the corresponding bucket from the data source party; using the query party key and the encryption algorithm to encrypt the received primary ciphertext of the source identifier a second time to obtain the corresponding secondary ciphertext of the source identifier; and determining the intersection query result for the queried identifier based on the secondary ciphertext of the queried identifier and the secondary ciphertext of the source identifier.
[0006] According to another aspect of the embodiments of this specification, a method for privacy-preserving intersection is also provided, wherein the source identifiers corresponding to each piece of data owned by the data source party are used by the data source party's data source key and an encryption algorithm that satisfies the commutative law to generate corresponding ciphertexts of the source identifiers. The generated ciphertexts of the source identifiers are distributed into a specified number of buckets according to a data bucketing method, with each bucket corresponding to a bucket number. The method includes: a querying party using its own querying party key to encrypt the identifier to be queried according to the encryption algorithm to obtain the corresponding ciphertext of the identifier to be queried; the querying party sending the ciphertext of the identifier to be queried, and the hash value of the identifier to be queried or the bucket number to be queried, to the data source party, wherein the bucket number is determined based on the hash value of the identifier to be queried according to the encryption algorithm. The data is obtained through a bucketing method; the data source obtains the bucket number to be queried and the ciphertext of the identifier to be queried sent by the querying party; the data source uses the data source key to perform secondary encryption on the ciphertext of the identifier to be queried according to the encryption algorithm to obtain the corresponding secondary ciphertext of the identifier to be queried; the data source obtains the ciphertext of the source identifier in the corresponding bucket according to the bucket number to be queried; the data source sends the secondary ciphertext of the identifier to be queried and the obtained ciphertext of the source identifier to the querying party; the querying party uses the querying party key to perform secondary encryption on the received ciphertext of the source identifier according to the encryption algorithm to obtain the corresponding secondary ciphertext of the source identifier; and the querying party determines the intersection query result for the identifier to be queried based on the secondary ciphertext of the identifier to be queried and the secondary ciphertext of the source identifier.
[0007] According to another aspect of the embodiments of this specification, an apparatus for privacy intersection is also provided, applied to a data source. The source identifiers corresponding to each piece of data owned by the data source are used by the data source key of the data source to generate corresponding primary ciphertexts of the source identifiers using a commutative encryption algorithm. The generated primary ciphertexts of the source identifiers are distributed into a specified number of buckets according to a data bucketing method, with each bucket corresponding to a bucket number. The apparatus includes: an obtaining unit, which obtains the primary ciphertext of the queried identifier corresponding to the queried identifier sent by the querying party, and the queried bucket number obtained according to the hash value of the queried identifier and the data bucketing method, wherein the primary ciphertext of the queried identifier is obtained by the querying party using its own... The query party's key is used to encrypt the identifier to be queried once using the encryption algorithm. An encryption unit uses the data source key to encrypt the ciphertext of the identifier to be queried a second time using the encryption algorithm to obtain the corresponding ciphertext of the identifier to be queried a second time. A ciphertext acquisition unit obtains the ciphertext of the source identifier in the corresponding bucket according to the bucket number to be queried. A sending unit sends the ciphertext of the identifier to be queried and the obtained ciphertext of the source identifier to the query party, so that the query party uses the query party's key to encrypt the ciphertext of the source identifier a second time using the encryption algorithm to obtain the corresponding ciphertext of the source identifier, and determines the intersection query result based on the ciphertext of the identifier to be queried and the ciphertext of the source identifier.
[0008] According to another aspect of the embodiments of this specification, an apparatus for privacy intersection is also provided, applied to a querying party. The source identifiers corresponding to various data owned by the data source party are used by the data source party's data source key and a commutative encryption algorithm to generate corresponding primary ciphertexts of the source identifiers. The generated primary ciphertexts of the source identifiers are distributed into a specified number of buckets according to a data bucketing method, with each bucket corresponding to a bucket number. The apparatus includes: a primary encryption unit, which uses the querying party's key to perform primary encryption on the identifier to be queried according to the encryption algorithm to obtain the corresponding primary ciphertext of the identifier to be queried; and a sending unit, which sends the primary ciphertext of the identifier to be queried, and the hash value of the identifier to be queried or the bucket number to be queried, to the data source party, so that the data source party can obtain the corresponding primary ciphertext of the identifier to be queried. The system comprises: a secondary ciphertext of the identifier to be queried and a primary ciphertext of the source identifier obtained from the corresponding bucket based on the bucket number to be queried, wherein the bucket number to be queried is obtained according to the data bucketing method based on the hash value of the identifier to be queried, and the secondary ciphertext of the identifier to be queried is obtained by the data source party using the data source key and the encryption algorithm to encrypt the primary ciphertext of the identifier to be queried a second time; a receiving unit receiving the secondary ciphertext of the identifier to be queried and the primary ciphertext of the source identifier obtained from the corresponding bucket from the data source party; a secondary encryption unit using the query party key and the encryption algorithm to encrypt the received primary ciphertext of the source identifier a second time to obtain the corresponding secondary ciphertext of the source identifier; and a query result determination unit determining the intersection query result for the identifier to be queried based on the secondary ciphertext of the identifier to be queried and the secondary ciphertext of the source identifier.
[0009] According to another aspect of the embodiments of this specification, an electronic device is also provided, comprising: at least one processor, a memory coupled to the at least one processor, and a computer program stored on the memory, wherein the at least one processor executes the computer program to implement the method for privacy intersection as described above.
[0010] According to another aspect of the embodiments of this specification, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the privacy intersection method as described above.
[0011] According to another aspect of the embodiments of this specification, a computer program product is also provided, including a computer program that, when executed by a processor, implements the method for privacy intersection as described above. Attached Figure Description
[0012] A further understanding of the nature and advantages of the embodiments described in this specification can be achieved by referring to the following accompanying drawings. In the drawings, similar components or features may have the same reference numerals.
[0013] Figure 1 A flowchart illustrating an example of a data source party performing data bucketing based on a source identifier according to an embodiment of this specification is shown.
[0014] Figure 2 A signaling diagram illustrating an example of a privacy intersection method according to an embodiment of this specification is shown.
[0015] Figure 3 A flowchart illustrating an example of a method for privacy intersection according to another embodiment of this specification is shown.
[0016] Figure 4 A flowchart illustrating an example of a method for privacy intersection according to another embodiment of this specification is shown.
[0017] Figure 5 A block diagram of an example of a privacy intersection apparatus according to another embodiment of this specification is shown.
[0018] Figure 6 A block diagram of an example of a privacy intersection apparatus according to another embodiment of this specification is shown.
[0019] Figure 7 A block diagram of an electronic device for implementing a privacy intersection method according to an embodiment of this specification is shown.
[0020] Figure 8 A block diagram of an electronic device for implementing a privacy intersection method according to another embodiment of this specification is shown. Detailed Implementation
[0021] The subject matter described herein will be discussed below with reference to exemplary embodiments. It should be understood that these embodiments are discussed merely to enable those skilled in the art to better understand and implement the subject matter described herein, and are not intended to limit the scope, applicability, or examples set forth in the claims. The function and arrangement of the elements discussed may be changed without departing from the scope of the embodiments described herein. Various processes or components may be omitted, substituted, or added as needed in the various examples. Furthermore, features described in some examples may be combined in other examples.
[0022] As used herein, the term "comprising" and its variations are open terms meaning "including but not limited to". The term "based on" means "at least partially based on". The terms "one embodiment" and "an embodiment" mean "at least one embodiment". The term "another embodiment" means "at least one other embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other definitions, whether explicit or implicit, may be included below. Unless explicitly indicated by the context, the definition of a term shall remain consistent throughout the specification.
[0023] Data privacy protection refers to measures taken to protect sensitive corporate data. Each company, as a data owner, possesses data distinct from other companies, and each company's private data must be protected from leakage to other companies. However, business interactions between different companies are inevitable. In business interactions involving data, it is necessary to first determine the intersection of the data held by both parties. At this point, a privacy intersection operation is required. Privacy intersection finds the intersection of the datasets held by multiple parties while protecting their privacy. Based on the privacy intersection, subsequent business operations can then proceed.
[0024] However, in current methods for finding the intersection of data from different parties, the large amount of data possessed by one or all parties means that performing privacy-preserving intersection operations on large volumes of data requires a significant amount of time, resulting in low efficiency and unreliable timeliness.
[0025] In view of the above, embodiments of this specification provide a method and apparatus for privacy intersection. The source identifiers corresponding to each piece of data owned by the data source provider are used to generate corresponding primary ciphertexts of the source identifiers using the data source provider's data source key and an encryption algorithm. The generated primary ciphertexts of the source identifiers are then distributed into multiple buckets according to a data bucketing method. In this method, the querying party uses its querying key to encrypt the identifier to be queried once using an encryption algorithm to obtain the corresponding ciphertext of the identifier to be queried; the querying party sends the ciphertext of the identifier to be queried, along with the hash value or bucket number of the identifier to be queried, to the data source party, where the bucket number is obtained according to the data bucketing method based on the hash value of the identifier to be queried; the data source party receives the bucket number and the ciphertext of the identifier to be queried sent by the querying party; the data source party uses its data source key to encrypt the ciphertext of the identifier to be queried a second time using an encryption algorithm to obtain the corresponding ciphertext of the identifier to be queried a second time; the data source party retrieves the ciphertext of the source identifier in the corresponding bucket based on the bucket number; the data source party sends the ciphertext of the identifier to be queried a second time and the retrieved ciphertext of the source identifier to the querying party; the querying party uses its querying key to encrypt the received ciphertext of the source identifier a second time using an encryption algorithm to obtain the corresponding ciphertext of the source identifier; and the querying party determines the intersection query result for the identifier to be queried based on the ciphertext of the identifier to be queried and the ciphertext of the source identifier. The technical solutions provided in the embodiments of this specification ensure data privacy protection for the data source through secondary encryption processing performed during the interaction between the data source and the querying party. Furthermore, the source identifier of the data source is assigned to multiple buckets, and each privacy intersection only requires processing the data in one bucket, significantly reducing the amount of data processing and thus improving the timeliness of the privacy intersection operation.
[0026] The method and apparatus for privacy-based intersection provided in the embodiments of this specification will now be described in detail with reference to the accompanying drawings.
[0027] The technical solutions provided in the embodiments of this specification may include a data source and a querying party, which may include different enterprises or organizations. For the same enterprise or organization, in different application scenarios, it may act as both a data source and a querying party.
[0028] The data source provider owns the data source, which may include a number of data points, including private data. The querying party may or may not own the data. If the querying party owns the data, it can also act as a data source in other business operations or scenarios.
[0029] In the embodiments of this specification, the intersection obtained by privacy intersection can be represented by data identifiers. Privacy intersection can process the data identifiers, and when the identifiers have an intersection, it can be determined that the corresponding data also have an intersection. When an intersection is obtained through privacy intersection, the query party can obtain the intersection, the data source party can obtain the intersection, or both the query party and the data source party can obtain the intersection.
[0030] In one application scenario of this specification's embodiments, the data source and the querying party use the privacy-preserving intersection method provided in this specification's embodiments to find the data intersection, thereby achieving data alignment. Here, data alignment refers to aligning data from parties targeting the same object, which may include users, etc., with each piece of data processed through data alignment targeting the same object. For example, the data source has data 'a' for user A, and the querying party has data 'b' for user A. Data 'a' and data 'b' are different. The privacy-preserving intersection method provided in this specification's embodiments can determine the intersection for user A, thereby aligning data 'a' and data 'b' for user A.
[0031] Once the data source and querying parties align their data, the aligned data can be used for business operations. For example, aligned data for each user can be used for user behavior modeling.
[0032] In another application scenario of the embodiments of this specification, the querying party's query identifier is targeted at a target user. The privacy-preserving intersection method provided in the embodiments of this specification is used to determine whether an intersection exists, thereby determining whether the data source contains data for the target user. If the querying party determines that the data source contains data for the target user, it can perform business operations targeting that target user. Thus, the privacy-preserving intersection method provided in the embodiments of this specification achieves the purpose of querying for the target user.
[0033] For example, the querying party is an advertiser, and the data source is an advertising platform. As an advertiser, the querying party needs to target ads to specific users. Therefore, the querying party needs to confirm whether the data source platform has target users and the size of those target users. Using the privacy-preserving intersection method provided in the embodiments of this specification, the querying party can determine whether there is an intersection of identifiers for the target users, thereby determining the existence and size of the target users within the data source platform. This allows for targeted ad delivery to the target users in subsequent ad delivery operations.
[0034] Figure 1 A flowchart of an example 100 of a data source party performing data bucketing based on a source identifier according to an embodiment of this specification is shown.
[0035] like Figure 1As shown in 110, the data source key of the data source party can be used to encrypt each source identifier once according to the encryption algorithm to obtain the ciphertext of the source identifier corresponding to each source identifier.
[0036] In the embodiments described in this specification, all data owned by the data source can be queried by the querying party, and the data source key of the data source can be the private key held by the data source.
[0037] Each piece of data owned by the data source can be associated with a source identifier. Different data have different source identifiers, which can be used to represent the corresponding data. Different types of data can use different objects as source identifiers. For example, enterprise data can use the unified social credit code as the source identifier, while user data can use their identity verification number as the source identifier.
[0038] All source identifiers owned by the data source are encrypted using the same data source key and encryption algorithm. Each source identifier can be associated with a ciphertext of the source identifier. The source identifiers owned by the data source and the ciphertexts of the source identifiers obtained are in one-to-one correspondence.
[0039] For example, the ciphertext of a source identifier id_b can be represented as E(Akey, id_b), where E represents the encrypted ciphertext and Akey represents the data source key. E(Akey, id_b) represents the ciphertext of the source identifier id_b obtained by encrypting the source identifier id_b once using the data source key Akey.
[0040] In the embodiments of this specification, the encryption algorithm satisfies the commutative law, that is, when the encryption algorithm is used for multiple encryptions, the encryption order does not affect the encryption result.
[0041] For example, in one encryption method, an identifier is first encrypted using the data source key from the data source provider, and then the corresponding ciphertext is further encrypted using the query key from the query provider to obtain a second ciphertext. In another encryption method, the same identifier can be first encrypted using the query key, and then the corresponding ciphertext can be further encrypted using the data source key to obtain a second ciphertext. The two ciphertexts obtained by the two encryption methods are identical. For example, the commutative law of an encryption algorithm can be expressed as follows:
[0042] E(Bkey, E(Akey, id))=E(Akey, E(Bkey, id))
[0043] Where Bkey represents the query key, E(Akey, id) represents the ciphertext obtained by encrypting the identifier id once using the data source key Akey, E(Bkey, id) represents the ciphertext obtained by encrypting the identifier id once using the query key Bkey, E(Bkey, E(Akey, id)) represents the ciphertext obtained by encrypting the ciphertext E(Akey, id) twice using the query key Bkey, and E(Akey, E(Bkey, id)) represents the ciphertext obtained by encrypting the ciphertext E(Bkey, id) twice using the data source key Akey.
[0044] In one example, commutative encryption algorithms may include RSA and ECC (Elliptic Curve Cryptography). The encryption algorithms used in the embodiments of this specification are the same.
[0045] In 120, the obtained source identifier can be distributed to multiple buckets in one go according to the data bucketing method.
[0046] In the embodiments of this specification, data bucketing involves distributing a number of data points into multiple buckets according to a predetermined method, which is a type of data preprocessing. Here, a bucket can be understood as a category of data; data assigned to the same bucket belongs to the same category.
[0047] The number of buckets can be customized; for example, a specified number of buckets can be preset. Each bucket can have a corresponding bucket number, and different buckets have different bucket numbers. In one example, the bucket numbers can be consecutive integers starting from 0, such as 0, 1, 2, ...
[0048] In the embodiments of this specification, data bucketing is performed based on identifiers. For example, the data bucketing method used by the data source is based on the source identifier, and the data bucketing method used by the querying party is based on the identifier to be queried. The objects allocated according to the data bucketing method are the encrypted source identifiers. After allocating the encrypted source identifiers according to the data bucketing method, each bucket can include multiple encrypted source identifiers. The number of encrypted source identifiers included in each bucket can be the same or different. In one example, the data bucketing method used in the various embodiments of this specification is the same.
[0049] In one example, a hash algorithm can be used to calculate the source identifier hash value corresponding to each source identifier.
[0050] In this example, a hash algorithm is used to calculate the hash value of each source identifier, which can be represented as hash(id_b), where hash represents the hash value of the source identifier and id_b represents the source identifier.
[0051] After obtaining the source identifier hash value corresponding to each source identifier, the ciphertext of each source identifier can be distributed to various buckets according to the obtained source identifier hash value and a specified number. The specified number represents the total number of buckets. Each source identifier is distributed to one bucket.
[0052] Based on the one-to-one correspondence between the source identifier and its ciphertext, there is a one-to-one correspondence between the data, the source identifier, and its ciphertext. Furthermore, each source identifier corresponds to a source identifier hash value.
[0053] In one example, for each source identifier hash value, the remainder of that hash value divided by a specified number can be calculated, represented as: (hash(id_b))%(bucket_count), where % represents the remainder calculation and bucket_count represents the specified number of buckets. Then, the resulting remainder value is compared with each bucket number to determine the bucket number that matches the remainder value. The source identifier corresponding to the source identifier hash value is then encrypted and assigned to the bucket corresponding to the determined bucket number.
[0054] For example, if the specified number of buckets is 100, and the bucket numbers are 0, 1, ..., 99, and the source identifier has a hash value of 2 obtained by hashing, then the source identifier hash value 2 is divided by the specified number 100 and the remainder is calculated. The remainder value is 2, so the source identifier corresponding to the source identifier hash value can be ciphertextally assigned to the bucket with bucket number 2.
[0055] In one example, after each source identifier is assigned a ciphertext, the number of source identifiers assigned in each bucket may differ. When the number of source identifiers assigned in each bucket is inconsistent, data padding can be performed on buckets where the number of source identifiers assigned in each bucket is less than the maximum number, so that the number of source identifiers assigned in each bucket is consistent.
[0056] In this example, the maximum number can be obtained by comparing the number of source identifier ciphertexts allocated in each bucket. The maximum number can be the bucket with the most source identifier ciphertexts. Different buckets can have different amounts of data added. For example, if the maximum number is 1000, and a bucket only contains 300 source identifier ciphertexts, then 700 data points need to be added to bring the total number of source identifier ciphertexts in that bucket to 1000.
[0057] In one example, dummy data can be used for data completion. The dummy data is only used to maximize the amount of ciphertext in the source identifier bucket and is not used in subsequent privacy intersection operations.
[0058] By performing data completion operations, it can be ensured that the amount of data in each bucket is consistent, avoiding the risk of information leakage caused by the leakage of statistical information reflected in the allocation results obtained by data bucketing, thereby ensuring data security.
[0059] In one example Figure 1 The illustrated scheme involves the data source performing a data bucketing operation on the encrypted data for each source identifier. This data bucketing operation can be performed offline. Therefore, once both the data source and the querying party are online, and in the online phase, the encrypted data for each source identifier has been allocated. The data source can then directly retrieve the encrypted data from the corresponding bucket and provide it to the querying party. This improves the timeliness of privacy-preserving intersection operations during the online phase.
[0060] Figure 2 A signaling diagram of an example 200 of a privacy intersection method according to an embodiment of this specification is shown.
[0061] exist Figure 2 In the illustrated scheme, the data source and the querying party interact. In one example, Figure 2 The privacy-preserving intersection operation shown can be applied in the online phase, so that both the data source and the querying party performing the privacy-preserving intersection operation are online. By having the data source and the querying party interact in the online phase to execute the privacy-preserving intersection operation provided in the embodiments of this specification, the timeliness of the privacy-preserving intersection operation in the online phase is ensured.
[0062] When the data source is online, the source identifiers corresponding to each piece of data owned by the data source are used to generate corresponding ciphertexts of the source identifiers through the data source key and encryption algorithm. The generated ciphertexts of the source identifiers are then distributed to a specified number of buckets according to the data bucketing method.
[0063] like Figure 2 As shown in 210, the querying party can use its own querying key to encrypt the identifier to be queried once according to the encryption algorithm to obtain the corresponding ciphertext of the identifier to be queried.
[0064] In the embodiments of this specification, the identifier to be queried can be determined by the querying party. The querying party determines whether there is an identifier in the source identifier of the data source party that is the same as the identifier to be queried by performing a privacy intersection operation, thereby determining whether there is data in the data owned by the data source party that matches the data corresponding to the identifier to be queried.
[0065] The query key is a key owned by the querying party, which can be their private key. The encryption algorithm used by the querying party is consistent with the encryption algorithm used by the data source party, and both satisfy the commutative law. In one example, the encryption algorithm can be determined through negotiation between the data source party and the querying party.
[0066] The identifier to be queried corresponds to the ciphertext of the identifier to be queried. The ciphertext of the identifier to be queried can be represented as: E(Bkey, id1), where Bkey represents the query key and id1 represents the identifier to be queried.
[0067] In section 220, the querying party can send the encrypted identifier to be queried, along with the hash value of the identifier or the bucket number to be queried, to the data source.
[0068] In one example, the querying party can send the encrypted queried identifier and the corresponding hash value of the queried identifier to the data source party.
[0069] In this example, the hash value of the queried identifier can be derived from the queried identifier itself. In one example, the querying party can use a hash algorithm to perform a hash calculation on the queried identifier to obtain the corresponding hash value. In one example, the hash algorithm used by the querying party can be the same as the hash algorithm used by the data source. The hash algorithm can be determined through negotiation between the data source and the querying party.
[0070] In another example, the querying party can send the obtained query identifier in encrypted form along with the query bucket number to the data source party.
[0071] In this example, before sending the query, the querying party can process the hash value of the identifier to be queried according to the data bucketing method to obtain the bucket number corresponding to the identifier to be queried.
[0072] In one example of data bucketing, the queryer can use a hash algorithm to perform a hash calculation on the identifier to be queried to obtain the hash value corresponding to that identifier. Then, based on the obtained hash value and a specified number used to represent the total number of buckets, the bucket number corresponding to the hash value of the identifier to be queried is determined. This bucket number is the bucket number corresponding to the identifier to be queried.
[0073] In this example, after obtaining the bucket number corresponding to the identifier to be queried, the querying party can send the encrypted identifier and the bucket number to the data source. In this example, the querying party obtains the bucket number, allowing the data source to directly access and use it without needing to perform the operation of obtaining the bucket number according to the data bucketing method. This reduces the data processing workload for the data source and improves its execution efficiency, thereby enhancing the overall efficiency of privacy-preserving intersection.
[0074] At 230, the data source can obtain the bucket number to be queried and the identifier to be queried once in encrypted form.
[0075] The queried identifier is sent by the querying party in encrypted form. The queried bucket number can be sent directly by the querying party, or it can be obtained by the data source based on the hash value of the queried identifier sent by the querying party.
[0076] In one example, when the querying party sends the encrypted identifier and hash value of the identifier to the data source, the data source receives the hash value and encrypted identifier of the identifier from the querying party. Then, it can process the hash value of the identifier according to the data bucketing method to obtain the bucket number corresponding to the identifier.
[0077] In this example, the querying party can use a hash algorithm to perform a hash calculation on the identifier to be queried to obtain the hash value of the identifier to be queried. The data source party then determines the bucket number corresponding to the hash value of the identifier to be queried based on the obtained hash value and a specified number used to represent the total number of buckets. This bucket number is the bucket number corresponding to the identifier to be queried.
[0078] In step 240, the data source can use the data source key to perform secondary encryption on the ciphertext of the identifier to be queried according to the encryption algorithm, so as to obtain the corresponding secondary ciphertext of the identifier to be queried.
[0079] In one example, the ciphertext of the queried identifier is E(Bkey, id1). The data source uses the data source key Akey to encrypt the ciphertext of the queried identifier a second time. The resulting ciphertext of the queried identifier can be represented as E(Akey, E(Bkey, id1)).
[0080] In 250, the data source can obtain the source identifier of the corresponding bucket once based on the bucket number to be queried.
[0081] The data source can compare the bucket number to be queried with the bucket numbers of various buckets to determine the bucket corresponding to the bucket number that matches the bucket number to be queried. For example, if the bucket number to be queried is 2, then the determined bucket is bucket number 2.
[0082] We can assume that when the data source has a source identifier that is the same as the identifier to be queried, this source identifier will be assigned to the same bucket as the identifier to be queried because they are identical. Based on this, a source identifier identical to the identifier to be queried can only exist in the bucket corresponding to the bucket number of the bucket to be queried. Therefore, subsequent processing only needs to be performed once on the ciphertext of the source identifier in the identified bucket, without processing other buckets. This significantly reduces the data processing workload on the data source side, correspondingly reducing data processing time and improving the timeliness of privacy intersection operations. When privacy intersection operations are applied in the online phase, the timeliness of privacy intersection operations in the online phase can be further improved.
[0083] After identifying the bucket corresponding to the bucket number that matches the bucket number to be queried, the source identifier can be retrieved once from the identified bucket. In one example, if the data source uses fake data to perform data completion, and fake data used for data completion exists in the identified bucket, the fake data can be removed, and only the source identifier in that bucket can be retrieved once.
[0084] It should be noted that the order of operations 240 and 250 is not limited. Figure 2 The operation sequence of 240 and 250 shown is only an example.
[0085] In 260, the data source can send the secondary ciphertext of the identifier to be queried and the primary ciphertext of the obtained source identifier to the querying party.
[0086] After receiving the secondary ciphertext of the identifier to be queried and the primary ciphertext of the source identifier, at 270, the querying party can use its key to perform secondary encryption on each of the received primary ciphertexts of the source identifier according to the encryption algorithm to obtain the corresponding secondary ciphertext of the source identifier.
[0087] In one example, the primary ciphertext of the source identifier received by the querying party is represented as E(Akey, id_b). The querying party then uses its querying key Bkey to perform secondary encryption on the primary ciphertext E(Akey, id_b), resulting in the secondary ciphertext of the source identifier, which can be represented as E(Bkey, E(Akey, id_b)). For instance, on the data source side, the source identifiers corresponding to the determined buckets are id_b1, id_b2, ..., so the primary ciphertexts of the source identifiers received by the querying party are E(Akey, id_b1), E(Akey, id_b2), ..., and correspondingly, the secondary ciphertexts of the source identifiers obtained by the querying party are E(Bkey, E(Akey, id_b1)), E(Bkey, E(Akey, id_b2)), ...
[0088] In 280, the querying party can determine the intersection query result for the target identifier based on the secondary ciphertext of the identifier to be queried and the secondary ciphertext of the source identifier.
[0089] In the embodiments of this specification, the querying party can compare the secondary ciphertext of the identifier to be queried with the secondary ciphertext of each source identifier to determine whether the secondary ciphertext of each source identifier is the same as the secondary ciphertext of the identifier to be queried, and then determine the intersection query result for the data to be queried.
[0090] In one example, if there is a source identifier secondary ciphertext that is the same as the source identifier secondary ciphertext to be queried in the source identifier secondary ciphertext, it can be determined that there is a source identifier in the source identifiers owned by the data source party that is the same as the source identifier to be queried, thereby determining that there is data in the data owned by the data source party that matches the data corresponding to the source identifier to be queried (which can also be called the data to be queried).
[0091] When the secondary ciphertext of the source identifier and the secondary ciphertext of the identifier to be queried are different, it can be determined that the source identifier and the identifier to be queried owned by the data source party are different. Therefore, it can be determined that there is no data in the data owned by the data source party that matches the data corresponding to the identifier to be queried.
[0092] In the embodiments of this specification, the data owned by the data source that matches the data to be queried (hereinafter referred to as source data) and the data to be queried have the same identifier, that is, both are query identifiers. The source data and the data to be queried can be the same data or different data. When the source data and the data to be queried are different, the data features included in the source data can differ from the data features included in the data to be queried. For example, if the data to be queried is for user A, then the source data matching the data to be queried is also for user A, but the data features included in the data to be queried are user A's behavioral characteristics before 2020, while the data features included in the source data matching the data to be queried are user A's behavioral characteristics after 2020.
[0093] Figure 3 A flowchart of an example 300 of a method for privacy intersection according to another embodiment of this specification is shown.
[0094] Figure 3 The privacy intersection method shown is executed by the data source. The source identifiers corresponding to each data owned by the data source are used to generate corresponding ciphertexts of the source identifiers through the data source key of the data source and an encryption algorithm that satisfies the commutative law. The generated ciphertexts of the source identifiers are distributed to a specified number of buckets according to the data bucketing method, and each bucket has a corresponding bucket number.
[0095] like Figure 3As shown in section 310, the ciphertext of the queried identifier corresponding to the queried identifier sent by the querying party, and the queried bucket number obtained according to the data bucketing method based on the hash value of the queried identifier, can be obtained. The ciphertext of the queried identifier is obtained by the querying party encrypting the queried identifier once using its own querying key according to the encryption algorithm.
[0096] In one example, the querying party can receive the encrypted text of the query identifier and the query bucket number corresponding to the query identifier. The query bucket number is obtained by the querying party processing the hash value of the query identifier corresponding to the query identifier according to the data bucketing method.
[0097] In one example, the queryer can receive the hash value of the identifier to be queried and the encrypted one-time ciphertext of the identifier to be queried; and process the hash value of the identifier to be queried according to the data bucketing method to obtain the corresponding bucket number to be queried.
[0098] In 320, the data source key can be used to encrypt the ciphertext of the identifier to be queried a second time according to the encryption algorithm to obtain the corresponding ciphertext of the identifier to be queried.
[0099] At 330, the source identifier of the corresponding bucket can be obtained once based on the bucket number to be queried.
[0100] In step 340, the secondary ciphertext of the identifier to be queried and the primary ciphertext of the source identifier obtained can be sent to the querying party, so that the querying party can use the querying party's key to perform secondary encryption on the primary ciphertext of the source identifier according to the encryption algorithm to obtain the corresponding secondary ciphertext of the source identifier, and determine the intersection query result based on the secondary ciphertext of the identifier to be queried and the secondary ciphertext of the source identifier.
[0101] In one example, the data source and the querying party are online when they interact to perform a privacy-preserving intersection operation.
[0102] In one example, the data source can also use the data source key to encrypt each of its source identifiers once according to the encryption algorithm to obtain the ciphertext of the source identifier corresponding to each source identifier; and distribute the obtained ciphertext of the source identifier to multiple buckets according to the data bucketing method.
[0103] In one example, the data source can also use a hash algorithm to calculate the source identifier hash value corresponding to each source identifier; and distribute each source identifier to each bucket in one ciphertext according to the obtained source identifier hash value and a specified number.
[0104] In one example, after the ciphertext of each source identifier is allocated, if the number of ciphertexts allocated to each bucket is inconsistent, the data source can also fill in the data in the buckets where the number of ciphertexts allocated to each source identifier has not reached the maximum number, so that the number of ciphertexts allocated to each bucket is consistent.
[0105] Figure 4 A flowchart of an example 400 of a method for privacy intersection according to another embodiment of this specification is shown.
[0106] Figure 4 The privacy-preserving intersection method shown is executed by the querying party. The source identifiers corresponding to each piece of data owned by the data source party are used to generate corresponding ciphertexts of the source identifiers through the data source party's data source key and an encryption algorithm that satisfies the commutative law. The generated ciphertexts of the source identifiers are distributed to a specified number of buckets according to the data bucketing method, and each bucket has a corresponding bucket number.
[0107] like Figure 4 As shown in 410, the query key can be used to encrypt the identifier to be queried once according to the encryption algorithm to obtain the corresponding ciphertext of the identifier to be queried.
[0108] At 420, the primary ciphertext of the identifier to be queried, as well as the hash value of the identifier to be queried or the bucket number to be queried, can be sent to the data source. This allows the data source to obtain the secondary ciphertext of the identifier to be queried based on the primary ciphertext and to retrieve the primary ciphertext of the source identifier in the corresponding bucket based on the bucket number. The bucket number to be queried is obtained by the data source using the hash value of the identifier to be queried according to the data bucketing method. The secondary ciphertext of the identifier to be queried is obtained by the data source using the data source key and the encryption algorithm to encrypt the primary ciphertext of the identifier to be queried a second time.
[0109] At 430, the secondary ciphertext of the identifier to be queried and the primary ciphertext of the source identifier obtained from the corresponding bucket can be received from the data source.
[0110] At 440, the query key can be used to encrypt the received source identifier ciphertext a second time according to the encryption algorithm to obtain the corresponding source identifier ciphertext a second time.
[0111] At 450, the intersection query result for the target identifier can be determined based on the secondary ciphertext of the target identifier and the secondary ciphertext of the source identifier.
[0112] In one example, when the querying party sends the bucket number to be queried to the data source, before sending the bucket number, the querying party can also process the hash value of the identifier to be queried according to the data bucketing method to obtain the bucket number corresponding to the identifier to be queried.
[0113] In one example, if the source identifier's secondary ciphertext contains a source identifier secondary ciphertext that is identical to the query identifier's secondary ciphertext, the querying party can also determine that the data source contains data matching the query identifier. If both the source identifier's secondary ciphertext and the query identifier's secondary ciphertext are different, the querying party can also determine that the data source does not contain data matching the query identifier.
[0114] In one example, the data source and the querying party are online when they interact to perform a privacy-preserving intersection operation.
[0115] Figure 5 A block diagram of an example of a privacy intersection apparatus 500 according to another embodiment of this specification is shown.
[0116] Figure 5 The privacy intersection device 500 shown is applied to the data source side. The source identifiers corresponding to each data owned by the data source side are used to generate corresponding ciphertexts of the source identifiers through the data source key of the data source side and the encryption algorithm that satisfies the commutative law. The generated ciphertexts of the source identifiers are distributed to a specified number of buckets according to the data bucketing method, and each bucket has a corresponding bucket number.
[0117] like Figure 5 As shown, the privacy-seeking device 500 includes: an acquisition unit 510, an encryption unit 520, a ciphertext acquisition unit 530, and a sending unit 540.
[0118] The obtaining unit 510 can be configured to obtain the ciphertext of the query identifier corresponding to the query identifier sent by the querying party, and the query bucket number obtained according to the hash value of the query identifier and the data bucketing method. The ciphertext of the query identifier is obtained by the querying party encrypting the query identifier once using its own querying key according to the encryption algorithm.
[0119] In one example, the obtaining unit 510 can also be configured to: receive from the querying party the encrypted text of the query identifier corresponding to the query identifier and the query bucket number, wherein the query bucket number is obtained by the querying party processing the hash value of the query identifier corresponding to the query identifier according to the data bucketing method.
[0120] In one example, the obtaining unit 510 can also be configured to: receive the hash value of the identifier to be queried and the ciphertext of the identifier to be queried from the querying party; and process the hash value of the identifier to be queried according to the data bucketing method to obtain the corresponding bucket number to be queried.
[0121] The encryption unit 520 can be configured to use the data source key to perform secondary encryption on the ciphertext of the query identifier according to the encryption algorithm, so as to obtain the corresponding secondary ciphertext of the query identifier.
[0122] The ciphertext acquisition unit 530 can be configured to retrieve the source identifier of the corresponding bucket once based on the bucket number to be queried.
[0123] The sending unit 540 can be configured to send the secondary ciphertext of the identifier to be queried and the primary ciphertext of the source identifier obtained to the querying party, so that the querying party can use the querying party key to perform secondary encryption on the primary ciphertext of the source identifier according to the encryption algorithm to obtain the corresponding secondary ciphertext of the source identifier, and determine the intersection query result based on the secondary ciphertext of the identifier to be queried and the secondary ciphertext of the source identifier.
[0124] In one example, the data source and the querying party are online when they interact to perform a privacy-preserving intersection operation.
[0125] In one example, the privacy-sharing device 500 may also include: a data bucketing unit.
[0126] The encryption unit 520 can also be configured to encrypt each source identifier using the data source key and an encryption algorithm to obtain a ciphertext of the source identifier corresponding to each source identifier. The data bucketing unit can be configured to distribute the obtained ciphertext of the source identifier into multiple buckets according to the data bucketing method.
[0127] In one example, the data bucketing unit can also be configured to: calculate the source identifier hash value corresponding to each source identifier using a hash algorithm; and distribute each source identifier to each bucket in one ciphertext according to the obtained source identifier hash value and a specified number.
[0128] In one example, the privacy intersection device 500 may further include a data completion unit. This data completion unit can be configured to: after the ciphertext allocation of each source identifier is completed, when the number of ciphertexts allocated to each bucket is inconsistent, perform data completion on the buckets where the number of ciphertexts allocated to the source identifier has not reached the maximum number, so that the number of ciphertexts allocated to the source identifier in each bucket is consistent.
[0129] Figure 6 A block diagram of an example of a privacy intersection apparatus 600 according to another embodiment of this specification is shown.
[0130] Figure 6 The privacy intersection device 600 shown is applied to the querying party. The source identifiers corresponding to each piece of data owned by the data source party are generated into corresponding ciphertexts using the data source key of the data source party and an encryption algorithm that satisfies the commutative law. The generated ciphertexts of each source identifier are distributed into a specified number of buckets according to the data bucketing method, and each bucket has a corresponding bucket number.
[0131] like Figure 6As shown, the privacy query device 600 includes: a primary encryption unit 610, a sending unit 620, a receiving unit 630, a secondary encryption unit 640, and a query result determination unit 650.
[0132] The encryption unit 610 can be configured to encrypt the identifier to be queried once using the query key according to the encryption algorithm to obtain the corresponding ciphertext of the identifier to be queried.
[0133] The sending unit 620 can be configured to send the primary ciphertext of the identifier to be queried, as well as the hash value of the identifier to be queried or the bucket number to be queried, to the data source party, so that the data source party can obtain the secondary ciphertext of the identifier to be queried based on the primary ciphertext of the identifier to be queried, and obtain the primary ciphertext of the source identifier in the corresponding bucket based on the bucket number to be queried. The bucket number to be queried is obtained according to the data bucketing method based on the hash value of the identifier to be queried. The secondary ciphertext of the identifier to be queried is obtained by the data source party using the data source key to encrypt the primary ciphertext of the identifier to be queried twice according to the encryption algorithm.
[0134] The receiving unit 630 can be configured to receive the secondary ciphertext of the identifier to be queried and the primary ciphertext of the source identifier obtained from the corresponding bucket from the data source.
[0135] The secondary encryption unit 640 can be configured to perform secondary encryption on the received primary ciphertext of the source identifier using the query key and an encryption algorithm to obtain the corresponding secondary ciphertext of the source identifier. In one example, the primary encryption unit 610 and the secondary encryption unit 640 can be the same encryption unit.
[0136] The query result determination unit 650 can be configured to determine the intersection query result for the identifier to be queried based on the secondary ciphertext of the identifier to be queried and the secondary ciphertext of the source identifier.
[0137] In one example, the privacy intersection device 600 may further include a data bucketing unit, which can be configured to: when the querying party sends the bucket number to be queried to the data source, process the hash value of the identifier to be queried according to the data bucketing method before sending the bucket number to be queried, so as to obtain the bucket number to be queried corresponding to the identifier to be queried.
[0138] In one example, the query result determination unit 650 can also be configured to: when a source identifier secondary ciphertext exists that is identical to the queried identifier secondary ciphertext, the querying party can further determine that the data source party contains data matching the queried identifier. When both the source identifier secondary ciphertext and the queried identifier secondary ciphertext are different, the querying party can further determine that the data source party does not contain data matching the queried identifier.
[0139] In one example, the data source and the querying party are online when they interact to perform a privacy-preserving intersection operation.
[0140] Reference above Figures 1 to 6 Embodiments of the method and apparatus for privacy intersection according to the embodiments of this specification have been described.
[0141] The privacy intersection apparatus described in this specification can be implemented in hardware, software, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of its host device reading the corresponding computer program instructions from the memory into memory and executing them. In the embodiments of this specification, the privacy intersection apparatus can be implemented, for example, using an electronic device.
[0142] Figure 7 A block diagram of an electronic device 700 for implementing a privacy intersection method according to an embodiment of this specification is shown.
[0143] like Figure 7 As shown, the electronic device 700 may include at least one processor 710, a memory (e.g., non-volatile memory) 720, a RAM 730, and a communication interface 740, and the at least one processor 710, memory 720, RAM 730, and communication interface 740 are connected together via a bus 750. The at least one processor 710 executes at least one computer-readable instruction (i.e., the elements implemented in software above) stored or encoded in the memory.
[0144] In one embodiment, computer-executable instructions are stored in a memory, which, when executed, cause at least one processor 710 to: obtain a primary ciphertext of a query identifier corresponding to a query identifier sent by a querying party, and a query bucket number obtained according to the hash value of the query identifier and a data bucketing method, wherein the primary ciphertext of the query identifier is obtained by the querying party encrypting the query identifier once using its own querying key according to an encryption algorithm; encrypting the primary ciphertext of the query identifier a second time using a data source key according to an encryption algorithm to obtain a corresponding secondary ciphertext of the query identifier; obtaining the primary ciphertext of the source identifier in the corresponding bucket according to the query bucket number; and sending the secondary ciphertext of the query identifier and the obtained primary ciphertext of the source identifier to the querying party, so that the querying party encrypts the primary ciphertext of the source identifier a second time using its querying key according to an encryption algorithm to obtain a corresponding secondary ciphertext of the source identifier; and determining an intersection query result based on the secondary ciphertext of the query identifier and the secondary ciphertext of the source identifier.
[0145] Figure 8 A block diagram of an electronic device 800 for implementing a privacy intersection method according to another embodiment of this specification is shown.
[0146] like Figure 8As shown, the electronic device 800 may include at least one processor 810, a memory (e.g., non-volatile memory) 820, a RAM 830, and a communication interface 840, and the at least one processor 810, memory 820, RAM 830, and communication interface 840 are connected together via a bus 850. The at least one processor 810 executes at least one computer-readable instruction (i.e., the elements implemented in software above) stored or encoded in the memory.
[0147] In one embodiment, computer-executable instructions are stored in a memory, which, when executed, cause at least one processor 810 to: encrypt the query identifier using a query key and an encryption algorithm to obtain a corresponding ciphertext of the query identifier; send the ciphertext of the query identifier, along with the hash value of the query identifier or the bucket number corresponding to the query identifier, to a data source, so that the data source obtains a corresponding secondary ciphertext of the query identifier based on the ciphertext of the query identifier and retrieves the ciphertext of the source identifier in the corresponding bucket based on the bucket number, wherein the bucket number is obtained according to the hash value of the query identifier in a data bucketing manner, and the secondary ciphertext of the query identifier is obtained by the data source using a data source key and an encryption algorithm to encrypt the ciphertext of the query identifier a second time; receive the secondary ciphertext of the query identifier and the ciphertext of the source identifier retrieved from the corresponding bucket from the data source; encrypt the received ciphertext of the source identifier a second time using the query key and an encryption algorithm to obtain a corresponding secondary ciphertext of the source identifier; and determine the intersection query result for the query identifier based on the secondary ciphertext of the query identifier and the secondary ciphertext of the source identifier.
[0148] It should be understood that the computer-executable instructions stored in the memory, when executed, cause at least one processor 710 and processor 810 to perform the above combinations as described in the various embodiments of this specification. Figure 1-7 The description includes various operations and functions.
[0149] According to one embodiment, a program product, such as a machine-readable medium, is provided. The machine-readable medium may have instructions (i.e., the elements implemented in software as described above), which, when executed by a machine, cause the machine to perform the above-described combinations of the various embodiments of this specification. Figure 1-7 The description includes various operations and functions.
[0150] Specifically, a system or apparatus equipped with a readable storage medium may be provided, on which software program code implementing the functions of any of the embodiments described above is stored, and the computer or processor of the system or apparatus can read and execute the instructions stored in the readable storage medium.
[0151] In this case, the program code itself, which can be read from the readable medium, can perform the functions of any of the above embodiments. Therefore, the machine-readable code and the readable storage medium storing the machine-readable code constitute part of the present invention.
[0152] The computer program code required for the operation of each part of this manual can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB, .NET, and Python; conventional procedural programming languages such as C, Visual Basic 2003, Perl, COBOL 2002, PHP, and ABAP; dynamic programming languages such as Python, Ruby, and Groovy; or other programming languages. This program code can run on the user's computer, or as a standalone software package on the user's computer, or partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any network, such as a local area network (LAN) or wide area network (WAN), or connected to an external computer (e.g., via the Internet), or in a cloud computing environment, or used as a service, such as Software as a Service (SaaS).
[0153] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer or the cloud via a communication network.
[0154] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0155] Not all steps and units in the above process and system structure diagrams are mandatory; some steps or units can be omitted as needed. The execution order of each step is not fixed and can be determined as required. The device structure described in the above embodiments can be a physical structure or a logical structure. That is, some units may be implemented by the same physical entity, or some units may be implemented by multiple physical entities, or they may be jointly implemented by certain components in multiple independent devices.
[0156] The term "exemplary" as used throughout this specification means "serving as an example, instance, or illustration" and does not imply that it is "preferred" or "advantageous" over other embodiments. Detailed descriptions are included for the purpose of providing an understanding of the described techniques. However, these techniques may be practiced without these detailed descriptions. In some instances, well-known structures and apparatuses are shown in block diagram form to avoid obscuring the concepts of the described embodiments.
[0157] The optional embodiments of the present specification have been described in detail above with reference to the accompanying drawings. However, the embodiments of the present specification are not limited to the specific details in the above embodiments. Within the scope of the technical concept of the embodiments of the present specification, various simple modifications can be made to the technical solutions of the embodiments of the present specification, and these simple modifications all fall within the protection scope of the embodiments of the present specification.
[0158] The foregoing description of this specification is provided to enable any person skilled in the art to implement or use the content of this specification. Various modifications to the content of this specification will be apparent to those skilled in the art, and the general principles defined herein can be applied to other variations without departing from the scope of protection of this specification. Therefore, this specification is not limited to the examples and designs described herein, but is consistent with the widest scope of the principles and novel features disclosed herein.
Claims
1. A method for privacy-preserving intersection, wherein, The method is executed by the data source provider. Each data source identifier corresponding to a given data source is used by the data source provider's data source key and a commutative encryption algorithm to generate a corresponding ciphertext. These ciphertexts are then distributed into a specified number of buckets according to a data bucketing method, with each bucket having a corresponding bucket number. The method includes: The query party obtains the ciphertext of the query identifier corresponding to the query identifier and the query bucket number obtained according to the data bucketing method based on the hash value of the query identifier corresponding to the query identifier. The ciphertext of the query identifier is obtained by the query party encrypting the query identifier once using its own query key according to the encryption algorithm. Using the data source key, the ciphertext of the identifier to be queried is encrypted twice according to the encryption algorithm to obtain the corresponding ciphertext of the identifier to be queried. Based on the bucket number to be queried, obtain the source identifier ciphertext from the corresponding bucket; and The secondary ciphertext of the identifier to be queried and the primary ciphertext of the source identifier obtained are sent to the querying party, so that the querying party can use the querying party key to perform secondary encryption on the primary ciphertext of the source identifier according to the encryption algorithm to obtain the corresponding secondary ciphertext of the source identifier, and determine the intersection query result based on the secondary ciphertext of the identifier to be queried and the secondary ciphertext of the source identifier.
2. The method as described in claim 1, wherein, The data includes the encrypted message of the query identifier sent by the querying party, and the query bucket number obtained according to the hash value of the query identifier and the data bucketing method. The querying party receives the encrypted text of the identifier to be queried corresponding to the identifier to be queried and the bucket number to be queried, wherein the bucket number to be queried is obtained by the querying party processing the hash value of the identifier to be queried corresponding to the identifier to be queried according to the data bucketing method.
3. The method as described in claim 1, wherein, The data includes the encrypted message of the query identifier sent by the querying party, and the query bucket number obtained according to the hash value of the query identifier and the data bucketing method. The querying party receives the hash value of the query identifier corresponding to the query identifier and a ciphertext of the query identifier; and The hash value of the identifier to be queried is processed according to the data bucketing method to obtain the corresponding bucket number to be queried.
4. The method as described in any one of claims 1 to 3, wherein, The data source and the querying party are online when they interact to perform a privacy-preserving intersection operation.
5. The method of claim 1, further comprising: Using the data source key, each source identifier is encrypted once according to the encryption algorithm to obtain the ciphertext of the source identifier corresponding to each source identifier; as well as The obtained source identifier is distributed into the multiple buckets in one go according to the data bucketing method.
6. The method of claim 5, wherein, The process of distributing the obtained source identifier ciphertext to the multiple buckets in one go according to the data bucketing method includes: A hash algorithm is used to calculate the source identifier hash value corresponding to each source identifier; and Based on the obtained source identifier hash value and the specified quantity, each source identifier is distributed to the respective bucket in a single ciphertext.
7. The method of claim 5, further comprising: After the ciphertext of each source identifier is allocated, if the number of ciphertexts allocated to each bucket is inconsistent, the buckets with fewer ciphertexts allocated to each source identifier will be padded with data to make the number of ciphertexts allocated to each bucket consistent.
8. The method as described in any one of claims 5 to 7, wherein, The data source performs a data bucketing operation on each source identifier's encrypted data during the offline phase.
9. A method for privacy-preserving intersection, wherein, The method is executed by the querying party. The source identifiers corresponding to each piece of data owned by the data source party are used to generate corresponding ciphertexts using the data source party's data source key and a commutative encryption algorithm. These ciphertexts are then distributed into a specified number of buckets according to a data bucketing method, with each bucket corresponding to a bucket number. The method includes: The query identifier is encrypted once using the query key according to the encryption algorithm to obtain the corresponding ciphertext of the query identifier; The ciphertext of the identifier to be queried, and the hash value or bucket number of the identifier to be queried, are sent to the data source, so that the data source can obtain the corresponding ciphertext of the identifier to be queried based on the ciphertext of the identifier to be queried and obtain the ciphertext of the source identifier in the corresponding bucket based on the bucket number. The bucket number is obtained according to the hash value of the identifier to be queried according to the data bucketing method. The ciphertext of the identifier to be queried is obtained by the data source using the data source key to encrypt the ciphertext of the identifier to be queried twice according to the encryption algorithm. Receive the secondary ciphertext of the identifier to be queried and the primary ciphertext of the source identifier obtained from the corresponding bucket from the data source. Using the query key, the received source identifier primary ciphertext is encrypted a second time according to the encryption algorithm to obtain the corresponding source identifier secondary ciphertext; and The intersection query result for the target identifier is determined based on the secondary ciphertext of the identifier to be queried and the secondary ciphertext of the source identifier.
10. The method of claim 9, wherein, When the querying party sends the bucket number to be queried to the data source, the method further includes, before sending the bucket number to be queried: The hash value of the identifier to be queried is processed according to the data bucketing method to obtain the bucket number corresponding to the identifier to be queried.
11. The method of claim 9, wherein, The intersection query result for the target identifier is determined based on the secondary ciphertext of the identifier to be queried and the secondary ciphertext of the source identifier, including: If a source identifier secondary ciphertext exists in the source identifier secondary ciphertext that is identical to the query identifier secondary ciphertext, it is determined that the data source contains data that matches the query identifier. When the secondary ciphertext of the source identifier is different from the secondary ciphertext of the identifier to be queried, it is determined that the data source does not have data matching the identifier to be queried.
12. The method as described in any one of claims 9 to 11, wherein, The data source and the querying party are online when they interact to perform a privacy-preserving intersection operation.
13. A method for privacy-preserving intersection, wherein, Each data source possessed by the data source provider generates a corresponding ciphertext using the data source provider's data source key and a commutative encryption algorithm. These ciphertexts are then distributed into a specified number of buckets according to a data bucketing method, with each bucket having a corresponding bucket number. The method includes: The querying party uses its own querying key to encrypt the identifier to be queried once according to the encryption algorithm to obtain the corresponding ciphertext of the identifier to be queried; The querying party sends the encrypted text of the identifier to be queried, as well as the hash value of the identifier to be queried or the bucket number to be queried corresponding to the identifier to be queried, to the data source party. The bucket number to be queried is obtained according to the data bucketing method based on the hash value of the identifier to be queried. The data source receives the bucket number to be queried and the encrypted identifier sent by the querying party. The data source provider uses the data source key to perform secondary encryption on the ciphertext of the identifier to be queried according to the encryption algorithm, so as to obtain the corresponding secondary ciphertext of the identifier to be queried. The data source party obtains the source identifier encrypted once from the corresponding bucket based on the bucket number to be queried. The data source sends the secondary ciphertext of the identifier to be queried and the primary ciphertext of the obtained source identifier to the querying party; The querying party uses its key to perform secondary encryption on the received source identifier ciphertext according to the encryption algorithm, to obtain the corresponding source identifier secondary ciphertext; and The querying party determines the intersection query result for the target identifier based on the secondary ciphertext of the identifier to be queried and the secondary ciphertext of the source identifier.
14. An apparatus for privacy-preserving intersection, applied to a data source, wherein the source identifiers corresponding to each piece of data owned by the data source are used to generate corresponding ciphertexts of the source identifiers using the data source key of the data source and an encryption algorithm that satisfies the commutative law; the generated ciphertexts of the source identifiers are distributed into a specified number of buckets according to a data bucketing method, and each bucket has a corresponding bucket number. The device includes: The obtaining unit obtains the first ciphertext of the query identifier corresponding to the query identifier sent by the querying party and the query bucket number obtained according to the hash value of the query identifier corresponding to the query identifier and the data bucketing method. The first ciphertext of the query identifier is obtained by the querying party encrypting the query identifier once using its own querying key according to the encryption algorithm. The encryption unit uses the data source key to perform secondary encryption on the ciphertext of the identifier to be queried according to the encryption algorithm, so as to obtain the corresponding secondary ciphertext of the identifier to be queried. The ciphertext acquisition unit retrieves the source identifier ciphertext from the corresponding bucket based on the bucket number to be queried; and The sending unit sends the secondary ciphertext of the identifier to be queried and the primary ciphertext of the source identifier obtained to the querying party, so that the querying party can use the querying party key to perform secondary encryption on the primary ciphertext of the source identifier according to the encryption algorithm to obtain the corresponding secondary ciphertext of the source identifier, and determine the intersection query result based on the secondary ciphertext of the identifier to be queried and the secondary ciphertext of the source identifier.
15. An apparatus for privacy-preserving intersection, applied to a querying party, wherein source identifiers corresponding to various data owned by a data source party are used to generate corresponding ciphertexts of source identifiers using the data source party's data source key and a commutative encryption algorithm; the generated ciphertexts of source identifiers are distributed into a specified number of buckets according to a data bucketing method, and each bucket has a corresponding bucket number. The device includes: A single encryption unit uses the query key to encrypt the identifier to be queried once according to the encryption algorithm to obtain the corresponding ciphertext of the identifier to be queried; The sending unit sends the primary ciphertext of the identifier to be queried, and the hash value of the identifier to be queried or the bucket number to be queried corresponding to the identifier to be queried, to the data source party, so that the data source party can obtain the secondary ciphertext of the identifier to be queried based on the primary ciphertext of the identifier to be queried, and obtain the primary ciphertext of the source identifier in the corresponding bucket based on the bucket number. The bucket number to be queried is obtained according to the hash value of the identifier to be queried according to the data bucketing method. The secondary ciphertext of the identifier to be queried is obtained by the data source party using the data source key to encrypt the primary ciphertext of the identifier to be queried twice according to the encryption algorithm. The receiving unit receives the secondary ciphertext of the identifier to be queried and the primary ciphertext of the source identifier obtained from the corresponding bucket from the data source. The secondary encryption unit uses the query key to perform secondary encryption on the received primary ciphertext of the source identifier according to the encryption algorithm, so as to obtain the corresponding secondary ciphertext of the source identifier; and The query result determination unit determines the intersection query result for the target identifier based on the secondary ciphertext of the identifier to be queried and the secondary ciphertext of the source identifier.
16. An electronic device comprising: At least one processor, a memory coupled to the at least one processor, and a computer program stored on the memory, wherein the at least one processor executes the computer program to implement the method as described in any one of claims 1-12.
17. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method as described in any one of claims 1-12.
18. A computer program product comprising a computer program that, when executed by a processor, implements the method as described in any one of claims 1-12.
Citation Information
Patent Citations
Inter-institution privacy data query and early warning method based on multi-party security computing
CN113515538A
Privacy-protecting data providing and querying method, device and system
CN115098868A