Method and apparatus for private set intersection

In the privacy submission process of multi-party security calculation, the data is distinguished by using markers and multi-level privacy submission, the problem of excessive data volume and communication volume in the existing technology is solved, and efficient privacy submission is achieved.

CN114036572BActive Publication Date: 2025-06-24SASI DIGITAL TECHNOLOGY (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111436332.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2025-06-24
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

In the privacy request process of multi-party security calculations, the amount of data and communications are too large, resulting in high usage of computer resources and low efficiency.

Method used

By using the mark code to modulo the predetermined numerical values, the data identification of the modulo value is obtained, and privacy submission is conducted in a secure manner to obtain a common identification. Then, further privacy submission is carried out on the data under the same distinction identification and split into multi-level processes to reduce the data processing volume and complexity.

Benefits of technology

It significantly reduces the processing volume and complexity of data for privacy submissions, improves efficiency, and is particularly effective in the case of huge data disparity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114036572B_ABST
    Figure CN114036572B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a private set intersection method and apparatus. During the private set intersection process of multi-party secure computing, the two data parties performing the private set intersection can respectively perform modulo operations on the local data using a marking code for a predetermined value to distinguish the data according to the modulo value. Then, the two data parties perform a private set intersection on the obtained modulo values in a secure manner, and the intersection is the same modulo value. Then, the data under the same modulo value can be subjected to private set intersection. This method splits the private set intersection process of the data into at least two levels. One level is to perform private set intersection on the modulo values to filter out the data that may exist in the intersection with the other party, and the other level is to perform private set intersection on the data under the same modulo value. At this time, the amount of data has been greatly reduced after filtering. This method can reduce the data processing amount and complexity of private set intersection and improve the efficiency of private set intersection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of computer technology, and in particular, to a method and apparatus for private set intersection. Background Art

[0002] Secure multi-party computation (MPC for short), also known as multi-party secure computation, means that multiple parties jointly calculate the result of a function without revealing the input data of each party to this function, and the calculation result is disclosed to one or more of the parties. For example, a typical application of secure multi-party computation is private set intersection. Private set intersection (PSI) can be understood as determining the data intersection among multiple parties on the premise of privacy protection. Private set intersection is often the core of multi-party collaborative training of machine learning algorithms and doing businesses such as multi-head lending. The core idea of private set intersection is that at the end of the protocol interaction, one or more parties should obtain the correct intersection and not obtain any other data in the data sets of other parties outside the intersection. During the process of private set intersection, the amount of data and communication directly affect the usage of computer resources and the efficiency of private set intersection. Summary of the Invention

[0003] One or more embodiments of this specification describe a method and apparatus for private set intersection to solve one or more problems mentioned in the background art.

[0004] According to a first aspect, a method for private set intersection is provided, which is used to determine the data intersection of N first data held by a first party and M second data held by a second party while protecting data privacy; the method includes: the first party respectively takes the modulo of a numerical value L with N marking codes corresponding one-to-one to the N first data to obtain n modulo values corresponding to n first identifiers; the second party respectively takes the modulo of the numerical value L with M marking codes corresponding one-to-one to the M second data to obtain m modulo values corresponding to m second identifiers; the first party and the second party perform a private set intersection operation on the n first identifiers and the m second identifiers to obtain s common identifiers among the n first identifiers and the m second identifiers; the first party and the second party perform a private set intersection operation on the P first data and the Q second data corresponding to their respective s common identifiers to obtain the data intersection of the N first data and the M second data.

[0005] In one embodiment, the private set intersection can be implemented by at least one of secret sharing, homomorphic encryption, garbled circuits, and oblivious transfer.

[0006] In one embodiment, the first party and the second party perform a private intersection operation on the P first data and the Q second data respectively corresponding to the s common identifiers, and the obtained data intersection of the N first data and the M second data includes: the first party and the second party respectively determine the number P of the first data and the number Q of the second data corresponding to the s common identifiers locally; the first party and the second party detect whether the difference between P and Q meets a predetermined condition through a secure comparison; the first party and the second party perform a private intersection operation on the P first data and the Q second data based on the detection result.

[0007] In a further embodiment, the predetermined condition is that the ratio of the smaller value to the larger value of P and Q is greater than a first threshold, or the ratio of the larger value to the smaller value is less than a second threshold.

[0008] In one embodiment, when P and Q do not meet the predetermined condition, the first party and the second party performing a private intersection operation on the P first data and the Q second data based on the detection result further includes: the first party respectively takes the modulo of the P marker codes corresponding to the P first data with respect to the numerical value L', and the obtained p modulo values correspond to p first identifiers; the second party respectively takes the modulo of the Q marker codes corresponding to the Q second data with respect to the numerical value L', and the obtained q modulo values correspond to q second identifiers; the first party and the second party perform a private intersection operation on the p first identifiers and the q second identifiers, and obtain s' common identifiers that are the same among the p first identifiers and the q second identifiers; the first party and the second party perform a private intersection operation on the a first data and the b second data respectively corresponding to the s' common identifiers, and obtain the data intersection of the N first data and the M second data.

[0009] In one embodiment, an integer between N and M.

[0010] In one embodiment, L is positively correlated with the product of N and M.

[0011] In one embodiment, a single marker code is one of the hash value of the corresponding single piece of data, the data identifier, and the hash value of the data identifier.

[0012] In one embodiment, for a single common identifier, the first party corresponds to at least one first data, and the second party corresponds to at least one second data. The first party and the second party perform a private intersection operation on the P first data and the Q second data respectively corresponding to the s common identifiers, and the obtained data intersection of the N first data and the M second data includes: the first party and the second party respectively perform a private intersection on the corresponding at least one first data and at least one second data for the s common identifiers, and obtain s subsets of the data intersection; the first party and / or the second party determine the data intersection according to the s subsets.

[0013] According to an embodiment of the second aspect, a method for private set intersection is provided, which is used to determine the data intersection of N first data held by a first party and M second data held by a second party while protecting data privacy; the method is executed by the first party and includes: respectively taking the modulo of a numerical value L with N marking codes corresponding one-to-one to the N first data to obtain n modulo values corresponding to n first identifiers; performing a private set intersection operation based on the n first identifiers and m second identifiers of the second party to obtain s common identifiers among the n first identifiers and the m second identifiers, where the m second identifiers correspond to m modulo values obtained by respectively taking the modulo of a numerical value L with M marking codes corresponding one-to-one to the M second data by the second party; and the second party performing a private set intersection operation on P first data and Q second data corresponding to the s common identifiers to obtain the data intersection of the N first data and the M second data.

[0014] In one embodiment, the step that the second party performs a private set intersection operation on P first data and Q second data corresponding to the s common identifiers to obtain the data intersection of the N first data and the M second data includes: determining the number P of first data corresponding to the s common identifiers locally; detecting with the second party through a secure comparison whether the difference between P and the number Q of second data corresponding to the s common identifiers meets a predetermined condition; and based on the detection result, performing a private set intersection operation on the P first data and the Q second data together with the second party.

[0015] In one embodiment, when P and Q do not meet the predetermined condition, the step that based on the detection result, performing a private set intersection operation on the P first data and the Q second data together with the second party further includes: respectively taking the modulo of the P marking codes corresponding to the P first data with a numerical value L' to obtain p modulo values corresponding to p first identifiers; performing a private set intersection operation on the p first identifiers and q second identifiers with the second party to obtain s' common identifiers that are the same among the p first identifiers and the q second identifiers, where the q second identifiers correspond to q modulo values obtained by respectively taking the modulo of the Q marking codes corresponding to the Q second data with a numerical value L' by the second party; and performing a private set intersection operation on a first data and b second data corresponding to the s' common identifiers respectively with the second party to obtain the data intersection of the N first data and the M second data.

[0016] In one embodiment, for a single common identifier, the first party corresponds to at least one piece of first data, and the second party corresponds to at least one piece of second data. The first party and the second party perform a private intersection operation on P pieces of first data and Q pieces of second data corresponding to the s common identifiers, and the data intersection of N pieces of first data and M pieces of second data obtained includes: for the s common identifiers, the first party and the second party respectively perform a private intersection operation on the corresponding at least one piece of first data and at least one piece of second data to obtain s subsets of the data intersection; and determining the data intersection according to the s subsets.

[0017] According to a third aspect, there is provided an apparatus for private intersection, configured to determine the data intersection of N pieces of first data held by a first party and M pieces of second data held by a second party while protecting data privacy; the apparatus is disposed in the first party and includes:

[0018] A dimensionality reduction unit, configured to respectively take the modulus of a numerical value L using N marker codes that are in one-to-one correspondence with N pieces of first data, to obtain n modulus values corresponding to n first identifiers;

[0019] A first private intersection unit, configured to perform a private intersection operation based on the n first identifiers and m second identifiers of the second party to obtain s common identifiers among the n first identifiers and the m second identifiers, where the m second identifiers correspond to m modulus values obtained by the second party by respectively taking the modulus of a numerical value L using M marker codes that are in one-to-one correspondence with M pieces of second data;

[0020] A second private intersection unit, configured to perform a private intersection operation with the second party on P pieces of first data and Q pieces of second data corresponding to the s common identifiers to obtain the data intersection of N pieces of first data and M pieces of second data.

[0021] According to a fourth aspect, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is caused to execute the method of the second aspect.

[0022] According to a fifth aspect, there is provided a computing device, including a memory and a processor, characterized in that the memory stores executable code, and when the processor executes the executable code, the method of the second aspect is implemented.

[0023] Through the method and device provided by the embodiments of this specification, during the private set intersection process of multi-party secure computing, the two data parties performing the private set intersection can respectively perform modulo operation on a predetermined value for the local data by using a marking code to distinguish the data according to the modulus value, such as grouping or partitioning and storing the data according to the modulus value. Then, the two data parties perform private set intersection on the distinguishing identifiers corresponding to the obtained modulus values in a secure manner, obtain the same distinguishing identifiers of the two parties, and further perform private set intersection on the data under the same distinguishing identifiers. This method splits the private set intersection process of the data into at least two levels. One level is to perform private set intersection on the distinguishing identifiers of the data, which is equivalent to finding the data that may be the same as that of the other party. The other level is to perform private set intersection on the data that may be the same as each other. It can be understood that on the one hand, the amount of data of the distinguishing identifiers in the first-level private set intersection process can be greatly reduced compared with the initial amount of data. On the other hand, when performing private set intersection on the data under the same distinguishing identifiers, the amount of data has been greatly reduced after filtering. In this way, the method and device provided by the embodiments of this specification can greatly reduce the data processing amount and complexity of the private set intersection and improve the efficiency of the private set intersection. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 Schematic diagram of the private set intersection implementation architecture showing the technical concept of this specification;

[0026] Figure 2 Flowchart of the private set intersection method for two-party interaction according to an embodiment;

[0027] Figure 3 Flowchart of the private set intersection method implemented on one party according to an embodiment;

[0028] Figure 4 Schematic block diagram of the private set intersection device provided on one party according to an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The technical solutions provided by this specification will be described below with reference to the drawings.

[0030] First, describe a private set intersection scenario for two data parties. It can be understood that the goal of private set intersection is for at least one data party to obtain the intersection of the business data of both parties, that is, the same or corresponding business data, without either data party disclosing the privacy of their local business data to each other. For example, in the case where the business data is bank credit data, two banks can be the two data parties for private set intersection, and the determined data intersection can be the same users. Bank users can be identified, for example, by mobile phone numbers, ID numbers, etc. Correspondingly, the data intersection can be determined through secure matching identified by mobile phone numbers, ID numbers, etc.

[0031] As an example of private set intersection, assume that A and B are two data parties, Party A holds a set of data X, and Party B holds a set of data Y. Then a conventional private set intersection method can be, for example, for B to privately compare each element y of data Y with each element x in A's data X through an oblivious pseudorandom function, and then obtain the intersection of X and Y. Assume that the number of elements in data X and Y is both n. Then the process of private set intersection is, for example: (1) A constructs n seeds k of oblivious pseudorandom functions i (i = 0, 1, 2......n - 1); (2) For each element y in Y, B executes a corresponding oblivious pseudorandom function F to obtain the set H B = {F(k i , y i ) | y i ∈ Y}; (3) For each element x in X, A executes each oblivious pseudorandom function F to obtain the set H A = {F(k i , x) | x ∈ X}; (4) A sends the set H A to B, and B determines the intersection of H A and H B , and then maps the intersection back to Y to obtain the intersection of X and Y.

[0032] Although this method is intuitive, it has a large overhead. Assume that the number of elements in sets X and Y are |X| and |Y| respectively, and the amount of data to be compared is O(|X| × |Y|). When the sizes of sets |X| and |Y| increase, the amount of data transmission in the private set intersection process grows rapidly. In conventional technologies, methods such as Elliptic Curve Diffie - Hellman encryption (ECDH), RSA, etc. can be used for private set intersection to reduce network overhead, but each piece of data needs to be encrypted twice, and the computational amount is still at least 2(|X| + |Y|) units. In the case where the order of magnitude of |X| and |Y| is large, the computational amount is still very large.

[0033] In view of this, this specification provides a technical concept for reducing the complexity of private set intersection based on multi-party secure computing, reducing the amount of data processing and communication volume, and improving the efficiency of private set intersection.

[0034] Figure 1 A specific implementation architecture of the technical concept of this specification is shown. To clarify the effect of the technical concept of this specification, specific numerical values are marked in Figure 1 . However, in practice, these numerical values vary according to the actual situation and are not limited by the numerical values marked in the figure. As Figure 1 shown, in this specific implementation architecture, a specific implementation scenario is assumed. Data party A (which can also be referred to as the first party or party A) holds 10,000 pieces of data (which can be extended to N pieces, where N is any positive integer), and data party B (which can also be referred to as the second party or party B) holds 1 million pieces of data (which can be extended to M pieces, where M is any positive integer). Using the conventional method for private set intersection requires securely comparing 10,000 pieces of data with 1 million pieces of data, and the complexity of the private set intersection is 10,000 × 1 million. In the private set intersection method that reduces network overhead, the resulting smaller amount of computation is, for example, 2 × (10,000 + 1 million) units.

[0035] The following combines Figure 2 the private set intersection process of an embodiment shown and Figure 1 the implementation architecture shown to further describe the technical concept of this specification.

[0036] Figure 2 The process shown is a flowchart of the private set intersection process completed by the interaction of two data parties. The first party and the second party involved in this figure can both be computers, devices, or servers with certain computing capabilities. For example, they can respectively correspond to Figure 1 data party A and data party B in Figure 2 . Under the condition of protecting data privacy, this process determines the data intersection of N pieces of first data held by the first party and M pieces of second data held by the second party. The following details each step with reference to

[0037] Those skilled in the art can understand that when performing private set intersection, the first party and the second party can each distinguish their local data through a consensus method, such as grouped or partitioned storage, etc. The role of the reference value L is to limit the maximum number of distinguishing identifiers of each party on the basis that the same data corresponds to the same distinguishing identifier. It is easy to understand that the possible modulus values of the relevant numerical values corresponding to the data (such as the marker code in this specification) for the reference value L are from 0 to L - 1. If these modulus values are made to correspond to the distinguishing identifiers of the data, the data can be effectively distinguished, and the same data of the two parties corresponds to the same distinguishing identifier.

[0038] Thus, the first party and the second party can first negotiate a reference value L. The two parties can randomly specify a value among the values between N and M using a secure computing method, or determine the average value of N and M, or determine a value that is positively correlated with the product of N and M as the reference value L. In an alternative embodiment, L can also be a fixed value (such as the fixed prime number 13). In another embodiment, L can also be specified by one of the parties, and this specification does not limit this in this manual. As Figure 1 shown in

[0039] After determining the value L, on the one hand, the first party, through step 201, takes the modulus of the value L respectively using N marking codes corresponding one-to-one to N pieces of first data to obtain n modulus values corresponding to n first identifiers. On the other hand, the second party, through step 202, takes the modulus of the value L respectively using M marking codes corresponding one-to-one to M pieces of second data to obtain m modulus values corresponding to m second identifiers.

[0040] Among them, step 201 and step 202 can be independently executed by the first party and the second party respectively, and the two steps can be independent of each other. For example, they can be executed in parallel, or one of the parties can pre-execute one of the steps. The marking code of the data can be the hash value of the data, the data identifier, or the hash value of the data identifier, etc. Among them, the hash value has randomness, which can ensure that the data is relatively evenly corresponding to each discrimination identifier. Therefore, when the hash value of the data or the hash value of the data identifier is used as the marking code, the hash methods of the first party and the second party are the same to ensure that the hash values of the same data are the same. And if the data identifier is a value with randomness and can ensure that the two parties use the same numerical identifier for the same data, it can also be used as the data for taking the modulus. For example, both the first party and the second party mark their registered users with unique mobile phone numbers and ID card numbers. Since mobile phone numbers and ID card numbers have a certain degree of randomness, they can be used in the process of taking the modulus of L for discrimination.

[0041] When the marking code is the hash value, the first party and the second party can pre-negotiate the hash method to ensure that the same data corresponds to the same hash value. Taking the modulus of the marking code with the value L respectively can obtain values between 0 and L - 1. A single first identifier corresponds to at least one piece of first data, and a single second identifier corresponds to at least one piece of second data. The first identifier can be regarded as a discrimination identifier for distinguishing each piece of first data (such as grouping, partitioning, etc.), and the second identifier can be regarded as a discrimination identifier for distinguishing each piece of second data (such as grouping, partitioning, etc.).

[0042] The distinguishing identifier can be a possible modulus value between 0 and L-1 obtained by taking the modulus of the data's tag code with respect to the reference value L, or it can be other string identifiers that are in one-to-one correspondence with 0 to L-1. In this way, the data of the first party and the second party each correspond to at most L distinguishing identifiers between 0 and L-1. In some embodiments, when data is distinguished by means such as grouped or partitioned storage, the distinguishing identifier can be a value between 0 and L-1. In this case, the data can be stored corresponding to the distinguishing identifier. At this time, the first identifier and the second identifier can be the distinguishing identifiers such as the groups or partitions into which the first party and the second party store the data, respectively. In some embodiments, the concept of grouping or partitioning may not be involved, and the data is simply distinguished according to the modulus result of taking the modulus with respect to L. Or rather, the data with the same modulus value obtained by taking the modulus with respect to a pre-determined value L is grouped together. At this time, the first identifier / second identifier can be the distinguishing identifier corresponding to the modulus value between 0 and L-1 obtained by taking the modulus of the data tag code of the first party / second party with respect to L.

[0043] It can be understood that since the distinguishing identifiers corresponding to the L modulus values for a single data party are not necessarily all occupied, both n and m are integers less than or equal to L. For example, Figure 1 in the specific example of, assuming that the negotiated reference value L is 100,000, since Party A has only 10,000 pieces of data, it corresponds to at most 10,000 distinguishing identifiers, while Party B holds 1,000,000 pieces of data, corresponding to at most 100,000 distinguishing identifiers.

[0044] Next, in step 203, the first party and the second party perform a private intersection operation on the n first identifiers and the m second identifiers to obtain s common identifiers among the n first identifiers and the m second identifiers. Among them, the method of private intersection can be any known or unknown private intersection method. For example, private intersection can be implemented through at least one of secret sharing, homomorphic encryption, garbled circuits, oblivious transfer, etc.

[0045] It can be understood that the intersection data is the data that exists in both the first party and the second party. Therefore, in the same data distinguishing process, in the first party and the second party, it will correspond to the same distinguishing identifier (i.e., different modulus values). Further, the data with different distinguishing identifiers must be different, and only the data with the same distinguishing identifier has the same possibility. If the distinguishing identifiers recorded by the first party are compared with those of the second party, the data corresponding to the distinguishing identifiers that do not exist in the other party can be filtered out. For example, if 10 pieces of data of the first party correspond to the distinguishing identifier "L-3", and there is no data in the second party corresponding to the distinguishing identifier "L-3" or there is no distinguishing identifier "L-3", then these 10 pieces of data must not be the intersection data of the two parties, and thus can be filtered out during the private intersection process.

[0046] Therefore, in step 203, a PSI (i.e., private set intersection) for differentiation identification can be performed. In this way, the private set intersection process of the data is temporarily transformed into the private set intersection of the first identification and the second identification. Refer to Figure 1 As shown, the private set intersection is performed through at most 10,000 (the actual quantity can be denoted as n) differentiation identifications of Party A and at most 100,000 (the actual quantity can be denoted as m) differentiation identifications of Party B. The method of private set intersection will not be elaborated here.

[0047] It is easy to understand that the computational complexity of the private set intersection of differentiation identifications is much smaller than that of directly performing private set intersection on the data, especially when the data volumes N and M of the two parties differ greatly (such as differing by 2 or more orders of magnitude). Suppose the number of identical identifications among n first identifications and m second identifications is s, then s common identifications can be obtained. Here, the s common identifications can be obtained by both parties to confirm the data that may be the same as that of the other party. The computational complexity of this private set intersection varies according to the method of private set intersection. For the purpose of comparison, in Figure 1 , the same measurement mechanism as directly performing private set intersection on 10,000 pieces of data and 1 million pieces of data in the previous text is adopted, which is at least 2×(10,000 + 100,000) units. In Figure 1 , it is assumed that the number of intersections of differentiation identifications is 5,000 (the value range of the actual quantity s is from 1 to L - 1).

[0048] Furthermore, in step 204, the first party and the second party perform private set intersection operations on the P first data and Q second data corresponding to their respective s common identifications to obtain the data intersection of the N first data and the M second data. It can be understood that both the first party and the second party can filter the local data according to the intersection of the differentiation identifications. That is, filter out the data that is necessarily not the same as the data of the other party. More specifically, filter out the data corresponding to other differentiation identifications not in the intersection of the differentiation identifications, or only obtain the data corresponding to the intersection of the differentiation identifications. In this way, the data corresponding to the differentiation identifications that do not appear or are corresponding data in the first party are filtered out in the second party, and the data corresponding to the differentiation identifications that do not appear or are corresponding data in the second party are filtered out in the first party. Therefore, the corresponding data of a single second identification that does not exist among the n first identifications is filtered out, and similarly, the corresponding data of a single first identification that does not exist among the m second identifications is also filtered out.

[0049] It can be understood that in order to ensure that some data is filtered out during the first privacy intersection, the differentiated data should not be too concentrated or too dispersed, as both situations are not conducive to filtering data through the privacy intersection of differentiated identifiers. Moreover, the filtering effect is better when the data gap between the first party and the second party is larger. Therefore, during the negotiation of L between the two parties as described above, L is balanced between a smaller quantity and a larger quantity. For example, L can take a value between N and M, that is, greater than the smaller data volume and less than the larger data volume. In one embodiment, L can take the average value of N and M. In another embodiment, L can take a number positively correlated with the product of N and M, such as (N·M). 1 / 2 etc.

[0050] Here, it can be assumed that the first party can have P pieces of first data corresponding to s common identifiers (such as Figure 1 6,000 pieces in the example), and the second party can have Q pieces of second data corresponding to s common identifiers (such as Figure 1 50,000 pieces in the example). In this way, the privacy intersection process of N pieces of first data and M pieces of second data is transformed into the privacy intersection of P pieces of first data and Q pieces of second data. In other words, the data intersection of these P pieces of first data and Q pieces of second data is the data intersection of N pieces of first data and M pieces of second data.

[0051] In one embodiment, the first party can determine P pieces of first data corresponding to s common identifiers in the local data, and the second party can determine Q pieces of second data corresponding to s common identifiers in the local data. Then, the first party and the second party jointly determine the data intersection of P pieces of first data and Q pieces of second data. The computational complexity at this time is O(P×Q), the computational amount is at least, for example, 2(P + Q) units, and the number of communication rounds is at least 1. As Figure 1 in the example, Party A and Party B can perform a privacy intersection on the filtered 6,000 pieces of data and 50,000 pieces of data to obtain the data intersection. For example Figure 1 the data intersection shown in the example is 1,000 pieces of data. Similarly, the computational amount of this privacy intersection is, for example, 2×(6,000 + 50,000) units.

[0052] In another embodiment, the first party and the second party can perform s privacy intersections for s common identifiers. Assume that the data corresponding to the s common identifiers of the first party are p1, p2... p s pieces (p1 + p2 +... + p s = P), and the data corresponding to the s common identifiers of the first party are q1, q2... q s pieces (q1 + q2 +... + q sIf (=Q), the first party and the second party can respectively perform private intersection on p1 pieces of data and q1 pieces of data, p2 pieces of data and q2 pieces of data... to obtain s subsets of the data intersection. Then, based on these s subsets, the data intersection can be obtained, which is the data intersection of P pieces of first data and Q pieces of second data. It can be understood that there may be empty sets among these s subsets. Therefore, according to a specific example, these s subsets can be filtered to remove the empty sets and then merged to obtain the data intersection. The computational complexity at this time is O(p1×q1)+O(p2×q2)+……+O(p s ×q s ), and the amount of computation is as small as 2(P + Q) for example, and the number of communication rounds is as small as s. Compared with the previous embodiment, the process of performing the private intersection operation on P pieces of first data and Q pieces of second data corresponding to s common identifiers in this embodiment has a lower complexity but an increased number of communication rounds.

[0053] In more embodiments, the first party and the second party can also use other methods to perform the private intersection operation on P pieces of first data and Q pieces of second data corresponding to s common identifiers to obtain the data intersection of N pieces of first data and M pieces of second data, which will not be elaborated here.

[0054] It should be noted that the data intersection of N pieces of first data and M pieces of second data can be known to both the first party and the second party, or can be known to one party while the other party cannot know, depending on the privacy protection requirements and / or business requirements, which are not limited here.

[0055] According to the above description, in Figure 1 the specific example shown, the amount of computation for directly performing private intersection on data is approximately 2×(10,000 + 1,000,000) = 2,020,000 units, while using the technical concept of this specification, the amount of computation is approximately 2×(10,000 + 100,000)+2×(6,000 + 50,000) = 332,000 units. In comparison, the amount of computation is reduced by one order of magnitude, which is much smaller than the amount of computation in the direct private intersection of data method. The complexity is also greatly reduced. In the above process, the comparison of the amount of computation is carried out under the same private intersection method, because the technical concept of this specification improves the architecture of the implementation process of private intersection of data, rather than involving the specific data setting details of private intersection of data.

[0056] It can be understood that the process of private intersection of the first identifier and the second identifier is to reduce the amount of data for private intersection, and the amount of data for private intersection can be reduced once or multiple times, depending on the specific business requirements. For example Figure 1As shown in FIG. 1 , after one dimensionality reduction, the privacy intersection of 6,000 pieces of data and 50,000 pieces of data is obtained. At this time, the two pieces of data differ by one order of magnitude, and the privacy intersection can be directly performed. If the number of pieces of data on the two sides is still very different after filtering, a second dimensionality reduction can be further performed, that is, step 201, step 202 and step 203 are executed again.

[0057] In view of this, in a possible design, after the first party and the second party each determine the number of first data items P and the number of second data items Q corresponding to the s common identifiers locally, they can detect whether the difference between P and Q meets the predetermined condition through security comparison, and perform a privacy intersection operation on the P first data and the Q second data based on the detection result. Specifically, if the predetermined condition is met, it can be considered that the dimensionality reduction is successful, and the privacy intersection of the P first data and the Q second data is directly performed. Otherwise, the dimensionality reduction can be continued until the predetermined condition is met. In one embodiment, the predetermined condition can be determined according to the ratio of P and Q. For example, the predetermined condition can be: the ratio of the smaller value to the larger value in P and Q is greater than a first threshold value (such as 1 / 10), or the ratio of the larger value to the smaller value is less than a second threshold value (such as 10). In another embodiment, the predetermined condition can be that the difference in order of magnitude between P and Q is less than 2. In other embodiments, the predetermined condition can also be other conditions, which will not be repeated here. In practice, when the number of data items of the first party and the second party differs greatly (for example, the difference is 2 or more orders of magnitude), the dimensionality reduction effect is particularly significant.

[0058] In the second dimension reduction process, in order to distinguish the result of the previous modulus, the value L or the data mark code can be changed. For example, in an optional implementation, the reference value of the modulus can be changed to L', and the determination method of L' is similar to that of L. In one example, At this time, the first party can take the modulus of the value L' for the P marking codes corresponding to the P pieces of first data, and the obtained p modulus values ​​correspond to p first identifiers (consistent with step 201), and the second party can take the modulus of the value L' for the Q marking codes corresponding to the Q pieces of second data, and the obtained q modulus values ​​correspond to q second identifiers (consistent with step 202). Then, the first party and the second party perform a privacy intersection operation on the p first identifiers and the q second identifiers, and obtain s' common identifiers that are the same among the p first identifiers and the q second identifiers (consistent with step 203). Assuming that the number of first data corresponding to the first party and the s' common identifiers is a, and the number of second data corresponding to the second party and the s' common identifiers is b, the first party and the second party can perform a privacy intersection operation on a pieces of first data and b pieces of second data, and the obtained data intersection is the data intersection of N pieces of first data and M pieces of second data.

[0059] exist Figure 2In the illustrated embodiment, the implementation process under the implementation framework of this specification is described based on the interaction between the first party and the second party. For any data party performing private set intersection, such as Figure 2 the first party or the second party in Figure 1 or the data party A and data party B in Figure 3 the operations they perform can be as shown in Figure 2 The first party and the second party in Figure 3 are only used to distinguish between the two data parties and have no substantial limiting effect on the technical solution. Therefore, for the sake of convenience of description, it is assumed that Figure 3 the party executing the operations shown in

[0060] Step 301: Take the modulus of the value L with respect to each of the N marking codes corresponding to the N first data items respectively, to obtain n modulus values corresponding to n first identifiers;

[0061] Step 302: Perform a private set intersection operation based on the n first identifiers and the m second identifiers of the second party, to obtain s common identifiers among the n first identifiers and the m second identifiers, where the m second identifiers correspond to the m modulus values obtained by taking the modulus of the value L with respect to each of the M marking codes corresponding to the M second data items respectively by the second party;

[0062] Step 303: Perform a private set intersection operation with the second party on the P first data items and Q second data items corresponding to the s common identifiers, to obtain the data intersection of the N first data items and the M second data items.

[0063] The operation process of the second party is opposite to that of the first party and will not be elaborated here. It can be understood that Figure 1 what is shown in Figure 2 is the implementation framework of this specification, Figure 1 and Figure 2 what is shown in Figure 3 is the private set intersection process of the interaction between two parties in an embodiment under the implementation framework of this specification. Therefore, the descriptions for Figure 1 and Figure 2 can be adapted to each other. Figure 1 and Figure 2 The embodiment shown in Figure 3 can be the architecture executed by any party during the private set intersection process. Therefore, the methods described for the corresponding data parties in

[0064] In the embodiments described above, for the private set intersection method provided in this specification, the two data parties performing private set intersection can distinguish local data according to the modulo results for a predetermined value. In this way, by performing private set intersection on the distinguishing identifiers corresponding to the moduli obtained by the two parties, data that cannot be the same between the two parties is filtered out, and only the filtered data is used for private set intersection, greatly reducing the amount of data. Thus, the private set intersection method provided in this specification reduces the computational complexity and the amount of data calculation, and improves the efficiency of private set intersection through the ingenious setting of distinguishing by modulus and the implementation framework of first performing private set intersection on the distinguishing identifiers for data dimensionality reduction and quantity reduction. Experiments show that for the case where the amounts of data of the two parties are very different, the technical concept provided in this specification has particularly remarkable effects.

[0065] According to an embodiment of another aspect, there is also provided an apparatus for private set intersection. It is used for two data parties to perform private set intersection during the multi-party secure computing process. The apparatus for private set intersection can be disposed on any one of the two data parties performing private set intersection. The two data parties performing private set intersection can also be regarded as a system for private set intersection.

[0066] Figure 4 The private set intersection apparatus 400 of an embodiment is shown. As Figure 4 shown, the apparatus 400 includes:

[0067] A dimensionality reduction unit 401, configured to respectively modulo a value L by using N marking codes corresponding one-to-one to N pieces of first data to obtain n modulus values corresponding to n first identifiers;

[0068] A first private set intersection unit 402, configured to perform a private set intersection operation based on the n first identifiers and the m second identifiers of the second party to obtain s common identifiers among the n first identifiers and the m second identifiers, where the m second identifiers correspond to the m modulus values obtained by the second party by respectively modulo a value L by using M marking codes corresponding one-to-one to M pieces of second data;

[0069] A second private set intersection unit 403, configured to perform a private set intersection operation with the second party on P pieces of first data and Q pieces of second data corresponding to the s common identifiers to obtain the data intersection of the N pieces of first data and the M pieces of second data.

[0070] It should be noted that Figure 4 the shown apparatus 400 corresponds to Figure 3 the described method, Figure 3The corresponding descriptions in the method embodiments are equally applicable to the apparatus 400 and will not be elaborated here. In addition, those skilled in the art can understand that the first private intersection unit 402 and the second private intersection unit 403 are both used to perform private intersection with the opposite data party, and the difference lies in the different data targeted. Here, considering that the results of the two private intersections may be different. For example, the result of the first private intersection of the modulus is publicly disclosed to both parties, while the result of the second private intersection may only be known to one party, two private intersection units 401 and 402 are shown. In practice, the first private intersection unit 402 and the second private intersection unit 403 may represent the same private intersection unit. In this way, the apparatus 400 includes only one private intersection unit, which is used to perform the operations of the first private intersection unit 402 and the second private intersection unit 403.

[0071] According to an embodiment of another aspect, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in connection with Figure 3 etc.

[0072] According to an embodiment of still another aspect, a computing device is further provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in connection with Figure 3 etc. is implemented.

[0073] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0074] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the technical concept of this specification. It should be understood that the above is only the specific embodiments of the technical concept of this specification and is not used to limit the protection scope of the technical concept of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions in the embodiments of this specification should be included in the protection scope of the technical concept of this specification.

Claims

1. A method for private set intersection, which is used to determine the data intersection of N first data held by a first party and M second data held by a second party while protecting data privacy; The method includes: The first party takes the modulus of the value L with each of the N marking codes corresponding to the N pieces of first data one by one, obtaining n modulus values corresponding to n first identifiers; The second party takes the modulus of the value L with each of the M marking codes corresponding to the M pieces of second data one by one, obtaining m modulus values corresponding to m second identifiers; The first party and the second party perform a private intersection operation on the n first identifiers and the m second identifiers, obtaining s common identifiers among the n first identifiers and the m second identifiers; The first party and the second party perform a private intersection operation on the P pieces of first data and the Q pieces of second data respectively corresponding to the s common identifiers, obtaining the data intersection of the N pieces of first data and the M pieces of second data.

2. The method according to claim 1, wherein, The private intersection is implemented by at least one of secret sharing, homomorphic encryption, garbled circuits, and oblivious transfer.

3. The method according to claim 1, wherein, The first party and the second party perform a private intersection operation on the P pieces of first data and the Q pieces of second data respectively corresponding to the s common identifiers, obtaining the data intersection of the N pieces of first data and the M pieces of second data includes: The first party and the second party respectively determine the number P of first data and the number Q of second data corresponding to the s common identifiers locally; The first party and the second party detect whether the difference between P and Q meets a predetermined condition through a secure comparison; The first party and the second party perform a private intersection operation on the P pieces of first data and the Q pieces of second data based on the detection result.

4. The method according to claim 3, wherein The predetermined condition is that the ratio of the smaller value to the larger value of P and Q is greater than a first threshold, or the ratio of the larger value to the smaller value is less than a second threshold.

5. The method according to claim 3, wherein In the case where P and Q do not meet the predetermined condition, the first party and the second party performing a private intersection operation on the P pieces of first data and the Q pieces of second data based on the detection result further includes: The first party takes the modulus of each of the P marking codes corresponding to the P pieces of first data with the value L' one by one, obtaining p modulus values corresponding to p first identifiers; The second party takes the modulus of each of the Q marking codes corresponding to the Q pieces of second data with the value L' one by one, obtaining q modulus values corresponding to q second identifiers; The first party and the second party perform a private intersection operation on the p first identifiers and the q second identifiers, obtaining s' common identifiers that are the same among the p first identifiers and the q second identifiers; The first party and the second party perform a private intersection operation on the a pieces of first data and the b pieces of second data respectively corresponding to the s' common identifiers, obtaining the data intersection of the N pieces of first data and the M pieces of second data.

6. The method according to claim 1, wherein L is an integer between N and M.

7. The method according to claim 1 or 6, wherein, L is positively correlated with the product of N and M.

8. The method according to claim 1, wherein, A single marking code is one of the hash value of the corresponding single piece of data, the data identifier, and the hash value of the data identifier.

9. The method according to claim 1, wherein For a single common identifier, the first party corresponds to at least one piece of first data, and the second party corresponds to at least one piece of second data. The first party and the second party perform a private intersection operation on the P pieces of first data and the Q pieces of second data respectively corresponding to the s common identifiers, obtaining the data intersection of the N pieces of first data and the M pieces of second data includes: The first party and the second party perform private intersection on at least one first data and at least one second data respectively for s common identifiers, and obtain s subsets of the data intersection; The first party and / or the second party determine the data intersection according to the s subsets.

10. A method for private set intersection, which is used to determine the data intersection of N first data held by a first party and M second data held by a second party while protecting data privacy; The method is executed by the first party and includes: Taking modulo of a numerical value L respectively with N marking codes corresponding one-to-one to N pieces of first data, and obtaining n modulo values corresponding to n first identifiers; Performing a private intersection operation based on the n first identifiers and m second identifiers of the second party to obtain s common identifiers among the n first identifiers and the m second identifiers, where the m second identifiers correspond to m modulo values obtained by taking modulo of a numerical value L respectively with M marking codes corresponding one-to-one to M pieces of second data of the second party; And performing a private intersection operation with the second party on P pieces of first data and Q pieces of second data corresponding to the s common identifiers to obtain the data intersection of N pieces of first data and M pieces of second data.

11. The method according to claim 10, wherein, The performing a private intersection operation with the second party on P pieces of first data and Q pieces of second data corresponding to the s common identifiers to obtain the data intersection of N pieces of first data and M pieces of second data includes: Determining the number P of first data corresponding to the s common identifiers locally; Detecting with the second party through a secure comparison whether the difference between the number Q of second data corresponding to the P and the s common identifiers meets a predetermined condition; Based on the detection result, performing a private intersection operation with the second party on P pieces of first data and Q pieces of second data.

12. The method according to claim 11, wherein, In the case where P and Q do not meet the predetermined condition, the performing a private intersection operation with the second party on P pieces of first data and Q pieces of second data based on the detection result further includes: Taking modulo of the P marking codes corresponding to the P pieces of first data respectively with a numerical value L', and obtaining p modulo values corresponding to p first identifiers; Performing a private intersection operation with the second party on the p first identifiers and q second identifiers to obtain s' common identifiers that are the same among the p first identifiers and the q second identifiers, where the q second identifiers correspond to q modulo values obtained by taking modulo of the Q marking codes corresponding to the Q pieces of second data of the second party respectively with the numerical value L'; Performing a private intersection operation with the second party on a pieces of first data and b pieces of second data respectively corresponding to the s' common identifiers to obtain the data intersection of the N pieces of first data and the M pieces of second data.

13. The method according to claim 10, wherein, For a single common identifier, the first party corresponds to at least one piece of first data, and the second party corresponds to at least one piece of second data. The performing a private intersection operation with the second party on P pieces of first data and Q pieces of second data corresponding to the s common identifiers to obtain the data intersection of N pieces of first data and M pieces of second data includes: Performing a private intersection operation with the second party on at least one piece of first data and at least one piece of second data respectively corresponding to the s common identifiers to obtain s subsets of the data intersection; Determining the data intersection according to the s subsets.

14. A device for private set intersection is used to determine the data intersection of N first data held by a first party and M second data held by a second party while protecting data privacy; The device is disposed in the first party and includes: A dimensionality reduction unit configured to respectively perform modulo operations on a value L using N marker codes that are in one-to-one correspondence with N pieces of first data, to obtain n modulo values corresponding to n first identifiers; A first private intersection unit configured to perform a private intersection operation based on the n first identifiers and m second identifiers of a second party, to obtain s common identifiers among the n first identifiers and the m second identifiers, where the m second identifiers correspond to m modulo values obtained by respectively performing modulo operations on a value L using M marker codes that are in one-to-one correspondence with M pieces of second data of the second party; A second private intersection unit configured to perform a private intersection operation with the second party on P pieces of first data and Q pieces of second data corresponding to the s common identifiers, to obtain a data intersection of the N pieces of first data and the M pieces of second data.

15. A computer-readable storage medium having stored thereon a computer program, which when executed on a computer, causes the computer to execute the method according to any one of claims 10-13.

16. A computing device, comprising a memory and a processor, characterized in that, Executable code is stored in the memory, and when the processor executes the executable code, the method according to any one of claims 10-13 is implemented.

Citation Information

Patent Citations

  • Privacy protection method and system based on knowledge migration under collaborative learning framework

    CN110647765A

  • Privacy intersection method and device

    CN112597524A