De-identification privacy set intersection method and system
By converting the identification data into random strings and de-identified set interception in a trusted execution environment, the problem of identity leakage of some users in the intersection in the existing technology is solved, and a safe and efficient privacy set interception is achieved, which is suitable for big data analysis and machine learning.
Patent Information
- Application Number
- CN202510345944.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
The existing privacy collection interception technology has leaked user identity identification information in the intersection part, resulting in insufficient privacy protection.
By converting the identification data into a random string sequence, deterministic random algorithm and cuckoo hash algorithm are used, combined with the inadvertent transmission protocol, deidentified set intersecting is performed in a trusted execution environment to ensure that the intersection result does not contain real identity information.
It realizes the safe and efficient calculation of intersection results without leaking user identity information, enhances privacy protection, and is suitable for scenarios such as big data analysis, machine learning, and vertical federated learning.
Smart Images

Figure CN120277712A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data privacy protection, and more specifically, to a method and system for de-identifying private set intersection. Background Art
[0002] Private Set Intersection (PSI) is a cryptographic technique that allows two or more parties to calculate their intersection without revealing the non-intersecting parts of their respective sets. The PSI technique is widely used in scenarios such as sample alignment and data matching. For example, in vertical federated learning, it is used to achieve collaborative computing between data holders without sharing the original data.
[0003] In existing PSI methods, the IDs (identity identifiers) of the non-intersecting parts can be protected from being leaked, but the IDs of the intersecting parts are still known to both parties. This means that after obtaining the intersection, both Party A and Party B can know which IDs are common, resulting in the privacy leakage of the IDs in the intersecting part. In vertical federated learning, if two data holders obtain the intersection through the PSI technique, both parties will know the user IDs and their related information in the intersection, which may violate the requirements of privacy protection.
[0004] For example, assume that in vertical federated learning, Party A holds user IDs and their gender information, and Party B holds user IDs and their address information. Through traditional PSI techniques, both parties can obtain the user IDs in the intersecting part (such as Bob) and know Bob's gender and address. However, this result will cause the ID of Bob and its related information to be known to both parties, posing a risk of privacy leakage. For example, Party A knows Bob's address, and Party B knows Bob's gender, which may be used by both parties to infer more sensitive information, thus threatening user privacy.
[0005] Therefore, existing PSI techniques have deficiencies in protecting the privacy of IDs in the intersecting part and need to be further improved. Summary of the Invention
[0006] The purpose of the present invention is to provide a method and system for de-identifying private set intersection, which solves the problem of leakage of sensitive identifier data, especially user identity identifier information, in existing private set intersection technologies.
[0007] To achieve the above purpose, the present invention provides a method for de-identifying private set intersection, including the following steps:
[0008] Convert the identification data of the first party and the second party into a random string sequence;
[0009] Using a deterministic random algorithm, calculate using the random string sequences of the first participant and the second participant as inputs respectively, and use the calculation results as address indexes;
[0010] The first participant and the second participant respectively reorder the local data according to the address indexes, construct a de-identified set, and send the constructed de-identified set to the trusted execution environment;
[0011] In the trusted execution environment, perform an intersection operation on the de-identified sets of the first participant and the second participant to obtain the de-identified data intersection result;
[0012] Wherein, the deterministic random algorithm refers to an algorithm that outputs the same result when the inputs are the same, but the output result is unpredictable.
[0013] In some embodiments, the step of converting the identification data of the first participant and the second participant into random string sequences further includes the following steps:
[0014] Perform binary representation on the identification data of the first participant and the second participant, and generate two random strings for each bit of the identification data of the first participant for selection;
[0015] The first participant and the second participant execute an oblivious transfer protocol, and the second participant sequentially selects corresponding random strings for each bit of its own identification data;
[0016] The first participant and the second participant generate random string sequences based on the randomly selected strings of each of them.
[0017] In some embodiments, the first participant and the second participant execute an oblivious transfer protocol, and the second participant sequentially selects corresponding random strings for each bit of its own identification data, which further includes:
[0018] The first participant sends a first random sequence data set to the second participant based on the oblivious transfer protocol, and the first random sequence data set is composed of two random strings corresponding to each bit of the identification data of the first participant in sequence;
[0019] The second participant sequentially selects for each bit of its own identification data the random string at the corresponding position in the first random sequence data set.
[0020] In some embodiments, the oblivious transfer protocol is an oblivious transfer protocol based on symmetric cryptography;
[0021] The step that the first participant and the second participant execute an oblivious transfer protocol, and the second participant sequentially selects corresponding random strings for each bit of its own identification data further includes the following steps:
[0022] The first participating party and the second participating party share a symmetric key;
[0023] The first participating party encrypts each random string in the first set of random sequence data using the symmetric key, and sends the encrypted first set of random sequence data to the second participating party;
[0024] The second participating party selects the corresponding random string in the first set of random sequence data according to each bit of its own identification data, and decrypts it using the symmetric key to obtain the required random string.
[0025] In some embodiments, the oblivious transfer protocol is an oblivious transfer protocol based on public key cryptography;
[0026] When the first participating party and the second participating party execute the oblivious transfer protocol, and the second participating party successively selects the corresponding random string for each bit of its own identification data, the steps further include the following steps:
[0027] The first participating party and the second participating party respectively generate their own public keys and private keys;
[0028] The first participating party encrypts each random string in the first set of random sequence data using the public key of the second participating party, and sends the encrypted first set of random sequence data to the second participating party;
[0029] The second participating party selects the corresponding random string in the first set of random sequence data according to each bit of its own identification data, and decrypts it using its own private key to obtain the required random string.
[0030] In some embodiments, the first participating party and the second participating party use the same deterministic random algorithm;
[0031] The deterministic random algorithm is a hash function, and the de-identified set is a hash table.
[0032] In some embodiments, the first participating party and the second participating party use the hash function to calculate the address index with the random string sequence as the input;
[0033] The first participating party and the second participating party respectively insert their corresponding local data into their own hash tables based on the cuckoo hash algorithm according to the address index, and send the hash tables to the trusted execution environment.
[0034] In some embodiments, when the first participating party and the second participating party respectively reorder the local data according to the address index to construct the de-identified set, the steps further include:
[0035] Initialize the hash table;
[0036] Based on a hash function, with a random string sequence as the input, calculate to obtain an address index;
[0037] Based on the address index, obtain the candidate positions of the local data in the hash table and insert them;
[0038] If the candidate position is already occupied, kick out the data at the occupied position to other candidate positions, and then insert the local data into the candidate position;
[0039] Repeat the above insertion process for all the local data of the first party and the second party until all the data are inserted into the hash table.
[0040] In some embodiments, the identification data at least includes identity identification data.
[0041] In some embodiments, the random string is generated by a pseudo-random number generator;
[0042] The length of the random string includes 128 bits or 256 bits.
[0043] To achieve the above object, the present invention provides a de-identified private set intersection system, which at least includes a first party, a second party, and a trusted execution environment:
[0044] The first party converts its own identification data into a random string sequence, uses a deterministic random algorithm, calculates with its own random string sequence as the input, uses the calculation result as its own address index, reorders the local data according to the address index, constructs a de-identified set, and sends the constructed de-identified set to the trusted execution environment;
[0045] The second party converts its own identification data into a random string sequence, uses a deterministic random algorithm, calculates with its own random string sequence as the input, uses the calculation result as its own address index, reorders the local data according to the address index, constructs a de-identified set, and sends the constructed de-identified set to the trusted execution environment;
[0046] The trusted execution environment is respectively communicatively connected to the first party and the second party, and performs an intersection operation on the de-identified sets of the first party and the second party to obtain a de-identified data intersection result;
[0047] Wherein, the deterministic random algorithm refers to an algorithm that outputs the same result when the input is the same, but the output result is unpredictable.
[0048] In some embodiments, the first participating party represents its own identification data in binary and generates two random strings for each bit of its own identification data for selection;
[0049] The second participating party represents its own identification data in binary, executes an oblivious transfer protocol with the first participating party, and sequentially selects corresponding random strings for each bit of its own identification data;
[0050] The first participating party and the second participating party generate corresponding random string sequences based on the randomly selected strings of each of them.
[0051] In some embodiments, the first participating party sends a first set of random sequence data to the second participating party based on the oblivious transfer protocol, and the first set of random sequence data is composed of two random strings corresponding to each bit of the first participating party's identification data in sequence;
[0052] The second participating party sequentially selects, for each bit of its own identification data, the random string at the corresponding position in the first set of random sequence data.
[0053] In some embodiments, the oblivious transfer protocol is an oblivious transfer protocol based on symmetric cryptography;
[0054] The first participating party and the second participating party share a symmetric key;
[0055] The first participating party encrypts each random string in the first set of random sequence data using the symmetric key and sends the encrypted first set of random sequence data to the second participating party;
[0056] The second participating party selects the random string corresponding to the first set of random sequence data according to each bit of its own identification data and decrypts it using the symmetric key to obtain the required random string.
[0057] In some embodiments, the oblivious transfer protocol is an oblivious transfer protocol based on public key cryptography;
[0058] The first participating party and the second participating party respectively generate their own public keys and private keys;
[0059] The first participating party encrypts each random string in the first set of random sequence data using the public key of the second participating party and sends the encrypted first set of random sequence data to the second participating party;
[0060] The second participating party selects the random string corresponding to the first set of random sequence data according to each bit of its own identification data and decrypts it using its own private key to obtain the required random string.
[0061] In some embodiments, the first participating party and the second participating party use the same deterministic random algorithm;
[0062] The deterministic random algorithm is a hash function, and the de-identified set is a hash table.
[0063] In some embodiments, the first participating party and the second participating party calculate an address index based on a hash function with a random string sequence as the input;
[0064] The first participating party and the second participating party respectively insert their corresponding local data into their own hash tables based on the address index using the cuckoo hash algorithm, and send them to the trusted execution environment.
[0065] In some embodiments, the first participating party and the second participating party insert local data into their own hash tables in the following manner:
[0066] Initialize the hash table;
[0067] Based on the hash function, calculate an address index with a random string sequence as the input;
[0068] Based on the address index, obtain the candidate position of the local data in the hash table and insert it;
[0069] If the candidate position is already occupied, kick out the data at the occupied position to other candidate positions, and then insert the local data into the candidate position;
[0070] Repeat the above insertion process for all local data of the first participating party and the second participating party until all data are inserted into the hash table.
[0071] In some embodiments, the identification data at least includes identity identification data.
[0072] In some embodiments, the random string is generated by a pseudorandom number generator;
[0073] The length of the random string includes 128 bits or 256 bits.
[0074] The present invention provides a method and system for de-identified private set intersection. Through a deterministic random algorithm, combined with the cuckoo hash algorithm and the oblivious transfer protocol, it effectively removes identification data such as user IDs, and realizes efficient and secure de-identified private set intersection calculation on the premise of ensuring that the private data of each party is not leaked. Brief Description of the Drawings
[0075] The above and other features, properties, and advantages of the present invention will become more apparent from the following description in conjunction with the accompanying drawings and embodiments, in which like reference numerals always denote like features, where:
[0076] Figure 1 Disclosed is a flowchart of a method for intersection of de-identified private sets according to an embodiment of the present invention;
[0077] Figure 2 Disclosed is a schematic diagram of the target for intersection of private sets according to an embodiment of the present invention;
[0078] Figure 3 Disclosed is a schematic diagram of performing an intersection operation using the cuckoo hash algorithm and a trusted environment according to an embodiment of the present invention;
[0079] Figure 4 Disclosed is a block diagram of the principle of a system for intersection of de-identified private sets according to an embodiment of the present invention. Detailed Embodiments
[0080] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the invention and are not used to limit the invention.
[0081] The present invention discloses a method and system for intersection of de-identified private sets, providing a more secure and privacy-protected set intersection technology, which is applicable to various privacy computing scenarios involving sensitive data, especially in the fields of big data analysis, machine learning, and joint computing, and has broad application prospects.
[0082] Existing PSI technologies are prone to privacy leakage when processing the IDs in the intersection part. For example, assume that Table 1 and Table 2 respectively record user IDs and their attribute information (e.g., gender and address).
[0083] Table 1
[0084] User ID User Gender Bob Male Alice Female Duck Male
[0085] Table 2
[0086]
[0087]
[0088] Representing the above Table 1 and Table 2 as sets A and B respectively for intersection operation, private set intersection is to find the intersection of two sets without revealing data privacy. The intersection result of sets A and B above is shown in Table 3 below:
[0089] Table 3
[0090] User ID User Gender User Address Bob Male Shanghai
[0091] As shown in Table 3, during the existing private set intersection process, although the intersection part (i.e., Bob) can be obtained, the holder of set B cannot know other data in set A except the intersection (such as Alice, Duck). Similarly, set A cannot learn other data in set B (such as Mary, Fucy). However, in the existing PSI technology, there are still potential privacy leakage problems. Especially during the intersection calculation process, both parties may indirectly expose some user ID information.
[0092] For example, the holder of set B can know the ID (Bob) in the intersection and its associated gender information, while the holder of set A may know Bob's address. Such information leakage may pose a threat to the privacy of data holders.
[0093] Aiming at the potential privacy leakage problem that may exist when the existing private set intersection (PSI) technology processes the user identifiers (IDs) in the intersection part, the present invention proposes a de-identification private set intersection method.
[0094] Figure 1 Disclosed is a flowchart of a de-identification private set intersection method according to an embodiment of the present invention. As Figure 1 shown, a de-identification private set intersection method proposed by the present invention includes at least the following steps:
[0095] Step S1, converting the identification data of the first participant and the second participant into a random string sequence;
[0096] Step S2, using a deterministic random algorithm to perform calculations respectively with the random string sequences of the first participant and the second participant as inputs, and using the calculation results as address indexes, where the deterministic random algorithm refers to an algorithm that outputs the same result when the inputs are the same, but the output result is unpredictable;
[0097] Step S3, the first participant and the second participant respectively reorder the local data according to the address indexes, construct de-identified sets, and send the constructed de-identified sets to the trusted execution environment;
[0098] Step S4, in the trusted execution environment, perform an intersection operation on the de-identified sets of the first participant and the second participant to obtain a de-identified data intersection result;
[0099] The proposed de-identification privacy set intersection method enables both parties to determine whether there is an intersection between two sets and to know the category or quantity of the intersection without knowing the specific identification information of the intersection, ensuring the privacy of the intersection operation.
[0100] In this embodiment, the identification data is ID (identity identification) data. The de-identification process mainly refers to removing the user ID to avoid exposing personal identity information.
[0101] For example, in machine learning applications, when tasks such as sample alignment are involved, the participating parties only need to know the category or set of the samples without knowing the specific IDs of the samples provided, thus effectively protecting the privacy of users.
[0102] In fact, this method is not limited to removing user IDs and can also be used to remove other identification data, such as usernames, phone numbers, email addresses, etc.
[0103] In the application scenario, any form of identification information that can be used to identify the identity of the user or data holder can be used as the object of de-identification. By removing this identification data, this method can further enhance the privacy protection and ensure that the sensitive information of the data holder will not be leaked during the privacy set intersection process.
[0104] Specifically, the proposed de-identification privacy set intersection method is based on an innovative design of hash functions, cuckoo hashing algorithms, and oblivious transfer (OT) technology. Without revealing the user IDs of both parties, it effectively realizes the intersection solution of the privacy set, thereby protecting the privacy information of the data holder.
[0105] These steps will be described in detail below. It should be understood that within the scope of the present invention, all the above technical features of the present invention and the technical features specifically described below (such as in the embodiments) can be combined with each other and are interrelated, thus constituting a preferred technical solution.
[0106] Step S1: Convert the identification data of the first participating party and the second participating party into a random string sequence;
[0107] In this embodiment, the identification data is ID data.
[0108] More specifically, the random string is generated by a pseudo-random number generator;
[0109] The length of the random string includes 128 bits or 256 bits.
[0110] In this embodiment, step S1 further includes the following steps:
[0111] Step S11: Represent the identity data of the first participant and the second participant in binary, and generate two random strings for each bit of the first participant's identity data for selection.
[0112] Step S12: The first participant and the second participant execute the oblivious transfer protocol, and the second participant sequentially selects the corresponding random string for each bit of its own identity data.
[0113] Step S13: The first participant and the second participant generate a random string sequence based on the randomly selected strings of each of them.
[0114] For step S11, in this embodiment, the identity (ID) data of the first participant is x, and the identity (ID) data of the second participant is y.
[0115] The identity (ID) data x of the first participant can be represented as a binary sequence [x1, x2,..., xl]. For each bit xi in the sequence, two random strings k_{i0} and k_{i1} are generated, corresponding to the two possible values of xi respectively:
[0116] When xi = 0, select k_{i0};
[0117] When xi = 1, select k_{i1}.
[0118] For example, for x1, if the corresponding random string is k_{10}, it means x1 = 0; if it is k_{11}, it means x1 = 1. Similarly, x2 also has two corresponding random strings k_{20} and k_{21}, and so on.
[0119] For step S12, the first participant and the second participant execute the oblivious transfer protocol, and the second participant sequentially selects the corresponding random string for each bit of its own identity data.
[0120] If one party can know the random string selected by the other party, it is possible to reverse the other party's identity (ID) through the known address index, thus revealing privacy. To prevent this from happening, this step uses the oblivious transfer (OT) protocol to prevent one party from obtaining the selection of the other party.
[0121] Oblivious transfer is a cryptographic protocol that mainly addresses the issues of privacy protection and selective information transfer. It allows one party (the sender) to send multiple messages to another party (the receiver), but the receiver can only receive their own selections without knowing the other messages, and the sender does not know the receiver's selections. This mechanism has important applications in secure multi-party computing, privacy protection, and distributed systems, which can ensure the confidentiality of information and the privacy of the receiver while preventing any information leakage.
[0122] In this embodiment, the first participating party sends a first set of random sequence data to the second participating party based on the oblivious transfer protocol, and the first set of random sequence data is composed of two random strings corresponding to each bit of the first participating party's identification data combined in sequence;
[0123] The second participating party sequentially makes selections for each bit of its own identification data according to the random strings at the corresponding positions in the first set of random sequence data.
[0124] For example, in this embodiment, an OT protocol is executed between the first participating party and the second participating party. The first participating party provides random strings k_{i0} and k_{i1} as inputs, where i takes values from 1, 2,..., l. The second participating party can finally obtain k_{i0} or k_{i1} corresponding to yi based on the binary sequence [y1, y2,..., yl] of its identity identification.
[0125] For the first participating party, the second participating party does not know other information except the random string corresponding to yi. Similarly, for the second participating party, the first participating party cannot obtain the specific value of yi either.
[0126] Under the OT (oblivious transfer) protocol, the privacy of both parties is guaranteed, and at the same time its operation purpose is achieved: the second participating party can make selections from the random strings provided by the first participating party according to its own yi value.
[0127] There are multiple implementation methods for the OT protocol. In this embodiment, two implementation methods are provided, namely the OT protocol based on public-key cryptography and the OT protocol based on symmetric cryptography.
[0128] The OT protocol based on public-key cryptography uses asymmetric encryption technologies (such as RSA, ECC, etc.) to protect the privacy of data. In this protocol, the sender and the receiver use a pair of public keys and private keys for encryption and decryption operations.
[0129] The OT protocol based on public-key cryptography further includes the following steps:
[0130] The first participating party and the second participating party respectively generate their own public keys and private keys;
[0131] The first party encrypts each random string in the first set of random sequence data (such as random strings \(k_{i0}\) and \(k_{i1}\)) using the public key of the second party, and sends the encrypted first set of random sequence data to the second party, and only the second party can decrypt it.
[0132] The second party selects one encrypted random string corresponding to the first set of random sequence data according to each bit of its own identity data (such as each bit of the identity identifier \(y\)), and decrypts it using its own private key to obtain the required random string.
[0133] Since the second party can only select one of the random strings and other information is invisible to the second party, the second party decrypts using its own private key to obtain the required random string.
[0134] The first party cannot know the choice of the second party because the choice of the second party is not disclosed to the first party; at the same time, the second party cannot know the content of other random strings to be selected and will not disclose its own choice information.
[0135] Different from the public-key cryptography OT protocol, the OT protocol of symmetric cryptography uses a shared symmetric key for encryption and decryption operations. The OT protocol based on symmetric cryptography is usually faster, but requires a pre-shared key between the participating parties or a trusted third party to manage the key.
[0136] The OT protocol based on symmetric cryptography further includes the following steps:
[0137] The first party and the second party share a symmetric key (possibly pre-exchanged through a secure channel or provided by a trusted party), and this symmetric key is used for encrypting and decrypting information;
[0138] The first party encrypts each random string in the first set of random sequence data (such as random strings \(k_{i0}\) and \(k_{i1}\)) using the symmetric key, and sends the encrypted first set of random sequence data to the second party;
[0139] The second party selects the random string corresponding to the first set of random sequence data according to each bit of its own identity data, and decrypts it using the symmetric key to obtain the required random string.
[0140] The first party cannot know the choice of the second party because the choice of the second party is based on the symmetric key and this key is private, and the second party will not know the content of other random strings to be selected.
[0141] In this process, due to the use of the OT protocol, the first party does not know the choice of the second party, and the second party cannot learn the choice of the first party, thus realizing the privacy protection of the data of both parties.
[0142] In step S13, the first party and the second party generate a random string sequence based on the randomly selected strings of each of them.
[0143] More specifically, the first party selects a random string according to the identification data value of each bit to generate a random string sequence of the identification data of the first party.
[0144] The second party selects a random string according to the identification data value of each bit to generate a random string sequence of the identification data of the second party.
[0145] In this embodiment, the first party denotes the sequence of k_{i0} or k_{i1} corresponding to xi as K1, which is used as the random string sequence of the first party, and the second party denotes the sequence of k_{i0} or k_{i1} corresponding to yi as K2, which is used as the random string sequence of the second party.
[0146] These random string sequences will be used as the input of the subsequent deterministic random algorithm to obtain the corresponding address indexes.
[0147] Step S2: Use the deterministic random algorithm to perform calculations with the random string sequences of the first party and the second party as inputs respectively, and use the calculation results as address indexes.
[0148] Step S3: The first party and the second party respectively reorder the local data according to the address indexes, construct a de-identified set, and send the constructed de-identified set to the trusted execution environment.
[0149] The deterministic random algorithm refers to an algorithm that outputs the same result when the input is the same, but the output result is unpredictable.
[0150] In this embodiment, the deterministic random algorithm is a hash function.
[0151] In step S2, the random string sequence converted from the identification data is used as the input, and the address index is obtained through hash function calculation, thereby hiding the true identity of the user and ensuring the privacy and security of data processing.
[0152] In this embodiment, the de-identified set is a hash table. In step S1, the user's real ID is replaced with a random string sequence K. In step S2, with the random string sequence K as the input, an address index is calculated through a hash function. The address index is used as the insertion position of the local data in the hash table in step S3, thus hiding the user's real identity and achieving de-identification processing. At the same time, since the random string sequence K is generated through the OT protocol, neither the first party nor the second party can directly obtain the K value of the other party, further protecting the privacy of both parties.
[0153] In this embodiment, the first party and the second party use the same deterministic random algorithm.
[0154] The first party and the second party use the same hash function and calculate the obtained address index with the random string sequence as the input;
[0155] The first party and the second party respectively insert their corresponding local data into their respective hash tables based on the address index using the cuckoo hashing algorithm and send the hash tables to the trusted execution environment.
[0156] In the cuckoo hashing algorithm, the hash table consists of multiple hash buckets, and each hash bucket is used to store a specific element or multiple conflicting elements.
[0157] To ensure that the hash bucket positions of the same ID in the two-party hash table are consistent so that the same IDs can be accurately matched during the intersection calculation, the first party and the second party must use the same hash function. In this way, the same ID data is mapped to the same position in the hash tables of both parties, ensuring that the IDs can be correctly matched when solving the intersection and avoiding problems caused by inconsistent mapping positions.
[0158] For example, for the data of a certain ID, the method of calculating the address index using the hash function corresponding to the cuckoo hashing algorithm is as follows:
[0159] First, based on the random string sequence K corresponding to the ID, calculate Hash(K) = h;
[0160] Then, determine the address index by taking the remainder of h divided by the hash bucket length, that is, h mod the hash bucket length, which will be used as its specific candidate position (hash bucket) in the hash table.
[0161] This calculation method ensures that the same ID must be mapped to the same position under the same Hash function. Conversely, if the positions are the same, the IDs must also be the same. This consistency is the key to ensuring that the first party and the second party can accurately match the same IDs during the intersection calculation.
[0162] In other embodiments, other deterministic random algorithms that meet the above conditions can also be used, such as SHA series hash functions, AES encryption algorithms, random permutation algorithms, etc. to calculate the address index.
[0163] As long as the above algorithms have the characteristics of determinism (the output is consistent when the input is the same) and pseudo-randomness (the output is unpredictable), they can be applied to the de-identification process of the present invention in the privacy protection scenario.
[0164] In this embodiment, the used cuckoo hash algorithm is an efficient implementation of a hash table, mainly used to solve the problem of collisions and achieve efficient element insertion and lookup.
[0165] In a traditional hash table, when multiple keys are mapped to the same location, it will cause collisions, affecting the lookup and insertion efficiency.
[0166] The cuckoo hash algorithm implements a "kicking out" algorithm by using two or more hash functions. When inserting an element, if the target location is already occupied, it will be "kicked out" to another location, so as to keep the load factor of the hash table low and ensure that the average lookup and insertion time is close to the constant time complexity. This method effectively reduces the impact of collisions and improves the performance of the hash table.
[0167] More specifically, the cuckoo hash algorithm is used to insert local data into the hash table, and the process includes the following steps:
[0168] Initialize the hash table;
[0169] Based on the hash function, using a random string sequence as the input, calculate the obtained address index;
[0170] Based on the address index, obtain the candidate position (hash bucket) of the local data in the hash table and insert it;
[0171] If the candidate position is already occupied, then kick out the data at the occupied position to other candidate positions (calculated by another hash function), and then insert the local data into the candidate position;
[0172] Repeat the above insertion process for all local data of the first party and the second party until all data are inserted into the hash table.
[0173] The cuckoo hash algorithm uses two hash functions to map data to different hash buckets and solves hash collisions through the "kicking out" operation to ensure that elements are finally inserted into the appropriate positions.
[0174] In this step, the cuckoo hash algorithm not only optimizes the storage and retrieval efficiency of data, but also ensures the accuracy and security of data processing.
[0175] After the cuckoo hash algorithm finishes inserting data into the hash buckets, the first participant and the second participant send the hash tables they constructed to the trusted execution environment for subsequent processing.
[0176] Step S4: In the trusted execution environment, perform an intersection operation on the de-identified sets of the first participant and the second participant to obtain the de-identified data intersection result.
[0177] The trusted execution environment is an isolated computing environment that can ensure data security and privacy protection, ensuring that data will not be leaked during processing.
[0178] The trusted execution environment ensures that during the intersection calculation, the identity data of the first participant and the second participant are not leaked. Even if the hash buckets are intercepted or tampered with during transmission, the trusted execution environment will encrypt and protect the data to prevent data leakage.
[0179] In the trusted execution environment, perform an intersection operation on the hash tables of the first participant and the second participant to obtain the de-identified data intersection result.
[0180] In this embodiment, in the trusted execution environment, perform an intersection solution on the hash tables of the first participant and the second participant, and finally obtain the de-identified intersection result, that is, the intersection part of the identity data. During this process, the calculation of the intersection only depends on the content of the hash buckets and does not disclose the identity information of any participant.
[0181] Finally, the trusted execution environment outputs the de-identified intersection result. This intersection result does not contain the real identity data of any user, only the relevant intersection identifiers or categories, ensuring the effective protection of the privacy of the participants.
[0182] It should be specifically noted that the intersection calculation performed in the trusted execution environment is "secure computing", which means that neither party can peek at the input data of the other party and cannot obtain the intermediate calculation results. The trusted execution environment ensures the integrity of the calculation process, and no other party can obtain sensitive data by accessing the computing resources.
[0183] If the intersection operation in step S4 is not performed in the trusted execution environment but unilaterally on the first participant or the second participant, the following risks and problems exist:
[0184] The operating party may be able to access information other than the intersection calculation result, and thus there is a risk of privacy leakage in obtaining the undisclosed part of the other party's data;
[0185] The transparency and fairness of the calculation cannot be guaranteed. One party may manipulate the calculation result by modifying the content in the hash bucket for a certain purpose, leading to the untrustworthiness of the calculation result.
[0186] A malicious third party may tamper with data during the communication process. For example, during the transmission of the hash bucket, a malicious party may intercept and modify the data, thus affecting the calculation result.
[0187] Next, an example is used to illustrate the de-identification private set intersection method proposed by the present invention.
[0188] Figure 2 Discloses a schematic diagram of the private set intersection target according to an embodiment of the present invention, as Figure 2 shown. Two different institutions (such as the tax bureau and the education bureau) perform private data intersection. The tax bureau has ID card numbers and salary data, while the education bureau has ID card numbers and educational attainment data. The private data intersection target is to find the salary and educational attainment information of the same ID card number based on the ID card number, while ensuring that the ID card number itself is not disclosed.
[0189] The tax bureau's data set is used as the first participating party, the education bureau's data set is used as the second participating party, and the identifying data is the ID card number. The identifying data (ID card number) of the data sets of the tax bureau and the education bureau will be converted into binary form.
[0190] Both parties execute the oblivious transfer protocol (OT). For example, the tax bureau sends its own set of random strings to the education bureau, and the education bureau selects the corresponding random strings according to each bit of its ID card number data.
[0191] After obtaining the random strings, the tax bureau and the education bureau generate corresponding random string sequences based on the randomly selected strings by themselves, and calculate the address index through a hash function.
[0192] Figure 3 Discloses a schematic diagram of performing the intersection operation according to the cuckoo hash algorithm and the trusted environment according to an embodiment of the present invention, as Figure 3 shown. The tax bureau and the education bureau insert the local data into the corresponding hash buckets in the hash table according to the cuckoo hash algorithm. The hash buckets are sorted according to the sequence corresponding to the address index, rather than according to the user ID. Each hash bucket position is associated with a data item, but the storage of the data is based on the correspondence between the address index and a specific hash bucket position, rather than according to the ID card number.
[0193] In this way, it can be ensured that the data corresponding to the same ID card number (such as ID 123) (such as salary 20 and educational attainment master) is stored in the same hash bucket.
[0194] For example, data with a salary of 20 and a master's degree are both placed in bucket at position 2. Through the hash function and the cuckoo hash algorithm, the user's real ID (such as 123) is hidden, and a digital label is used to ensure that the ID number is not leaked.
[0195] Finally, in the hash table, the data corresponding to position 2 is the user with ID 123, which can accurately associate the salary and education data. Although the participating parties cannot directly know each other's IDs, they can perform data intersection solving in the trusted execution environment based on these hidden positions, that is, the data intersection corresponding to ID 123 at position 2, so as to obtain the salary and education information (salary is 20 and education is master's degree) of the same ID number.
[0196] Through the above embodiments, it is shown that the proposed de-identification privacy set intersection method of the present invention uses the hash function, the cuckoo hash algorithm and the oblivious transfer protocol to ensure the privacy of data, and at the same time enables two parties to perform intersection operations on data based on the ID number to obtain the de-identified intersection result, which can effectively avoid the leakage of user identity information.
[0197] The de-identification privacy set intersection method proposed by the present invention has broader applicability and flexibility during the de-identification process and can meet various privacy computing requirements involving identification information.
[0198] The first participating party and the second participating party not only have user IDs, but may also have other identification information, such as user names, phone numbers or email addresses, etc. This method can still perform de-identification processing on different identification data according to the above steps to protect different types of sensitive information.
[0199] For example, the first participating party may have identification information such as user ID, user name, phone number, etc., while the second participating party may have information such as user ID, email, phone number, etc. In one requirement, taking the phone number as the identification information, the phone number can be binaryized, and a random string is generated for each bit, and then the intersection is calculated according to the same OT protocol, cuckoo hash algorithm and trusted execution environment, and the de-identified intersection result can be obtained without leaking any identification information.
[0200] Therefore, whether it is in the intersection of data sets based on a single identifier or in complex scenarios involving multiple identifiers, this method can ensure that no user privacy information is leaked during the intersection calculation. This method has strong privacy protection and can be widely applied to various data processing scenarios that need to protect user privacy, such as big data analysis, machine learning (especially vertical federated learning), joint computing and other fields.
[0201] Figure 4Disclosed is a principle block diagram of a de-identified private set intersection system according to an embodiment of the present invention. As Figure 4 shown, the present invention also proposes a de-identified private set intersection system, which at least includes a first participant 401, a second participant 402, and a trusted execution environment 403:
[0202] The first participant 401 and the second participant 402 can be the same hardware device, both including data storage, computing processing, encryption modules, and network interfaces, for storing and processing data, generating random strings, and performing private set intersection operations.
[0203] The trusted execution environment 403 is an independent isolated computing platform, usually including hardware that supports secure computing, such as TPM (Trusted Platform Module), HSM (Hardware Security Module), and Intel SGX (Intel Software Guard Extensions), to ensure the secure execution of computing tasks under the premise of privacy protection.
[0204] The first participant 401 converts its own identification data into a sequence of random strings, uses a deterministic random algorithm, takes its own sequence of random strings as input for calculation, uses the calculation result as its own address index, reorders the local data according to the address index, constructs a de-identified set, and sends the constructed de-identified set to the trusted execution environment 403;
[0205] The second participant 402 converts its own identification data into a sequence of random strings, uses a deterministic random algorithm, takes its own sequence of random strings as input for calculation, uses the calculation result as its own address index, reorders the local data according to the address index, constructs a de-identified set, and sends the constructed de-identified set to the trusted execution environment 403;
[0206] The trusted execution environment 403 is communicatively connected to the first participant 401 and the second participant 402 respectively, and performs an intersection operation on the de-identified sets of the first participant 401 and the second participant 402, so as to obtain a de-identified data intersection result;
[0207] Among them, the deterministic random algorithm refers to an algorithm that outputs the same result when the input is the same, but the output result is unpredictable.
[0208] In some embodiments, the first participant 401 represents its own identification data in binary, and generates two random strings for each bit of its own identification data for selection;
[0209] The second participating party 402 performs binary representation on its own identification data, executes an oblivious transfer protocol with the first participating party 401, and sequentially selects corresponding random strings for each bit of its own identification data;
[0210] The first participating party 401 and the second participating party 402 generate corresponding random string sequences based on the random strings they each select.
[0211] In some embodiments, the first participating party 401 sends a first random sequence data set to the second participating party 402 based on the oblivious transfer protocol, and the first random sequence data set is composed of two random strings corresponding to each bit of the identification data of the first participating party 401 combined in sequence;
[0212] The second participating party 402 sequentially selects, for each bit of its own identification data, the random string at the corresponding position in the first random sequence data set.
[0213] In some embodiments, the oblivious transfer protocol is an oblivious transfer protocol based on public key cryptography;
[0214] The first participating party 401 and the second participating party 402 respectively generate their own public keys and private keys;
[0215] The first participating party 401 encrypts each random string of the first random sequence data set using the public key of the second participating party 402, and sends the encrypted first random sequence data set to the second participating party 402;
[0216] The second participating party 402 selects the random string corresponding to the first random sequence data set according to each bit of its own identification data, and decrypts it using its own private key to obtain the required random string.
[0217] In some embodiments, the oblivious transfer protocol is an oblivious transfer protocol based on symmetric cryptography;
[0218] The first participating party 401 and the second participating party 402 share a symmetric key;
[0219] The first participating party 401 encrypts each random string of the first random sequence data set using the symmetric key, and sends the encrypted first random sequence data set to the second participating party;
[0220] The second participating party 402 selects the random string corresponding to the first random sequence data set according to each bit of its own identification data, and decrypts it using the symmetric key to obtain the required random string.
[0221] In some embodiments, the first participant 401 and the second participant 402 use the same deterministic random algorithm;
[0222] The deterministic random algorithm is a hash function, and the de-identified set is a hash table.
[0223] In some embodiments, the first participant 401 and the second participant 402 calculate an address index based on a hash function with a random string sequence as the input;
[0224] The first participant 401 and the second participant 402 respectively insert their corresponding local data into their own hash tables based on the address index using the cuckoo hash algorithm, and send the hash tables to the trusted execution environment 403.
[0225] In some embodiments, the first participant 401 and the second participant 402 insert local data into their own hash tables in the following manner:
[0226] Initialize the hash table;
[0227] Based on the hash function, calculate an address index with a random string sequence as the input;
[0228] Based on the address index, obtain the candidate position of the local data in the hash table and insert it;
[0229] If the candidate position is already occupied, kick out the data at the occupied position to other candidate positions, and then insert the local data into the candidate position;
[0230] Repeat the above insertion process for all the local data of the first participant 401 and the second participant 402 until all the data are inserted into the hash table.
[0231] In some embodiments, the identification data at least includes identity identification data.
[0232] In some embodiments, the random string is generated by a pseudo-random number generator;
[0233] The length of the random string includes 128 bits or 256 bits.
[0234] The specific implementation details of the de-identified private set intersection system of the present invention correspond to the foregoing de-identified private set intersection method, so the specific details are not repeated here.
[0235] The present invention provides a method and system for intersection calculation of de-identified private sets. Through a deterministic random algorithm, combined with the cuckoo hash algorithm and the oblivious transfer protocol, it effectively removes identification data such as user IDs, and realizes efficient and secure calculation of the intersection of private sets on the premise of ensuring that the private data of all parties is not leaked. At the same time, through the guarantee of a trusted execution environment, the security and data integrity of the calculation process are ensured, avoiding data tampering and information leakage, and having significant privacy protection and technical application value.
[0236] The present invention provides a method and system for intersection calculation of de-identified private sets. Through hash functions, the cuckoo hash algorithm, and the oblivious transfer protocol for de-identification processing, it realizes a double improvement in privacy protection and calculation efficiency, and specifically has the following beneficial effects:
[0237] 1) Enhanced privacy protection: Through the use of hash functions and oblivious transfer technology for de-identification processing, it ensures that both parties perform intersection calculation without disclosing their own IDs, effectively avoiding the leakage of IDs in the intersection part, thereby significantly improving the privacy security of data holders;
[0238] 2) Efficient calculation: Utilize the cuckoo hash algorithm to construct an efficient hash table, enabling fast ID matching even in the case of large-scale data sets, significantly improving the intersection efficiency and reducing the calculation overhead;
[0239] 3) Wide applicability: This solution is applicable to various vertical federated learning scenarios, has good generality and scalability, and can meet the requirements of different application scenarios.
[0240] Although the above methods are illustrated and described as a series of actions for simplicity of explanation, it should be understood and appreciated that these methods are not limited by the order of the actions, because according to one or more embodiments, some actions may occur in a different order and / or concurrently with other actions that are illustrated and described herein or that are not illustrated and described herein but are understandable to those skilled in the art.
[0241] As shown in this application and the claims, unless the context clearly indicates an exception, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list, and the method or device may also include other steps or elements.
[0242] Those skilled in the art will appreciate that information, signals, and data can be represented using any of a variety of different technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips described throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.
[0243] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and the design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.
[0244] The various illustrative logical modules and circuits described in connection with the embodiments disclosed herein can be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0245] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read from, and write to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0246] The above embodiments are provided for those skilled in the art to implement or use the present invention. Those skilled in the art can make various modifications or changes to the above embodiments without departing from the inventive concept of the present invention. Therefore, the protection scope of the present invention is not limited by the above embodiments, but should be the maximum scope that conforms to the innovative features mentioned in the claims.
Claims
1. A method for intersection of privacy sets after de-identification, characterized in that, Including the following steps: Convert the identification data of the first participant and the second participant into a random string sequence; Using a deterministic random algorithm, calculate respectively with the random string sequences of the first participant and the second participant as inputs, and use the calculation results as address indexes; The first participant and the second participant respectively reorder the local data according to the address indexes, construct a de-identified set, and send the constructed de-identified set to the trusted execution environment; In the trusted execution environment, perform an intersection operation on the de-identified sets of the first participant and the second participant to obtain the de-identified data intersection result; Wherein, the deterministic random algorithm refers to an algorithm that outputs the same result when the input is the same, but the output result is unpredictable.
2. The method for intersection of de-identified privacy sets according to claim 1, wherein The step of converting the identification data of the first participant and the second participant into a random string sequence further includes the following steps: Perform binary representation on the identification data of the first participant and the second participant, and generate two random strings for each bit of the identification data of the first participant for selection; The first participant and the second participant execute an oblivious transfer protocol, and the second participant sequentially selects corresponding random strings for each bit of its own identification data; The first participant and the second participant generate a random string sequence based on the randomly selected strings of each of them.
3. The de-identification privacy set intersection method according to claim 2, characterized in that, The first participant and the second participant execute an oblivious transfer protocol, and the second participant sequentially selects corresponding random strings for each bit of its own identification data, which further includes: The first participant sends a first random sequence data set to the second participant based on the oblivious transfer protocol, and the first random sequence data set is composed of two random strings corresponding to each bit of the identification data of the first participant in sequence; The second participant sequentially selects the random strings at the corresponding positions in the first random sequence data set for each bit of its own identification data.
4. The de-identification privacy set intersection method according to claim 3, wherein The oblivious transfer protocol is an oblivious transfer protocol based on public key cryptography; The first participant and the second participant execute an oblivious transfer protocol, and the second participant sequentially selects corresponding random strings for each bit of its own identification data, which further includes the following steps: The first participant and the second participant respectively generate their own public keys and private keys; The first participant encrypts each random string of the first random sequence data set with the public key of the second participant, and sends the encrypted first random sequence data set to the second participant; The second participant selects the random string corresponding to the first random sequence data set according to each bit of its own identification data, and decrypts it with its own private key to obtain the required random string.
5. The method for intersection of de-identified privacy sets according to claim 3, wherein The oblivious transfer protocol is an oblivious transfer protocol based on symmetric cryptography; The first participant and the second participant execute an oblivious transfer protocol, and the second participant sequentially selects corresponding random strings for each bit of its own identification data, which further includes the following steps: The first participant and the second participant share a symmetric key; The first participating party encrypts each random string in the first set of random sequence data using a symmetric key, and sends the encrypted first set of random sequence data to the second participating party; The second participating party selects the corresponding random string in the first set of random sequence data according to each bit of its own identification data, and decrypts it using the symmetric key to obtain the required random string.
6. The de-identification privacy set intersection method according to claim 1, wherein The first participating party and the second participating party use the same deterministic random algorithm; The deterministic random algorithm is a hash function, and the de-identified set is a hash table.
7. The method for intersection of de-identified privacy sets according to claim 6, wherein The first participating party and the second participating party use the hash function, with the random string sequence as the input, to calculate and obtain the address index; The first participating party and the second participating party respectively insert their respective local data into their respective hash tables based on the cuckoo hash algorithm according to the address index, and send the hash tables to the trusted execution environment.
8. The method for intersection of de-identified privacy sets according to claim 7, wherein The steps for the first participating party and the second participating party to reorder the local data according to the address index and construct the de-identified set further include: Initialize the hash table; Based on the hash function, with the random string sequence as the input, calculate and obtain the address index; Based on the address index, obtain the candidate position of the local data in the hash table and insert it; If the candidate position is already occupied, kick out the data at the occupied position to other candidate positions, and then insert the local data into the candidate position; Repeat the above insertion process for all the local data of the first participating party and the second participating party until all the data are inserted into the hash table.
9. A de-identified private set intersection system, characterized in that, At least includes the first participating party, the second participating party and the trusted execution environment: The first participating party converts its own identification data into a random string sequence, uses the deterministic random algorithm, calculates with its own random string sequence as the input, uses the calculation result as its own address index, reorders the local data according to the address index, constructs the de-identified set, and sends the constructed de-identified set to the trusted execution environment; The second participating party converts its own identification data into a random string sequence, uses the deterministic random algorithm, calculates with its own random string sequence as the input, uses the calculation result as its own address index, reorders the local data according to the address index, constructs the de-identified set, and sends the constructed de-identified set to the trusted execution environment; The trusted execution environment is respectively communicatively connected to the first participating party and the second participating party, and performs an intersection operation on the de-identified sets of the first participating party and the second participating party to obtain the de-identified data intersection result; Among them, the deterministic random algorithm refers to an algorithm that outputs the same result when the input is the same, but the output result is unpredictable.
10. The de-identified private set intersection system according to claim 9, wherein The first participating party represents its own identification data in binary, and generates two random strings for each bit of its own identification data for selection; The second participating party represents its own identification data in binary, executes the oblivious transfer protocol with the first participating party, and sequentially selects the corresponding random string for each bit of its own identification data; The first participating party and the second participating party generate corresponding random string sequences based on the randomly selected strings of each party.
11. The de-identified private set intersection system according to claim 10, wherein The first participating party sends a first set of random sequence data to the second participating party based on the oblivious transfer protocol. The first set of random sequence data is composed of two randomly selected strings corresponding to each bit of the first participating party's identification data, combined in sequence. The second participating party sequentially selects, for each bit of its own identification data, the randomly selected string at the corresponding position in the first set of random sequence data.
12. The de-identified private set intersection system according to claim 11, wherein The oblivious transfer protocol is an oblivious transfer protocol based on public key cryptography. The first participating party and the second participating party respectively generate their own public keys and private keys. The first participating party encrypts each randomly selected string in the first set of random sequence data using the public key of the second participating party, and sends the encrypted first set of random sequence data to the second participating party. The second participating party selects the randomly selected string corresponding to the first set of random sequence data according to each bit of its own identification data, and decrypts it using its own private key to obtain the required randomly selected string.
13. The de-identified private set intersection system according to claim 11, characterized in that, The oblivious transfer protocol is an oblivious transfer protocol based on symmetric cryptography. The first participating party and the second participating party share a symmetric key. The first participating party encrypts each randomly selected string in the first set of random sequence data using the symmetric key, and sends the encrypted first set of random sequence data to the second participating party. The second participating party selects the randomly selected string corresponding to the first set of random sequence data according to each bit of its own identification data, and decrypts it using the symmetric key to obtain the required randomly selected string.
14. The de-identified private set intersection system according to claim 9, characterized in that, The first participating party and the second participating party use the same deterministic random algorithm. The deterministic random algorithm is a hash function, and the de-identified set is a hash table.
15. The de-identified private set intersection system according to claim 14, wherein The first participating party and the second participating party calculate the address index based on the hash function, with the random string sequence as the input. The first participating party and the second participating party respectively insert their corresponding local data into their own hash tables based on the address index using the cuckoo hash algorithm, and send them to the trusted execution environment.
16. The de-identified private set intersection system according to claim 15, wherein The first participating party and the second participating party insert the local data into their own hash tables in the following way: Initialize the hash table. Calculate the address index based on the hash function, with the random string sequence as the input. Obtain the candidate position of the local data in the hash table based on the address index and insert it. If the candidate position is already occupied, kick out the data at the occupied position to other candidate positions, and then insert the local data into the candidate position. Repeat the above insertion process for all local data of the first participating party and the second participating party until all data are inserted into the hash table.