Multi-party sample alignment method and device for longitudinal federated learning
By adopting a combination method of random response mechanism, cuckoo hashing technology and OPRF protocol in vertical federated learning, the alignment sample set is sampled, encoded and reconstructed, which solves the shortcomings of the existing MPSI protocol in protecting the privacy of alignment samples and improving efficiency, and achieves efficient and secure data alignment.
Patent Information
- Application Number
- CN202510298357.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-17
AI Technical Summary
The existing multi-party privacy set interception (MPSI) protocol is difficult to effectively protect the privacy of aligned samples in vertical federated learning, and when the data set size is uneven, it is easy to lead to the risk of privacy leakage of weak parties, and the calculation and communication overhead is high and the efficiency is low.
The random response mechanism and cuckoo hashing technology are used to independently sample and encode the original data set, secret shares are reconstructed through the OPRF protocol, and randomly perturbed the index set to generate an aligned sample set to ensure the privacy protection and accuracy of the aligned samples.
It effectively protects the privacy of the aligned samples, reduces the risk of privacy leakage, improves the accuracy and efficiency of data alignment, and meets the requirements of differential privacy.
Smart Images

Figure CN120163265A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vertical federated learning, and in particular, to a multi-party sample alignment method and device for vertical federated learning. Background Art
[0002] In the torrent of the big data era, data has rapidly emerged due to its huge value and has become an essential key element in economic development, especially playing a crucial role in information matching services. In the scenario of vertical federated learning with multi-party collaboration, participants obtain common samples in the dataset by matching the same identifiers, establish the mapping relationship of different datasets, and ensure the smooth progress of subsequent steps of vertical federated learning. However, while the information matching service brings great convenience, it also hides serious privacy leakage risks. For example, a scholar at a university collected a large amount of user data on Facebook and shared it with an analysis company, resulting in the user information being used for political advertising. When data processors lack effective management and protection measures during the collection and sharing of user data, the privacy and security of users are at stake. Therefore, it is urgent to study how to protect user sensitive information in the information matching service for data alignment in vertical federated learning.
[0003] Existing information matching services use the private set intersection technology, abbreviated as PSI technology, to protect the user sensitive information involved in the private information matching process. PSI gives the ability to jointly calculate the intersection or the cardinality of the intersection of their respective private sets by two or more parties who do not trust each other, without revealing any other information except the intersection and the cardinality of the intersection. The privacy guarantee of the PSI scheme mainly relies on encryption technologies such as homomorphic encryption and public key encryption, and most of them have strong privacy protection effects, but the computational overhead and communication overhead are generally high. In order to reduce the computational and communication overhead, other technologies are introduced in PSI to complete the matching and transmission work more efficiently. And two-party PSI is unable to handle the data alignment work in vertical federated learning with multi-party collaboration, and a more complex multi-party private set intersection, that is, MPSI, is required to complete it.
[0004] Compared with two-party PSI protocols, the MPSI protocol brings problems such as participant collusion, more computational rounds, and diverse topological structures due to the addition of more participants. Extra security sharing, key splitting, and other links need to be considered, which brings additional security risks. Existing MPSI protocols have overcome these additional difficulties through techniques such as zero sharing, but they only have the ability to protect the privacy of non-aligned samples while ignoring the privacy protection of aligned samples. For example, based on public key cryptosystems, garbled circuits, and OT and other technologies, although privacy can be protected to a certain extent, when used for data alignment, each participant can still obtain information that some sample IDs exist in the intersection. When participants share aligned samples, it may lead to the leakage of sensitive information. Correspondingly, existing MPSI protocols are prone to higher privacy leakage risks for the weak party, that is, the participant with a smaller dataset, in the case of unbalanced dataset sizes. Different from two-party PSI, which has a lightweight solution for large datasets and can find a clever solution when only one party is large. However, existing MPSI protocols have many participants and are inefficient when the dataset is large. Even if most sets are small, they will be dragged down by the operations of a small number of large sets, making it difficult to meet the needs of practical applications. This results in some MPSI protocols facing asymmetric sets and requiring a large number of unexpected calculations and communication operations, leading to too long running times. In addition to efficiency issues, existing MPSI schemes applying differential privacy are difficult to ensure the accuracy of aligned samples while ensuring privacy protection. The distortion of aligned samples may lead to a decline in the model training effect. Traditional differential privacy PSI is inefficient in scenarios with large dataset participants on the one hand; on the other hand, perturbations are all carried out on the large data side, and the security of the data of the small data side cannot be guaranteed. Therefore, there is an urgent need for a multi-party sample alignment method for vertical federated learning to complete the data alignment work in multi-party collaborative vertical federated learning. Summary of the Invention
[0005] In view of this, it is necessary to provide a multi-party sample alignment method and device for vertical federated learning to effectively solve the technical problems in data alignment in multi-party collaborative vertical federated learning.
[0006] The present invention provides a multi-party sample alignment method for vertical federated learning, including the following steps:
[0007] Step S1, the randomly selected first weak party and other participants independently sample the original datasets they hold using the random response mechanism with different probabilities, the first weak party exchanges PRF seeds with other participants, and uses the seeds and the sampled subsets to generate the secret shares of each participant;
[0008] Step S2: The second weakest party randomly selected encodes the sample IDs into a variant Bloom filter using the cuckoo hashing technique, and other participating parties construct their own hash tables using the same hash function as the second weakest party;
[0009] Step S3: The third weakest party randomly selected executes the OPRF protocol using the Bloom filter and the hash table as inputs to reconstruct the secret shares of each participating party to obtain the reconstructed shares;
[0010] Step S4: The third weakest party extracts the reconstructed shares to obtain the intersection affected by the sampling selection and the index sets of other participating parties, randomly perturbs the index sets, and generates the aligned sample sets of each participating party;
[0011] Step S5: Each participating party extracts the aligned samples from its own dataset according to the corresponding aligned sample sets to complete the sample alignment.
[0012] Preferably, in step S1, the first weakest party randomly selected and other participating parties independently sample the original datasets they hold using the random response mechanism with different probabilities, specifically:
[0013] Sample the original dataset held by the first weakest party with probability p. Each element in the original dataset of the first weakest party is left with probability p and removed with probability 1 - p to obtain the sampled subset; sample the original datasets of other participating parties except the first weakest party with probability p0 to obtain the corresponding subsets.
[0014] Preferably, in step S1, the first weakest party and other participating parties exchange PRF seeds, and use the seeds and the sampled subsets to generate the secret shares of each participating party, specifically:
[0015] Each participating party randomly selects the seed of the pseudo-random function F to obtain the PRF seed as the key. The first weakest party and other participating parties exchange keys with each other and set the relevant keys, and both the key exchange and the key setting are completed in the channel constructed by the security algorithm;
[0016] For each element in the sampled subset, each participating party calculates the corresponding secret share using the improved zero-sharing algorithm.
[0017] Preferably, step S1 further includes:
[0018] Using the index of the element in the subset as the local identifier, each participating party shares the secret slice with the first weakest party as the key to calculate the hash value, and stores the local mapping relationship between the index and the element.
[0019] Preferably, step S2 is specifically:
[0020] The second weakest party encodes the subset using cuckoo hashing with multiple hash functions; other participating parties construct hash tables using the same hash functions as the second weakest party; wherein each element is repeatedly stored at multiple positions in multiple hash tables.
[0021] Preferably, step S3 is specifically as follows:
[0022] The third weakest party uses the OPRF protocol, with the Bloom filter and the hash tables of the participating parties whose shares are to be reconstructed as inputs, to obtain the set of positions of the samples of the participating parties whose shares are to be reconstructed in the third weakest party;
[0023] For each position in the set of positions, the third weakest party uses the OPRF protocol, with the value at the corresponding position in the hash table as the input, to obtain the position share of the sample of the participating party whose share is to be reconstructed at the corresponding position in the hash table;
[0024] The third weakest party uses the PRF seed and the position share to generate the reconstructed share of the participating party whose share is to be reconstructed.
[0025] Preferably, step S4 is specifically as follows:
[0026] The third weakest party checks the reconstructed shares, extracts the intersection affected by sampling and the corresponding set of indices of other participating parties;
[0027] The third weakest party randomly perturbs the set of indices and adds the set of indices to the perturbation set with a set perturbation probability;
[0028] The third weakest party randomly shuffles all the perturbation sets to generate the aligned sample sets of each participating party and synchronously sends them to the corresponding participating parties.
[0029] Preferably, step S5 is specifically as follows:
[0030] Other participating parties except the third weakest party extract aligned samples from their own datasets according to the received aligned sample sets to complete sample alignment.
[0031] The present invention also provides a multi-party sample alignment device for vertical federated learning, including a memory and a processor, where a computer program is stored on the memory, and when the computer program is executed by the processor, the multi-party sample alignment method for vertical federated learning is implemented.
[0032] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention combines differential privacy technology and random response strategy, and uses random response technology to sample and perturb the input and output of OPRF, meeting the requirements of differential privacy and ensuring the security of all aligned and unaligned data during the data alignment process. Thus, while protecting the privacy information of aligned samples, the accuracy of sample alignment is guaranteed. High accuracy and efficiency of sample alignment are achieved through efficient symmetric key primitives, simplifying the sample matching process, improving efficiency, and ensuring that the perturbation of the matching result does not interfere with the high precision of sample alignment, enhancing the performance in terms of both efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the illustrative embodiments and descriptions thereof are used to explain the present invention without unduly limiting the present invention. In the drawings:
[0034] Figure 1 It is a flowchart of an embodiment of a multi-party sample alignment method for vertical federated learning provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The following will specifically describe the preferred embodiments of the present invention in conjunction with the drawings, where the drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.
[0036] Embodiment 1
[0037] The purpose of this embodiment is to design a multi-party sample alignment differential privacy protection method for vertical federated learning. In a federated environment with poor symmetry, this method uses differential privacy technology to protect the security of sensitive samples of multiple weak data parties during the sample alignment stage, reduce the risk of aligned data leakage, and achieve high accuracy and efficiency of sample alignment through efficient symmetric key primitives. The risks of the so-called data leakage mainly refer to the risks of sample ID leakage, intersection privacy leakage, embedded representation leakage, etc. This method draws on differential privacy technology and random response strategies to perturb the multi-party sample matching results, thereby protecting the privacy information of the aligned samples while ensuring the accuracy of sample alignment. The random response technology is used to sample and perturb the input and output of the OPRF, meeting the requirements of differential privacy and ensuring the security of all aligned and unaligned data during the data alignment process. The above-mentioned matching result refers to the matching result in the ciphertext state calculated by the weak party. All data includes aligned data and unaligned data. This method integrates OT, OPRF, and zero sharing based on the Pseudorandom Function (PRF). On the premise of ensuring compatibility with multi-party scenarios, it realizes efficient primitives and algorithm processes, simplifies the sample matching process, improves efficiency, and ensures that the perturbation of the matching result does not interfere with the high precision of sample alignment, enhancing performance in terms of both efficiency and accuracy. Zero sharing refers to a technical solution for privacy protection. Through specific encoding and encryption technologies, participating parties cannot directly obtain the specific data content of other parties when sharing data, thereby protecting privacy.
[0038] The sample alignment method provided in this embodiment is applicable to the sample alignment work of vertical federated learning, especially having a good effect in the case of a large gap in set sizes, and can be separately deployed on the strong party and weak party servers participating in vertical federated learning.
[0039] Please refer to Figure 1 , a multi-party sample alignment method for vertical federated learning in this embodiment specifically includes the following steps:
[0040] Step S1: The randomly selected first weak party and other participating parties independently sample their respective original data sets using the random response mechanism with different probabilities. The first weak party exchanges the PRF seeds with other participating parties and uses the seeds and the sampled subsets to generate the secret shares of each participating party;
[0041] Step S2: The randomly selected second weak party encodes the sample IDs into a variant Bloom filter using the cuckoo hashing technique, and other participating parties construct their own hash tables using the same hash function as the second weak party;
[0042] Step S3: The randomly selected third weak party uses the Bloom filter and the hash table as input to execute the OPRF protocol, reconstructs the secret shares of each participant, and obtains the reconstructed shares;
[0043] Step S4: the third weak party extracts the reconstruction share to obtain the intersection affected by the sampling selection and the index set of other participants, randomly perturbs the index set, and generates an alignment sample set of each participant;
[0044] Step S5: Each participant extracts alignment samples from its own data set according to the corresponding alignment sample set to complete sample alignment.
[0045] The technical idea of this embodiment is to use differential privacy technology to protect the privacy of aligned samples, take into account the security of aligned samples by perturbing the grading results, and allow different participants to obtain different perturbed versions of aligned samples in combination with scenarios of unbalanced data sets. Finally, zero sharing, PRF-based OT technology, cuckoo hashing and other technologies are used to provide secure and reliable communication and alignment for VFL sample alignment.
[0046] This method is divided into five major steps: sampling and preprocessing, building a hash table, share reconstruction, extracting and perturbing aligned samples, and sample alignment.
[0047] Sampling and preprocessing: The data sets of the participants are sampled to reduce the size of the data sets of the participants, and are perturbed using random response technology to ensure the privacy protection of non-aligned samples. At the same time, the participants exchange PRF seeds with other participants and use the seeds and the sampled data sets to generate their own shares to generate each participant's share.
[0048] Construct a hash table: randomly select a weak party as the second weak party, use cuckoo hashing technology to encode the sample ID into a variant of the Bloom filter, and other participants use the same hash function to construct their own hash tables to provide input data for subsequent share reconstruction and sample alignment, thereby improving the efficiency of sample matching.
[0049] Share reconstruction: The third weak party uses the OPRF protocol to reconstruct the shares of other participants with the Bloom filter and hash table as input. Due to the nature of OPRF, if the shares of other participants are the same as the sample ID of the third weak party, the shares will be implicitly transmitted to the third weak party through the OPRF protocol, thereby ensuring the privacy of the reconstructed shares. The third weak party is a participant who is selected first and is different from the second weak party.
[0050] Extract and perturb aligned samples: The third weakest party extracts the reconstructed shares, obtains the intersection affected by the sampling selection and the corresponding index sets of other participating parties, then randomly perturbs the index sets, and generates aligned sample sets for different participating parties to ensure differential privacy protection of the aligned samples.
[0051] Sample alignment module: Each participating party extracts aligned samples from its own dataset according to the received index set, thus completing sample alignment.
[0052] This embodiment extends the advantages of random perturbation, applies the random perturbation technique to both dataset sampling and intersection calculation results simultaneously. Different from traditional perturbation only for sampled data, it can provide security protection without significantly affecting the output results regardless of whether the data is aligned. A pseudo-random function is adopted to reconstruct the shares of other participating parties to ensure the privacy of the shares. OPRF allows the sender to calculate the random function of the function value without revealing the input, thus realizing the secure reconstruction of the shares. The characteristics of Bloom filters and hash tables are fully utilized to store sample IDs. Combining with the encoding scheme of cuckoo hashing can not only ensure the accuracy of Bloom filters but also improve the efficiency of sample matching.
[0053] The present invention proposes a local differential privacy algorithm applicable to asymmetric multi-party set intersection, which provides the ability for participating parties to have different degrees of privacy protection. At the same time, the intersection of the results of all perturbed versions is extracted from the local private set according to the perturbed index set, effectively protecting the data security of the weak parties in the data alignment stage of vertical federated learning while solving the problem of low intersection accuracy of the MPSI protocol incorporating differential privacy. The multi-party private set intersection module of the present invention is based on an improved zero-sharing algorithm, introduces the OPRF protocol to reduce the interference of perturbation, and improves the overall communication and calculation efficiency by locally establishing a private mapping between set elements and special identifiers and extracting special identifiers to form a set when sharing results.
[0054] Specifically, taking three medical institutions as an example for illustration, the three medical institutions respectively have the medical record data, test data, and imaging data of patients. They want to jointly train a disease prediction model, but due to the requirements of data privacy protection, they cannot directly share the data. Therefore, all data must be kept secure and confidential during the data vertical federated learning alignment stage. This method is used to handle the data alignment problem in vertical federated learning among medical institutions.
[0055] Among them, Hospital A has tens of millions of patient data, while Hospital B and Hospital C each have a relatively small number of tens of thousands of patient data. Since the number of data entries in Hospital A is much larger than those in B and C, A is called the strong party in the interaction, and B and C are called the weak parties. Both the strong party and the weak parties have to undertake the work of data sampling, perturbation, and alignment. In particular, a randomly selected weak party needs to additionally undertake the perturbation of the intersection, that is, the result after alignment.
[0056] Specifically, it includes the following steps:
[0057] Step 11: Sampling and zero sharing. In this step, the randomly selected first weak party hospital and the strong party independently sample the original sample sets they hold with probabilities p and p0 respectively using the random response mechanism. Only the elements in the sampled sets will be used for zero sharing and subsequent operations. Then, each pair of participating parties exchanges the pseudo-random function keys, and at the same time each party generates secret shares locally. The specific steps are as follows:
[0058] Step 111: Each participating party P i Randomly selects the seed of the pseudo-random function F where j = i + 1, i + 2, …, n, u = 1, 2, …, n, and exchanges keys with the other n - 1 participating parties and sets the relevant keys to ensure that all key setting and exchange processes are completed in the channel constructed by the security algorithm.
[0059] Step 112: Randomly select a small data party, the first weak party, to execute. The probability that each element in the dataset held by the first weak party is sampled is p, that is, each element remains in the new sampled set with probability p and is removed from the set with probability 1 - p. The first weak party uses the sampled set as the input set for the subsequent steps of the protocol, and let X j be the subset sampled by the first weak party with probability p. The other participating parties sample the elements in the set with probability p0 to obtain a random subset n of X and set
[0060] Step 113: For each element in the set each participating party calculates its corresponding secret share using the improved zero-sharing algorithm and uses the index of the element in the local private set of the participating party as the local identifier μ to share with other participating parties as the hash value calculated by the key and stores the local mapping relationship between the index and the element. represents the u-th secret slice of the key sent by the first weak party to other participating parties.
[0061] Step 12. Confuse the set. In this step, select a random small data party different from the above-mentioned one performing perturbation, and use the cuckoo hash with 5 hash functions to encode the private set. Other participating parties use the same 5 hash functions to construct a hash table, where each element is stored repeatedly in 5 positions.
[0062] Step 13. Share reconstruction. In this step, the selected second-weakest party, Hospital B, and other hospitals execute the OPRF protocol using the Bloom filter and the hash table as inputs respectively. If the shares of the set elements in other parties are the same as the shares held by the second-weakest party, then these shares will be inadvertently transferred to the second-weakest party.
[0063] Steps for Hospital A to reconstruct shares:
[0064] Step 131. Hospital B uses the OPRF protocol with the Bloom filter B1 and the hash table T of Hospital A A as inputs to obtain the set S of positions of the sample IDs of Hospital A in B1 A .
[0065] Step 132. For each position i in S A , Hospital B uses the OPRF protocol with T A [i] as the input to obtain the share y A [i] of the sample ID of Hospital A in T A [i].
[0066] Step 133. Hospital B uses the PRF and y A [i] to generate the share ρ A [i] of Hospital A.
[0067] Steps for Hospital C to reconstruct shares:
[0068] Step 134. Hospital B uses the OPRF protocol with the Bloom filter B1 and the hash table T of Hospital C C as inputs to obtain the set S of positions of the sample IDs of Hospital C in B1 C .
[0069] Step 135. For each position i in S C , Hospital B uses the OPRF protocol with T C [i] as the input to obtain the share y C [i] of the sample ID of Hospital C in T C [i].
[0070] Step 136. Hospital B uses the PRF and y C [i] to generate the share ρ C [i] of Hospital C.
[0071] Steps for reconstructing the share of Hospital B:
[0072] Step 137: Hospital B uses the OPR protocol, with the Bloom filter B1 and its own hash table T B as inputs, to obtain the set S of positions of the sample IDs of Hospital B in B1 B .
[0073] Step 138: For each position i in S B , Hospital B uses the OPRF protocol, with T B [i] as the input, to obtain the share y B [i] of the sample ID of Hospital B in T B [i].
[0074] Step 139: Hospital B uses the PRF and y B [i] to generate its own share ρ B [i].
[0075] Step 14: Extract and perturb the aligned samples. In this step, the selected third-weakest party, Hospital B, extracts the reconstructed shares to obtain the intersection affected by the sampling operation and other corresponding index sets, adds each index outside the index set to the result index set with a certain perturbation probability to generate a perturbed index set to increase the possibility of successful matching, and shuffles all versions of the index sets and sends them to the participants in the vertical federated learning for preparation for the subsequent steps.
[0076] Step 141: The third-weakest party checks the reconstructed shares and extracts the intersection affected by the sampling operation and the index sets of other hospitals corresponding thereto. The share reconstruction refers to the traditional OPRF protocol.
[0077] Step 142: Hospital B randomly perturbs the index set and adds it to the set with probability q, where q is the perturbation parameter.
[0078] Step 143: Hospital B randomly shuffles all versions of the perturbed intersections and synchronously sends them to Hospital A and Hospital C.
[0079] Step 15: Sample alignment. Except for the selected weak parties, other hospitals extract the aligned samples from their own datasets according to the received index sets, obtain the aligned sample set and complete the sample alignment to support the subsequent collaborative training.
[0080] Example 2
[0081] This example provides a multi-party sample alignment device for vertical federated learning, including a memory and a processor. A computer program is stored on the memory, and when the computer program is executed by the processor, it implements the multi-party sample alignment method for vertical federated learning described in Example 1.
[0082] The multi-party sample alignment device for vertical federated learning provided in this embodiment is used to implement the multi-party sample alignment method for vertical federated learning. Therefore, the technical effects possessed by the multi-party sample alignment method for vertical federated learning are also possessed by the multi-party sample alignment device for vertical federated learning, which will not be elaborated here.
[0083] As mentioned above, only the specific preferred embodiments of the present invention are described, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the present invention.
Claims
1. A multi-party sample alignment method for vertical federated learning, characterized in that: The following steps are involved: Step S1: The randomly selected first weak party and other participating parties independently sample the original data sets they hold using a random response mechanism with different probabilities. The first weak party exchanges PRF seeds with other participating parties, and uses the seeds and the sampled subsets to generate the secret shares of each participating party. Step S2: The randomly selected second weak party uses the cuckoo hashing technique to encode the sample ID into the variant Bloom filter, and other participants use the same hash function as the second weak party to build their own hash tables; Step S3: The randomly selected third weak party uses the Bloom filter and the hash table as input to execute the OPRF protocol, reconstructs the secret shares of each participant, and obtains the reconstructed shares; Step S4: the third weak party extracts the reconstruction share to obtain the intersection affected by the sampling selection and the index set of other participants, randomly perturbs the index set, and generates an alignment sample set of each participant; Step S5: Each participant extracts alignment samples from its own data set according to the corresponding alignment sample set to complete sample alignment.
2. The multi-party sample alignment method for vertical federated learning according to claim 1, characterized in that: The first weak party and other participants randomly selected in step S1 use a random response mechanism with different probabilities to independently sample the original data sets they hold, specifically: Sampling the original data set held by the first weak party with probability p, retaining each element in the original data set of the first weak party with probability p and removing each element with probability 1-p, to obtain a sampled subset; The original data sets of the other parties except the first weak party are sampled with probability p0 to obtain corresponding subsets.
3. The multi-party sample alignment method for vertical federated learning according to claim 1, characterized in that: In step S1, the first weak party exchanges PRF seeds with other participants, and uses the seeds and the sampled subsets to generate secret shares of each participant, specifically: Each participant randomly selects a seed of a pseudo-random function F and obtains the PRF seed as a key. The first weak party exchanges keys with other participants and sets related keys. The key exchange and key setting are completed in a channel constructed by a security algorithm. For each element in the sampled subset, each participant uses the improved zero-sharing algorithm to calculate the corresponding secret share.
4. The multi-party sample alignment method for vertical federated learning according to claim 1, characterized in that: The step S1 further comprises: Using the index of the element in the subset as the local identifier, each participant shares the secret slice with the first weak party as the key to calculate the hash value, and stores the local mapping relationship between the index and the element.
5. The multi-party sample alignment method for vertical federated learning according to claim 1, characterized in that: The step S2 is specifically as follows: The second weak party encodes the subset using cuckoo hashing of multiple hash functions; other participants construct hash tables using the same hash function as the second weak party; wherein each element is repeatedly stored in multiple locations of the multiple hash tables.
6. The multi-party sample alignment method for vertical federated learning according to claim 1, characterized in that: The step S3 is specifically as follows: The third weak party uses the OPRF protocol, taking the Bloom filter and the hash table of the to-be-reconstructed share participant as input, to obtain a position set of the to-be-reconstructed share participant sample in the third weak party; For each position in the position set, the third weak party uses the OPRF protocol, takes the corresponding position value of the hash table as input, and obtains the position share of the to-be-reconstructed share participant sample in the corresponding position of the hash table; The third weak party generates a reconstruction share of a participant to be reconstructed using the PRF seed and the position share.
7. The multi-party sample alignment method for vertical federated learning according to claim 1, characterized in that: The step S4 is specifically: The third weak party checks the reconstruction share and extracts the intersection affected by the sampling and the corresponding index set of other participants; The third weak party randomly perturbs the index set, and adds the index set to the perturbation set with a set perturbation probability; The third weak party randomly shuffles all disturbance sets, generates an alignment sample set for each participant, and sends it to the corresponding participant synchronously.
8. The multi-party sample alignment method for vertical federated learning according to claim 1, characterized in that: The step S5 is specifically: The other participants except the third weak party extract alignment samples from their own data sets according to the received alignment sample set to complete sample alignment.
9. A multi-party sample alignment device for vertical federated learning, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, a multi-party sample alignment method for vertical federated learning as described in any one of claims 1 to 8 is implemented.