Model training method, system, device and electronic equipment based on federated learning

By obtaining the sample intersection identification list and privacy set intersection technology, combined with the dynamic pairing of the task controller, the problem of low sample alignment efficiency in vertical federated learning is solved, and efficient large-scale and real-time federated learning training of data is achieved.

CN118396083BActive Publication Date: 2025-09-12BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410702816.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-31
Publication Date
2025-09-12
Estimated Expiration
2044-05-31

AI Technical Summary

Technical Problem

In existing technologies, vertical federated learning systems are inefficient in the sample data alignment process and cannot support federated learning of large-scale and real-time data, especially when manual mapping is inefficient and time-consuming.

Method used

By obtaining the sample intersection identifier list, sample alignment is performed based on the privacy set intersection technology, and the task controller is used to realize dynamic pairing of trainers and two-stage sample alignment, thereby improving the sample alignment efficiency and model training efficiency.

Benefits of technology

It achieves efficient sample alignment under large-scale and real-time data conditions, supports sample sizes of tens of billions and streaming updates, is suitable for large-scale distributed training architectures, and improves the training efficiency of federated learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118396083B_ABST
    Figure CN118396083B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a model training method, system, device, and electronic device based on federated learning, which realizes automatic alignment of sample data of each participant's trainer through two-stage sample alignment, thereby improving the efficiency of sample alignment and the efficiency of model training based on federated learning. The method includes: obtaining a sample intersection identifier list; determining a first sample subset and a sample subset identifier corresponding to the first sample subset based on the sample intersection identifier list and the original sample subset, where the original sample subset is a sample distributed to the first trainer, and the original sample subset is a portion of the original sample of the first participant; sending the sample subset identifier to a second trainer paired with the first trainer, so that the second trainer determines a second sample subset based on the sample subset identifier and the original sample set of the second participant, and the first sample subset and the second sample subset are used for the federated learning-based model training of the first trainer and the second trainer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of federated learning, and in particular, to a model training method, system, device, and electronic device based on federated learning. Background Art

[0002] Federated learning is a technology that facilitates multiple data owners to collaborate on model training in a way that protects the security of original data. It is conducive to data sharing and cooperation and solves the problem of data silos.

[0003] Taking a vertical federated learning system as an example, within each participant, the participant's sample data is randomly distributed to multiple trainers by the participant's own distributed system. Therefore, strict data alignment is required to ensure the correctness of training metrics. Related technologies typically perform manual mapping of each participant's sample data before initiating training. This manual mapping is inefficient and time-consuming, and cannot support large-scale federated learning on real-time data. Summary of the Invention

[0004] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] In a first aspect, the present disclosure provides a model training method based on federated learning, applied to a first trainer, where the first trainer is a trainer of a first participant with a label, the method comprising:

[0006] Obtaining a sample intersection identifier list, the sample intersection identifier list including identifiers corresponding to samples in a sample intersection obtained after sample alignment between the first participant and the unlabeled second participant, wherein the sample intersection is obtained based on a privacy set intersection technique, the sample intersection including a first service feature and a second service feature for the same user group, the original sample of the first participant including the first service feature, and the original sample of the second participant including the second service feature;

[0007] Determining, based on the sample intersection identifier list and the original sample subset, a first sample subset and a sample subset identifier corresponding to the first sample subset; wherein the original sample subset is the sample distributed to the first trainer, and the original sample subset is a portion of the original samples of the first participant;

[0008] The sample subset identifier is sent to a second trainer paired with the first trainer, so that the second trainer determines a second sample subset based on the sample subset identifier and the original sample set of the second participant; wherein the second trainer is the trainer of the second participant, and the first sample subset and the second sample subset are used for federated learning-based model training of the first trainer and the second trainer.

[0009] In a second aspect, the present disclosure provides a model training method based on federated learning, which is applied to a second trainer, where the second trainer is a trainer of an unlabeled second participant, and the method includes:

[0010] Obtaining a sample subset identifier sent by a first trainer; wherein the first trainer is a trainer paired with the second trainer among the labeled first participants, the sample subset identifier is determined by the first trainer based on a sample intersection identifier list and an original sample subset, the original sample subset is the samples distributed to the first trainer, and the original sample subset is a portion of the original samples of the first participant, the sample intersection identifier list includes identifiers corresponding to each sample in a sample intersection obtained after sample alignment between the first participant and the second participant, wherein the sample intersection is obtained based on a privacy set intersection technique, the sample intersection includes a first business feature and a second business feature for the same user group, the original samples of the first participant include the first business feature, and the original samples of the second participant include the second business feature;

[0011] A second sample subset is determined based on the sample subset identifier and the original sample set of the second participant, and the first sample subset and the second sample subset are used for federated learning-based model training of the first trainer and the second trainer.

[0012] In a third aspect, the present disclosure provides a model training system based on federated learning, comprising a task controller, a plurality of first trainers belonging to a first participant with a label, and a plurality of second trainers belonging to a second participant without a label, wherein the first trainers are used to execute the method described in the first aspect above, the second trainers are used to execute the method described in the second aspect above, and the task controller is used to determine paired first trainers and second trainers.

[0013] In a fourth aspect, the present disclosure provides a model training device based on federated learning, which is applied to a first trainer, where the first trainer is a trainer of a first participant with a label, and the device includes:

[0014] a first acquisition module, configured to acquire a sample intersection identifier list, the sample intersection identifier list including identifiers corresponding to samples in a sample intersection obtained by aligning samples of the first participant with the unlabeled second participant, wherein the sample intersection is obtained based on a privacy set intersection technique, the sample intersection including a first service feature and a second service feature for the same user group, the original sample of the first participant including the first service feature, and the original sample of the second participant including the second service feature;

[0015] A first determining module is configured to determine, based on the sample intersection identifier list and the original sample subset, a sample subset identifier corresponding to the first sample subset and the first sample subset; wherein the original sample subset is the sample distributed to the first trainer, and the original sample subset is a portion of the original samples of the first participant;

[0016] A sending module is used to send the sample subset identifier to a second trainer paired with the first trainer, so that the second trainer determines a second sample subset based on the sample subset identifier and the original sample set of the second participant; wherein the second trainer is the trainer of the second participant, and the first sample subset and the second sample subset are used for federated learning-based model training of the first trainer and the second trainer.

[0017] In a fifth aspect, the present disclosure provides a model training device based on federated learning, which is applied to a second trainer, where the second trainer is a trainer of a second participant without labels, and the device includes:

[0018] a second acquisition module, configured to acquire a sample subset identifier sent by a first trainer; wherein the first trainer is a trainer paired with the second trainer among the labeled first participants, the sample subset identifier is determined by the first trainer based on a sample intersection identifier list and an original sample subset, the original sample subset is a sample distributed to the first trainer, and the original sample subset is a portion of the original samples of the first participant, the sample intersection identifier list includes identifiers corresponding to each sample in a sample intersection obtained after sample alignment between the first participant and the second participant, wherein the sample intersection is obtained based on a privacy set intersection technique, the sample intersection includes a first business feature and a second business feature for the same user group, the original samples of the first participant include the first business feature, and the original samples of the second participant include the second business feature;

[0019] A second determination module is used to determine a second sample subset based on the sample subset identifier and the original sample set of the second participant, and the first sample subset and the second sample subset are used for federated learning-based model training of the first trainer and the second trainer.

[0020] In a sixth aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in any one of the first or second aspects above.

[0021] In a seventh aspect, the present disclosure provides an electronic device, comprising:

[0022] a storage device having a computer program stored thereon;

[0023] A processing device is used to execute the computer program in the storage device to implement the steps of the method described in any one of the first aspect or the second aspect above.

[0024] In an eighth aspect, the present disclosure provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method described in any one of the first or second aspects above.

[0025] Through the above technical solution, the first trainer determines the first sample subset and the sample subset identifier corresponding to the first sample subset based on the sample intersection identifier list and the original sample subset, and sends the sample subset identifier to the second trainer paired with the first trainer, so that the second trainer can determine the second sample subset based on the sample subset identifier and the original sample set of the second participant. Using this method, the sample intersection identifier list obtained by performing the first sample alignment based on the original sample sets of each participant is first obtained, and then the first sample subset is determined based on the sample intersection identifier list and the partial samples distributed to itself, and the sample subset identifier corresponding to the first sample subset is sent to the paired second trainer, so that the second trainer can determine the corresponding second sample subset, thereby achieving the second sample alignment after the trainers are paired. Based on the two-stage sample alignment, the sample data of each participant's trainer is automatically aligned, improving the efficiency of sample alignment and the efficiency of model training based on federated learning, thereby supporting large-scale federated learning-based model training on real-time data.

[0026] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings:

[0028] Figure 1 1 is a schematic diagram of a process of vertical federated learning according to an exemplary embodiment of the present disclosure;

[0029] Figure 2 is a schematic diagram showing sample data distribution according to an exemplary embodiment of the present disclosure;

[0030] Figure 3 is a schematic diagram showing a “virtual” training sample set according to an exemplary embodiment of the present disclosure;

[0031] Figure 4 is a flowchart of a model training method based on federated learning according to an exemplary embodiment of the present disclosure;

[0032] Figure 5 is a schematic diagram of a model training method based on federated learning according to an exemplary embodiment of the present disclosure;

[0033] Figure 6 is a schematic diagram illustrating a process of determining a cache sample set according to an exemplary embodiment of the present disclosure;

[0034] Figure 7 is a schematic diagram of a model training system based on federated learning according to an exemplary embodiment of the present disclosure;

[0035] Figure 8 is a flowchart of a model training method based on federated learning according to an exemplary embodiment of the present disclosure;

[0036] Figure 9 is a flowchart of a model training method based on federated learning according to an exemplary embodiment of the present disclosure;

[0037] Figure 10 1 is a structural block diagram of a model training device based on federated learning according to an exemplary embodiment of the present disclosure;

[0038] Figure 11 1 is a structural block diagram of a model training device based on federated learning according to an exemplary embodiment of the present disclosure;

[0039] Figure 12 The figure is a schematic structural diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0040] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0041] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0042] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0043] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0044] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0045] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0046] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0047] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0048] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0049] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0050] At the same time, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0051] Federated learning can be categorized into horizontal federated learning, vertical federated learning, and federated transfer learning, depending on the distribution of sample datasets across multiple participants. Vertical federated learning refers to scenarios where there is significant sample overlap but minimal feature overlap between participants. For example, companies A and B share a common user base, but each possesses distinct business characteristics specific to that user base.

[0052] like Figure 1 As shown in the figure, in vertical federated learning, data is typically divided by features, and each participant has its own feature data. During training, the feature data of each participant is aligned using the same ID identifier. For example, in vertical federated learning with two participants, Party A provides the data label. Party A and Party B each train their own bottom model. Party A, based on the model training interaction layer, interactively connects Party B's forward embedding and backward gradient information to complete the training of the top model.

[0053] Continuing with the example of vertical federated learning between two participants, for data parallel multi-machine model training scenarios, for example, Party A and Party B each include two trainers, such as Figure 2As shown, Party A's sample data is randomly distributed to two trainers by its own distributed system. The first trainer holds the sample set corresponding to id1 and id4, while the second trainer holds the sample set corresponding to id2, id3, and id4. Similarly, Party B's sample data is also randomly distributed to two trainers by its own distributed system. The first trainer holds the sample set corresponding to id2, id5, id6, and id7, while the second trainer holds the sample set corresponding to id3.

[0054] According to the correctness and security requirements of federated learning, based on the common ID intersection obtained by private set intersection (PSI), the sample sets of both parties are logically formed as follows: Figure 3 The "virtual" training sample set shown is composed of the corresponding samples in the intersection of the two parties' samples, and the sample features and labels of each party need to be strictly aligned according to the ID.

[0055] Related technologies mainly restrict each participant to sequentially process all samples through a single trainer, and complete model training through federated learning on a single machine. This limits the concurrency of multiple machines and makes it difficult to achieve large-scale distributed concurrent training of massive sample data. Alternatively, the sample data of each participant is manually mapped before initiating training. However, manual mapping is inefficient and time-consuming, and cannot support large-scale federated learning of real-time data. In other words, the vertical federated learning system is limited by its own architecture, and the amount of data it supports is usually within the billions. The scale of the models it supports is also relatively limited. Large-scale vertical federated learning systems need to support large-scale distributed training architectures such as huge training sample volumes (such as tens of billions) and continuous streaming updates (such as billions of new samples added daily), and the training framework must be adapted to parameter servers.

[0056] In view of this, the present disclosure provides a model training method, system, device, and electronic device based on federated learning to solve the above technical problems. It should be noted that the model training method based on federated learning provided by the present disclosure can be applied to large-scale vertical federated learning scenarios, where each participant can be a distributed structure, that is, each participant includes multiple trainers.

[0057] The following further explains the embodiments of the present disclosure in conjunction with the accompanying drawings. For the convenience of description, the embodiment is described by taking Party A and Party B performing model training for vertical federated learning between two parties. In practical applications, it can be applied to model training for vertical federated learning of any number of parties.

[0058] The following further explains the embodiments of the present disclosure with reference to the accompanying drawings.

[0059] Figure 4This is a flow chart of a model training method based on federated learning according to an exemplary embodiment of the present disclosure. The method is applied to a first trainer, which is a trainer of a first participant with a label. Figure 4 , the method comprising:

[0060] S401: Obtain a sample intersection identifier list.

[0061] The sample intersection identification list includes the identification corresponding to each sample in the sample intersection obtained after the first participant and the unlabeled second participant align their samples, such as Figure 3 The sample IDs of the “virtual” training sample set are shown, where the sample intersection can be obtained based on the privacy set intersection technology.

[0062] S402: Determine the first sample subset and the sample subset identifier corresponding to the first sample subset based on the sample intersection identifier list and the original sample subset.

[0063] The original sample subset is the sample distributed to the first trainer, and the original sample subset is part of the original sample of the first participant, such as Figure 2 The participant shown distributes a sample set to a single trainer. The sample subset identifier can be the ID of the sample, which is not limited in the present disclosure.

[0064] In a possible manner, determining the first sample subset based on the sample intersection identifier list and the original sample subset may include: in the original sample subset, determining samples whose sample identifiers belong to the sample intersection identifier list as the first sample subset.

[0065] For example, Figure 2 and Figure 3 As an example of the sample data shown in the figure, the sample intersection identifier list includes Figure 3 The sample ids of the “virtual” training sample set are shown, and the first sample subset determined by the trainer A2 is the sample data corresponding to id2, id3 and id7.

[0066] It is worth noting that since the ID in the sample subset identification sent by the first trainer to the second trainer is obtained after filtering based on the privacy set intersection technology, there will be no ID outside the intersection of the two parties, and it does not involve the characteristic fields of the sample data, so it will not affect the security of federated learning.

[0067] S403: Send the sample subset identifier to the second trainer paired with the first trainer, so that the second trainer determines a second sample subset based on the sample subset identifier and the original sample set of the second participant.

[0068] The second trainer is a trainer of the second participant, and the first sample subset and the second sample subset are used for model training based on federated learning of the first trainer and the second trainer.

[0069] For example, taking the first trainer as trainer A2 and the second trainer as trainer B1, if the sample subset identified by trainer A2 is id2, id3 and id7, trainer B1 can filter out the sample data corresponding to id2, id3 and id7 from party B's original sample set, and use the sample data corresponding to id2, id3 and id7 as the second sample subset.

[0070] That is, the first trainer determines the first sample subset based on the sample intersection identifier list and the original sample subset it was distributed to. The second trainer does not use the original sample subset it was distributed to, but instead re-determines the second sample subset based on the sample subset identifier corresponding to the first sample subset determined by the first trainer. This enables dynamic sample alignment between the first and second trainers participating in federated learning model training, improving sample alignment efficiency. It is also applicable to large-scale distributed training architectures such as those with huge sample volumes (e.g., tens of billions) and continuous streaming updates (e.g., billions added daily), and those where the training framework is adapted to parameter servers.

[0071] Using this method, the first trainer first obtains a list of sample intersection identifiers from a first sample alignment based on the original sample sets of each participant. It then determines a first sample subset based on this list of sample intersection identifiers and some of its own distributed samples. The first trainer then sends the sample subset identifier corresponding to the first sample subset to the paired second trainer, allowing the second trainer to determine the corresponding second sample subset, thus achieving a second sample alignment after the trainers are paired. This two-stage sample alignment enables automatic alignment of sample data from each participant's trainer, improving the efficiency of both sample alignment and federated learning-based model training, thereby supporting large-scale federated learning-based model training on real-time data.

[0072] In a possible embodiment, the method further includes: after the first trainer is started, sending a registration request to the task controller, so that the task controller determines a second trainer paired with the first trainer in response to the registration request.

[0073] For example, after the first trainer is started, a registration request can be sent to the task controller. The task controller is used to pair the trainers of each participant participating in the model training based on federated learning. The task controller can be set on a security-neutral server, such as the coordinator in federated learning.

[0074] In a possible embodiment, the method further includes: sending a polling request regarding the pairing status to the task controller, so that the task controller, in response to the polling request, after determining the second trainer paired with the first trainer, sends the trainer identifier of the second trainer to the first trainer. Sending the sample subset identifier to the second trainer paired with the first trainer may include: sending the sample subset identifier to the second trainer corresponding to the trainer identifier.

[0075] For example, after sending a registration request to the task controller, the first trainer can periodically poll the task controller for pairing status. If the task controller has not yet determined the second trainer paired with the first trainer, it will not respond to the polling request. If the task controller has already determined the second trainer paired with the first trainer, it will send the trainer identifier of the second trainer to the first trainer. The first trainer can then exchange data with the second trainer based on the trainer identifier, thereby achieving collaborative training between the first and second trainers.

[0076] In a possible embodiment, the method further includes: when a polling request is sent to the task controller and no trainer identifier is received from the task controller within a first preset time period, or when the sample subset identifier is sent to a second trainer corresponding to the trainer identifier and no feedback message is received from the second trainer within a second preset time period, sending a new registration request to the task controller, so that the task controller redetermines the second trainer paired with the first trainer in response to the new registration request.

[0077] For example, after sending a registration request to the task controller, if the first trainer does not receive the trainer identifier sent by the task controller within a first preset time period, a pairing error may have occurred, or the task manager may not have received the registration request. In this case, a new registration request may be sent to the task controller so that the task controller can re-pair the trainers. Alternatively, the first trainer may be restarted and a new registration request may be sent to the task controller, although this disclosure is not limited thereto.

[0078] For example, if the data interaction between the first trainer and the second trainer is abnormal, for example, the second trainer does not respond to the message sent by the first trainer within the second preset time period, then the second trainer may have an abnormality, such as exiting the federated learning-based model training due to a fault. In this case, the first trainer may also resend a new registration request to the task controller so that the task controller can re-pair the trainers.

[0079] The first preset duration and the second preset duration may be equal or unequal, and this disclosure does not impose any restrictions thereon. Thus, when each trainer polls to request the current pairing status and a paired trainer performs a model training task, by setting a timeout, it is possible to re-pair the trainers in the event of a pairing anomaly or a trainer exiting due to a fault, thereby achieving flexible pairing of the trainers of all parties.

[0080] In a possible manner, the paired first trainer and second trainer are determined by the task controller in the following manner: when both the first list to be paired and the second list to be paired are non-empty, a first identifier is randomly selected from the first list to be paired, and a second identifier is randomly selected from the second list to be paired, and the trainer corresponding to the first identifier and the trainer corresponding to the second identifier are respectively used as the paired first trainer and second trainer; wherein the first list to be paired is used to store the identifier of the trainer of the first participant who sends the registration request, and the second list to be paired is used to store the identifier of the trainer of the second participant who sends the registration request.

[0081] For example, if trainers A1 and A2 of the first participant send a registration request to the task controller, the trainer identifiers A1 and A2 are stored in a first pending pairing list. Similarly, if trainers B1 and B2 of the second participant send a registration request to the task controller, the trainer identifiers B1 and B2 are stored in a second pending pairing list. The task controller randomly selects a trainer identifier from each of the first and second pending pairing lists, for example, A1 from the first pending pairing list and B1 from the second pending pairing list, marking them as paired.

[0082] Alternatively, taking the first participant's trainer A1 as an example, upon receiving a registration request from trainer A1, the task controller may first query whether there is a second participant's trainer to be paired in the second pending pairing list. If so, a trainer is randomly selected as the trainer to be paired with trainer A1. If not, the trainer identifier of trainer A1 is stored in the first pending pairing list, and pairing is performed after the second participant's trainer sends a registration request to the task controller. The specific pairing process can be determined based on needs, and this disclosure does not impose any restrictions on this.

[0083] It is worth noting that by introducing a task controller into the federated learning system, dynamic pairing of trainers can be achieved, and the information maintained by the task controller only involves the training task name, trainer identification and other information of each participant, and will not know the characteristic fields of any sample data of each participant, so it will not affect the security of federated learning.

[0084] Figure 5This is a flow chart of a model training method based on federated learning according to an exemplary embodiment of the present disclosure. The method is applied to the second trainer, which is a trainer of the second participant without labels. Figure 5 , the method comprising:

[0085] S501: Obtain a sample subset identifier sent by a first trainer.

[0086] Among them, the first trainer is the trainer of the labeled first participant that is paired with the second trainer, the sample subset identifier is determined by the first trainer based on the sample intersection identifier list and the original sample subset, the original sample subset is the sample distributed to the first trainer, and the original sample subset is part of the original samples of the first participant, and the sample intersection identifier list includes the identifier corresponding to each sample in the sample intersection obtained after the first participant and the second participant align their samples.

[0087] S502: Determine a second sample subset based on the sample subset identifier and the original sample set of the second participant.

[0088] The first sample subset and the second sample subset are used for model training based on federated learning by the first trainer and the second trainer.

[0089] By adopting the above method, after the first trainer of the labeled first participant determines the first sample subset, the second trainer can determine the corresponding second sample subset of the paired second trainer according to the sample subset identifier corresponding to the first sample subset. Based on the two-stage sample alignment, the sample data of the trainers of each participant are automatically aligned, thereby improving the sample alignment efficiency and the model training efficiency based on federated learning, thereby supporting large-scale federated learning-based model training on real-time data.

[0090] In a possible embodiment, the method further includes: obtaining a cached sample set, the cached sample set including samples whose sample identifiers belong to the sample intersection identifier list in the original sample set of the second participant. Determining the second sample subset based on the sample subset identifier and the original sample set of the second participant may include: determining, in the cached sample set, samples whose sample identifiers belong to the sample subset identifier as the second sample subset.

[0091] For example, refer to Figure 6Party A and Party B use the privacy set intersection technique and their full sample datasets (i.e., original sample sets) to obtain a sample intersection identifier list. After obtaining the sample intersection identifier list, Party B can filter its own full sample dataset based on the sample intersection identifier list to obtain the cached sample set corresponding to the sample intersection identifier list. This allows the second trainer to filter the cached sample set based on the sample subset identifiers when constructing the second sample subset. This improves the efficiency of the second trainer in constructing the second sample subset compared to filtering the full sample dataset.

[0092] In a possible manner, each sample in the cache sample set is stored in the form of a key-value pair, wherein the sample identifier of the sample is the key and the sample feature of the sample is the value corresponding to the key.

[0093] For example, refer to Figure 6 Each sample in the cache sample set can be stored in the form of a key-value pair, with the sample identifier of the sample as the Key and the sample feature of the sample as the Value corresponding to the Key. Storing in the form of a key-value pair can facilitate the second trainer to filter based on the sample identifier, further improving the efficiency of the second trainer in constructing the second sample subset.

[0094] It is worth noting that the present disclosure does not limit the storage component of the cache sample set. It can be a KV cache, a memory storage component such as redis, or a database system such as MySQL, and can be specifically set according to needs.

[0095] In a possible embodiment, the method further includes: after the second trainer is started, sending a registration request to the task controller, so that the task controller determines the first trainer paired with the second trainer in response to the registration request.

[0096] For example, after the second trainer is started, a registration request can be sent to the task controller. The task controller is used to pair the trainers of each participant in the federated learning. The specific pairing process has been described above and will not be repeated here in this disclosure.

[0097] In a possible embodiment, the method further includes: sending a polling request for the pairing status to the task controller, so that the task controller, in response to the polling request, sends the trainer identifier of the first trainer to the second trainer after determining the first trainer paired with the second trainer; and in response to receiving the trainer identifier of the first trainer, blocking and waiting for the sample subset identifier sent by the first trainer.

[0098] For example, after sending a registration request to the task controller, the second trainer can periodically poll the task controller for pairing status. If the task controller has not yet determined the first trainer paired with the second trainer, it does not respond to the polling request. If the task controller has already determined the first trainer paired with the second trainer, it sends the trainer identifier of the first trainer to the second trainer. After receiving the trainer identifier from the first trainer, the second trainer blocks and waits for the sample subset identifier sent by the first trainer. The first trainer can then exchange data with the second trainer based on the trainer identifier, thereby achieving collaborative training between the first and second trainers.

[0099] Figure 7 This is a framework of a model training system based on federated learning according to an exemplary embodiment of the present disclosure. Figure 7 The model training system 700 based on federated learning includes a task controller 701, multiple first trainers belonging to a labeled first participant 702, and multiple second trainers belonging to an unlabeled second participant 703. The first trainers are used to execute the above-mentioned method applied to the first trainers, and the second trainers are used to execute the above-mentioned method applied to the second trainers; the task controller 701 is used to determine the paired first trainers and second trainers.

[0100] Using this system, the first trainer can determine a first sample subset based on a list of sample intersection identifiers and some of its own distributed samples. The second trainer can then determine a corresponding second sample subset based on the sample subset identifier corresponding to the first sample subset. This two-stage sample alignment automatically aligns the sample data of each participant's trainer, improving the efficiency of both sample alignment and federated learning model training. This enables large-scale federated learning model training on real-time data.

[0101] The following uses the example of two parties participating in the model training based on federated learning to illustrate the interactive process of the model training method based on federated learning provided by the present disclosure.

[0102] Figure 8 1 is a flow chart of a model training method based on federated learning according to an exemplary embodiment of the present disclosure. Figure 8 , the method comprising:

[0103] S801: The first trainer obtains a sample intersection identifier list, and determines a first sample subset and a sample subset identifier corresponding to the first sample subset based on the sample intersection identifier list and the original sample subset, and sends the sample subset identifier to a second trainer paired with the first trainer.

[0104] Among them, the first trainer is the trainer of the first participant with a label, the second trainer is the trainer of the second participant without a label, the sample intersection identification list includes the identification corresponding to each sample in the sample intersection obtained after the first participant and the unlabeled second participant align the samples, the original sample subset is the sample distributed to the first trainer, and the original sample subset is part of the samples in the original sample of the first participant.

[0105] S802: The second trainer obtains the sample subset identifier sent by the first trainer, and determines a second sample subset based on the sample subset identifier and the original sample set of the second participant.

[0106] The first sample subset and the second sample subset are used for model training based on federated learning by the first trainer and the second trainer.

[0107] Using this method, we first obtain the sample intersection identifiers obtained from the first sample alignment based on the original sample sets of each participant. The first trainer can then determine the first sample subset based on the sample intersection identifier list and the portion of samples distributed to it. The second trainer can then determine the corresponding second sample subset based on the sample subset identifier corresponding to the first sample subset, achieving a second sample alignment after the trainers are paired. This two-stage sample alignment enables automatic alignment of sample data from each participant's trainer, improving the efficiency of sample alignment and federated learning-based model training, thereby supporting large-scale federated learning-based model training on real-time data.

[0108] For example, let's continue with the federated learning model training between Party A and Party B. First, we need to prepare the offline sample data preprocessing. Both parties export the ID sets corresponding to their respective original sample sets offline, and obtain the intersection ID, i.e., the sample intersection identification list, based on the privacy set intersection technology, to achieve the first sample alignment. Then, Party B, which does not provide labels, determines the cached sample set with the sample ID as the primary key based on the intersection ID and its own original sample set. You can refer to Figure 6 The data preprocessing process shown is not described in detail in this disclosure.

[0109] Then, both parties start the same number of trainers to train the model based on distributed federated learning. After a trainer on either side is started, it registers the trainer with the task controller. The registration content is its own trainer id, such as Figure 9As shown, trainer x on party A and trainer y on party B. After receiving a registration request from a party, the task controller places that trainer in its pending pairing list. If both pending pairing lists are not empty, it randomly selects a trainer ID from each party and identifies them as a pair. This pairs the trainers one-to-one, achieving dynamic pairing of trainers. For example, trainer x and trainer y are considered a pairing. After completing registration, each trainer can query the task controller to see if it has been paired. If the task controller has identified it as a pair, it notifies the trainer of the ID of its paired counterpart.

[0110] Continuing with the example of trainer x and trainer y as a pair of trainers, refer to Figure 9 After learning the ID of its paired trainer y, trainer x filters the original sample subset distributed to trainer x based on the intersection ID, retains the features and labels of the samples with IDs in the intersection and enters the bottom model of its own side, and records the ordered list of IDs of these samples, that is, the sample subset identifier, and sends the ordered list of IDs to the paired trainer y. After learning the ID of its paired trainer x, trainer y blocks and waits for trainer x to send the ordered list of IDs. After receiving the ordered list of IDs, it queries the corresponding samples in the cached sample set of party B based on the ordered list of IDs, obtains the second sample subset, and enters the bottom model of its own side in the same order as its own samples, realizing the second sample alignment after the trainers are paired. The two paired trainers in each group perform forward data interaction and reverse gradient propagation for model training based on the federated learning algorithm to complete the training process.

[0111] It's worth noting that the present disclosure doesn't rely on manual sample alignment. Instead, it enables regular incremental preprocessing and real-time online matching of multiple trainers, thus supporting continuous training of new sample data in a single federated learning session. Furthermore, participants in federated learning can launch any number of trainers to participate in training simultaneously, eliminating the need to serially process massive amounts of sample data on a single trainer, thereby improving overall training efficiency.

[0112] Furthermore, considering that trainer x may exit federated learning due to pairing failure, the original sample subset is filtered based on the intersection ID after successful pairing. However, given that task controller pairing takes a considerable amount of time, trainer x can also filter the original sample subset based on the intersection ID before pairing. Then, after learning the ID of its paired trainer y, trainer x can directly send the ordered list of IDs to trainer y. This can be configured as needed and is not limited by this disclosure.

[0113] This approach supports large-scale distributed federated learning through two-stage sample alignment and dynamic pairing of distributed trainers from multiple parties. Each participant can launch multiple trainers to simultaneously train federated learning-based models, improving training efficiency and expanding the amount of sample data and model parameter size supported by the federated learning system. By splitting the data preprocessing phase into two, real-time online training, a large-scale, horizontally scalable, multi-party fine-grained sample alignment mechanism is implemented. Dynamic pairing management of distributed trainers from multiple parties is implemented through a secure and neutral task controller, supporting efficient and stable data-parallel training across multiple parties.

[0114] Based on the same concept, the embodiment of the present disclosure also provides a model training device based on federated learning, which is applied to a first trainer, which is a trainer of a first participant with a label, referring to Figure 10 , the model training device 100 based on federated learning includes:

[0115] A first acquisition module 1001 is configured to acquire a sample intersection identifier list, wherein the sample intersection identifier list includes identifiers corresponding to samples in a sample intersection obtained after the first participant and the unlabeled second participant align their samples;

[0116] A first determining module 1002 is configured to determine, based on the sample intersection identifier list and the original sample subset, a sample subset identifier corresponding to the first sample subset and the first sample subset; wherein the original sample subset is the sample distributed to the first trainer, and the original sample subset is a portion of the original samples of the first participant;

[0117] A sending module 1003 is configured to send the sample subset identifier to a second trainer paired with the first trainer, so that the second trainer determines a second sample subset based on the sample subset identifier and the original sample set of the second participant; wherein the second trainer is the trainer of the second participant, and the first sample subset and the second sample subset are used for federated learning-based model training of the first trainer and the second trainer.

[0118] Optionally, the first determining module 1002 is configured to:

[0119] In the original sample subset, samples whose sample identifiers belong to the sample intersection identifier list are determined as the first sample subset.

[0120] Optionally, the model training device 100 based on federated learning further includes:

[0121] The first registration module is configured to send a registration request to the task controller after the first trainer is started, so that the task controller determines a second trainer paired with the first trainer in response to the registration request.

[0122] Optionally, the model training device 100 based on federated learning further includes:

[0123] a first polling module configured to send a polling request for a pairing status to the task controller, so that the task controller, in response to the polling request, sends a trainer identifier of the second trainer to the first trainer after determining a second trainer paired with the first trainer;

[0124] The sending module 1003 is configured to send the sample subset identifier to the second trainer corresponding to the trainer identifier.

[0125] Optionally, the paired first trainer and the second trainer are determined by the task controller in the following manner:

[0126] When both the first list to be paired and the second list to be paired are not empty, randomly selecting a first identifier from the first list to be paired and a second identifier from the second list to be paired, and using the trainer corresponding to the first identifier and the trainer corresponding to the second identifier as the first trainer and the second trainer for pairing, respectively;

[0127] The first to-be-paired list is used to store the identifier of the trainer of the first participant who sends the registration request, and the second to-be-paired list is used to store the identifier of the trainer of the second participant who sends the registration request.

[0128] Optionally, the model training device 100 based on federated learning further includes:

[0129] A timeout module is configured to send a new registration request to the task controller when a polling request is sent to the task controller and no trainer identifier is received from the task controller within a first preset time period, or when the sample subset identifier is sent to the second trainer corresponding to the trainer identifier and no feedback message is received from the second trainer within a second preset time period, so that the task controller redetermines the second trainer paired with the first trainer in response to the new registration request.

[0130] Based on the same concept, the embodiment of the present disclosure also provides a model training device based on federated learning, which is applied to the second trainer, which is a trainer of the second participant without labels. Figure 11 , the model training device 110 based on federated learning includes:

[0131] A second acquisition module 1101 is configured to acquire a sample subset identifier sent by a first trainer; wherein the first trainer is a trainer paired with the second trainer in a labeled first participant, the sample subset identifier is determined by the first trainer based on a sample intersection identifier list and an original sample subset, the original sample subset being samples distributed to the first trainer, and the original sample subset being a portion of the original samples of the first participant, and the sample intersection identifier list including identifiers corresponding to samples in the sample intersection obtained after sample alignment between the first participant and the second participant;

[0132] The second determination module 1102 is used to determine a second sample subset based on the sample subset identifier and the original sample set of the second participant, and the first sample subset and the second sample subset are used for federated learning-based model training of the first trainer and the second trainer.

[0133] Optionally, the model training device 110 based on federated learning further includes:

[0134] A cache module, configured to obtain a cached sample set, wherein the cached sample set includes samples in the original sample set of the second participant whose sample identifiers belong to the sample intersection identifier list;

[0135] The second determining module 1102 is configured to:

[0136] In the cached sample set, samples whose sample identifiers belong to the sample subset identifier are determined as the second sample subset.

[0137] Optionally, each sample in the cache sample set is stored in the form of a key-value pair, wherein the sample identifier of the sample is a keyword, and the sample feature of the sample is a value corresponding to the keyword.

[0138] Optionally, the model training device 110 based on federated learning further includes:

[0139] The second registration module is configured to send a registration request to the task controller after the second trainer is started, so that the task controller determines a first trainer paired with the second trainer in response to the registration request.

[0140] Optionally, the model training device 110 based on federated learning further includes:

[0141] a second polling model for sending a polling request for a pairing status to the task controller, so that the task controller, in response to the polling request, sends a trainer identifier of the first trainer to the second trainer after determining a first trainer paired with the second trainer;

[0142] The waiting module is configured to, in response to receiving the trainer identifier of the first trainer, block and wait for the sample subset identifier sent by the first trainer.

[0143] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0144] Based on the same concept, an embodiment of the present disclosure also provides a computer-readable medium on which a computer program is stored, which, when executed by a processing device, implements the steps of any model training method based on federated learning.

[0145] Based on the same concept, an embodiment of the present disclosure further provides an electronic device, which may include:

[0146] a storage device having a computer program stored thereon;

[0147] A processing device is used to execute the computer program in the storage device to implement the steps of any of the above-mentioned model training methods based on federated learning.

[0148] Based on the same concept, an embodiment of the present disclosure also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of any of the above-mentioned model training methods based on federated learning.

[0149] Reference below Figure 12 , which shows a schematic structural diagram of an electronic device 120 suitable for implementing the embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 12 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0150] like Figure 12As shown, electronic device 120 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 121, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 122 or programs loaded from a storage device 128 into a random access memory (RAM) 123. RAM 123 also stores various programs and data required for the operation of electronic device 120. Processing device 121, ROM 122, and RAM 123 are connected to each other via a bus 124. An input / output (I / O) interface 125 is also connected to bus 124.

[0151] Typically, the following devices may be connected to the I / O interface 125: an input device 126 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 127 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 128 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 129. The communication device 129 may allow the electronic device 120 to communicate with other devices wirelessly or by wire to exchange data. Figure 12 The electronic device 120 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0152] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 129, or installed from the storage device 128, or installed from the ROM 122. When the computer program is executed by the processing device 121, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0153] It should be noted that the computer-readable medium described above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.

[0154] In some embodiments, communications may be conducted using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0155] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0156] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains a sample intersection identification list, wherein the sample intersection identification list includes identifications corresponding to samples in the sample intersection obtained after the first participant and the unlabeled second participant perform sample alignment; determines the sample subset identifications corresponding to the first sample subset and the first sample subset based on the sample intersection identification list and the original sample subset; wherein the original sample subset is the sample distributed to the first trainer, and the original sample subset is part of the original samples of the first participant; sends the sample subset identification to the second trainer paired with the first trainer, so that the second trainer determines the second sample subset based on the sample subset identification and the original sample set of the second participant; wherein the second trainer is the trainer of the second participant, and the first sample subset and the second sample subset are used for the federated learning-based model training of the first trainer and the second trainer.

[0157] Alternatively, the computer-readable medium carries one or more programs, which, when executed by the electronic device, causes the electronic device to: obtain a sample subset identifier sent by a first trainer; wherein the first trainer is a trainer paired with the second trainer among the labeled first participants, and the sample subset identifier is determined by the first trainer based on a sample intersection identifier list and an original sample subset, the original sample subset is a sample distributed to the first trainer, and the original sample subset is a portion of the original samples of the first participant, and the sample intersection identifier list includes identifiers corresponding to each sample in the sample intersection obtained after sample alignment between the first participant and the second participant; determine a second sample subset based on the sample subset identifier and the original sample set of the second participant, and the first sample subset and the second sample subset are used for federated learning-based model training of the first trainer and the second trainer.

[0158] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0160] The modules described in the embodiments of the present disclosure may be implemented in software or hardware, wherein the name of a module does not necessarily limit the module itself.

[0161] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0162] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0163] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the scope of the above disclosure. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0164] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0165] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.

Claims

1. A model training method based on federated learning, characterized in that: Applied to a first trainer, where the first trainer is a trainer of a first participant with a label, the method includes: Obtaining a sample intersection identifier list, the sample intersection identifier list including identifiers corresponding to samples in a sample intersection obtained after sample alignment between the first participant and the unlabeled second participant, wherein the sample intersection is obtained based on a privacy set intersection technique, the sample intersection including a first service feature and a second service feature for the same user group, the original sample of the first participant including the first service feature, and the original sample of the second participant including the second service feature; Determining, based on the sample intersection identifier list and the original sample subset, a first sample subset and a sample subset identifier corresponding to the first sample subset; wherein the original sample subset is the sample distributed to the first trainer, and the original sample subset is a portion of the original samples of the first participant; The sample subset identifier is sent to a second trainer paired with the first trainer, so that the second trainer determines a second sample subset based on the sample subset identifier and the original sample set of the second participant; wherein the second trainer is the trainer of the second participant, and the first sample subset and the second sample subset are used for federated learning-based model training of the first trainer and the second trainer.

2. The method according to claim 1, characterized in that The determining of the first sample subset based on the sample intersection identifier list and the original sample subset includes: In the original sample subset, samples whose sample identifiers belong to the sample intersection identifier list are determined as the first sample subset.

3. The method according to claim 1 or 2, characterized in that The method further comprises: After the first trainer is started, a registration request is sent to the task controller, so that the task controller determines a second trainer paired with the first trainer in response to the registration request.

4. The method according to claim 3, characterized in that The method further comprises: sending a polling request for a pairing status to the task controller, so that the task controller, in response to the polling request, sends a trainer identifier of the second trainer to the first trainer after determining a second trainer paired with the first trainer; The sending the sample subset identifier to the second trainer paired with the first trainer includes: sending the sample subset identifier to the second trainer corresponding to the trainer identifier.

5. The method according to claim 3, characterized in that The paired first trainer and second trainer are determined by the task controller in the following manner: When both the first list to be paired and the second list to be paired are not empty, randomly selecting a first identifier from the first list to be paired and randomly selecting a second identifier from the second list to be paired, and using the trainer corresponding to the first identifier and the trainer corresponding to the second identifier as the first trainer and the second trainer for pairing, respectively; The first to-be-paired list is used to store the identifier of the trainer of the first participant who sends the registration request, and the second to-be-paired list is used to store the identifier of the trainer of the second participant who sends the registration request.

6. The method according to claim 4, characterized in that The method further comprises: In the case where a polling request is sent to the task controller and the trainer identifier sent by the task controller is not received within a first preset time period, or in the case where the sample subset identifier is sent to the second trainer corresponding to the trainer identifier and the feedback message sent by the second trainer is not received within a second preset time period, a new registration request is sent to the task controller so that the task controller redetermines the second trainer paired with the first trainer in response to the new registration request.

7. A model training method based on federated learning, characterized in that: Applied to a second trainer, where the second trainer is a trainer of a second participant without a label, the method includes: Obtaining a sample subset identifier sent by a first trainer; wherein the first trainer is a trainer paired with the second trainer among the labeled first participants, the sample subset identifier is determined by the first trainer based on a sample intersection identifier list and an original sample subset, the original sample subset is the samples distributed to the first trainer, and the original sample subset is a portion of the original samples of the first participant, the sample intersection identifier list includes identifiers corresponding to each sample in a sample intersection obtained after sample alignment between the first participant and the second participant, wherein the sample intersection is obtained based on a privacy set intersection technique, the sample intersection includes a first business feature and a second business feature for the same user group, the original samples of the first participant include the first business feature, and the original samples of the second participant include the second business feature; A second sample subset is determined based on the sample subset identifier and the original sample set of the second participant, and the first sample subset and the second sample subset are used for federated learning-based model training of the first trainer and the second trainer.

8. The method according to claim 7, characterized in that The method further comprises: Acquire a cached sample set, the cached sample set including samples in the original sample set of the second participant whose sample identifiers belong to the sample intersection identifier list; The determining the second sample subset based on the sample subset identifier and the original sample set of the second participant includes: In the cached sample set, samples whose sample identifiers belong to the sample subset identifier are determined as the second sample subset.

9. The method according to claim 8, characterized in that Each sample in the cache sample set is stored in the form of a key-value pair, wherein the sample identifier of the sample is a keyword and the sample feature of the sample is a value corresponding to the keyword.

10. The method according to any one of claims 7 to 9, characterized in that: The method further comprises: After the second trainer is started, a registration request is sent to the task controller, so that the task controller determines a first trainer paired with the second trainer in response to the registration request.

11. The method according to claim 10, characterized in that The method further comprises: sending a polling request for a pairing status to the task controller, so that the task controller, in response to the polling request, sends a trainer identification of the first trainer to the second trainer after determining a first trainer paired with the second trainer; In response to receiving the trainer identifier of the first trainer, blocking and waiting for the sample subset identifier sent by the first trainer.

12. A model training system based on federated learning, characterized in that: It includes a task controller, multiple first trainers belonging to labeled first participants and multiple second trainers belonging to unlabeled second participants, the first trainers are used to execute the method described in any one of claims 1 to 6, the second trainers are used to execute the method described in any one of claims 7 to 11, and the task controller is used to determine the paired first trainers and second trainers.

13. A model training device based on federated learning, characterized in that: Applied to a first trainer, the first trainer being a trainer of a first participant with a label, the apparatus comprises: a first acquisition module, configured to acquire a sample intersection identifier list, the sample intersection identifier list including identifiers corresponding to samples in a sample intersection obtained by aligning samples of the first participant with the unlabeled second participant, wherein the sample intersection is obtained based on a privacy set intersection technique, the sample intersection including a first service feature and a second service feature for the same user group, the original sample of the first participant including the first service feature, and the original sample of the second participant including the second service feature; A first determining module is configured to determine, based on the sample intersection identifier list and the original sample subset, a sample subset identifier corresponding to the first sample subset and the first sample subset; wherein the original sample subset is the sample distributed to the first trainer, and the original sample subset is a portion of the original samples of the first participant; A sending module is used to send the sample subset identifier to a second trainer paired with the first trainer, so that the second trainer determines a second sample subset based on the sample subset identifier and the original sample set of the second participant; wherein the second trainer is the trainer of the second participant, and the first sample subset and the second sample subset are used for federated learning-based model training of the first trainer and the second trainer.

14. A model training device based on federated learning, characterized in that: Applied to a second trainer, where the second trainer is a trainer of a second participant without a label, the apparatus comprises: a second acquisition module, configured to acquire a sample subset identifier sent by a first trainer; wherein the first trainer is a trainer paired with the second trainer among the labeled first participants, the sample subset identifier is determined by the first trainer based on a sample intersection identifier list and an original sample subset, the original sample subset is a sample distributed to the first trainer, and the original sample subset is a portion of the original samples of the first participant, the sample intersection identifier list includes identifiers corresponding to each sample in a sample intersection obtained after sample alignment between the first participant and the second participant, wherein the sample intersection is obtained based on a privacy set intersection technique, the sample intersection includes a first business feature and a second business feature for the same user group, the original samples of the first participant include the first business feature, and the original samples of the second participant include the second business feature; A second determination module is used to determine a second sample subset based on the sample subset identifier and the original sample set of the second participant, and the first sample subset and the second sample subset are used for federated learning-based model training of the first trainer and the second trainer.

15. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 11 are implemented.

16. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 11.

17. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Decision model training method, prediction method and device based on longitudinal federation learning

    CN111598186A

  • Longitudinal federated learning method under different sample identifiers, equipment and medium

    CN115630713A