Model training method based on longitudinal federated learning, medium, equipment and product

By adopting the publish subscription architecture and channel caching mechanism in vertical federated learning, the asynchronous model training is realized, the delay problem caused by synchronous training is solved, and the training efficiency and accuracy are improved.

CN120235271AActive Publication Date: 2025-07-01BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510573068.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-01
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The existing vertical federated learning methods have computational burden and communication overhead in model training, and synchronous training leads to a long training delay, affecting efficiency.

Method used

The publish subscription architecture and channel caching mechanism are adopted to realize asynchronous model training, and the publisher and subscriber roles are alternately trained, and the current batch of training samples and embed vectors are used to update the model, reducing synchronization dependencies.

Benefits of technology

Asynchronous training eliminates the training delay caused by synchronous dependencies, improves model training efficiency, and ensures the alignment of sample identification at any time, improving the accuracy of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235271A_ABST
    Figure CN120235271A_ABST
Patent Text Reader

Abstract

The invention discloses a longitudinal federated learning-based model training method, a medium, equipment and a product. The method comprises the steps of inputting a first training sample into a first feature extraction model to obtain a first embedding vector, putting the first embedding vector into a first channel corresponding to a current batch, and if a non-empty second channel exists in N second channels, extracting first gradient information stored earliest from the non-empty second channel to update the first feature extraction model, the N batches are in one-to-one correspondence with the N first channels and the N second channels, and the second channels are used for storing first gradient information of the first training samples of the corresponding batches; and if the N second channels are all empty, obtaining the next batch of training samples to execute the same process, and ending until the N batches are all executed. Through a publishing and subscribing architecture and a channel caching mechanism, a multi-party training process is decoupled, asynchronous training of the model is realized, and the model training efficiency is improved. And sample identifiers can be aligned at any time, and the accuracy of model training is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of federated learning, and in particular, to a model training method, medium, device, and product based on vertical federated learning. Background Art

[0002] In vertical federated learning (VFL), the sample identifiers (e.g., IDs) of data are aligned among different participants, but the feature sets are different. As Figure 1 shown, the labeled organization A has the first feature data of the first object (i.e., the private data of organization A), while the unlabeled organization B has the second feature data of the first object (which belongs to the private data of organization B). For example, organization A is a bank and organization B is an e-commerce. Among them, the bank has the financial data of users, and the e-commerce has the consumption records of users. The label can be whether the user will overdue repayment or the credit rating of the user. The two parties hope to cooperate in training the model, but cannot directly share the original feature data. Traditional VFL uses federated gradient transmission, that is, each party calculates the local gradient and exchanges information through encryption technologies (such as secure multi-party computation or homomorphic encryption), and finally optimizes the global model. However, this method has a large communication overhead and high computational cost.

[0003] In order to reduce the computational burden and communication overhead in the model training process, currently, the model training is usually based on VFL of split learning. However, this model training method usually adopts synchronous training, which may affect the model training efficiency and result in a large training delay. Summary of the Invention

[0004] This Summary of the Invention section is provided to introduce concepts in a concise form that will be described in detail in the subsequent Detailed Description section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.

[0005] In a first aspect, the present disclosure provides a model training method based on vertical federated learning. The multiple parties participating in the model training include a first party with labels and M second parties without labels, where M ≥ 1. The data held by each of the multiple parties is divided into N batches after sample identification alignment, where N ≥ 1. The N batches correspond one-to-one to N first channels and also correspond one-to-one to N second channels. The method is applied to any one of the M second parties without labels, and the method includes: obtaining a first training sample of the current batch, where the first training sample of the current batch is one of the N batches; inputting the first training sample into a first feature extraction model to obtain a first embedding vector; putting the first embedding vector into the first channel corresponding to the current batch; if there is a non-empty second channel among the N second channels, taking out the earliest stored first gradient information from the non-empty second channel; updating the first feature extraction model using the earliest stored first gradient information; where the second channel is used to store the first gradient information of the first training sample of the corresponding batch, and the first gradient information is obtained by the first party training an inference model based on a second embedding vector and M first embedding vectors taken out from the first channel corresponding to the corresponding batch, and the second embedding vector is obtained by the first party inputting the second training sample of the corresponding batch into a second feature extraction model; if all the N second channels are empty, obtaining the training sample of the next batch and executing the same process until all the N batches are executed and then ending.

[0006] Second aspect, the present disclosure provides a model training method based on vertical federated learning. The multiple parties participating in the model training include a first party with labels and M second parties without labels, where M ≥ 1. The data held by each of the multiple parties is divided into N batches after sample identification alignment, where N ≥ 1. The N batches correspond one-to-one to N first channels and also correspond one-to-one to N second channels. The method is applied to the first party, and the method includes: obtaining a second training sample of the current batch, where the second training sample of the current batch is one of the N batches; inputting the second training sample into a second feature extraction model to obtain a second embedding vector; taking out M first embedding vectors from the first channel corresponding to the current batch; where the M first embedding vectors are obtained by each of the M second parties without labels inputting the first training sample of the current batch into a local first feature extraction model and putting them into the first channel corresponding to the current batch; training an inference model based on the second embedding vector and the M first embedding vectors to obtain first gradient information, second gradient information, and third gradient information; putting the first gradient information into the second channel corresponding to the current batch, and using the second gradient information to update the second feature extraction model, and using the third gradient information to update the inference model.

[0007] Third aspect, the present disclosure provides a computer-readable medium, on which a computer program is stored. When the computer program is executed by a processing device, it implements the steps of the method provided in the first aspect of the present disclosure or the steps of the method provided in the second aspect of the present disclosure.

[0008] Fourth aspect, the present disclosure provides an electronic device, including: a storage device, on which a computer program is stored; a processing device, configured to execute the computer program in the storage device to implement the steps of the method provided in the first aspect of the present disclosure or the steps of the method provided in the second aspect of the present disclosure.

[0009] Fifth aspect, the present disclosure provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the method provided in the first aspect of the present disclosure or the steps of the method provided in the second aspect of the present disclosure.

[0010] In the above technical solution, before the M unlabeled second participants and the labeled first participant perform model training based on vertical federated learning, they first align the sample identifiers of the data they hold, divide the aligned sample sets into N batches respectively, and establish N first channels and N second channels corresponding to the N batches one by one. Among them, the first channel is used to store the embedding vectors of the first training samples in the corresponding batch, and the second channel is used to store the gradient information of the first training samples in the corresponding batch. When performing batch model training, the second participant can use the first training samples in the current batch to train the first feature extraction model to obtain the first embedding vector. Then, the second participant, as the publisher, puts the first embedding vector into the first channel corresponding to the current batch. At the same time, the first participant can use the second training samples in the current batch to train the second feature extraction model to obtain the second embedding vector. Then, the first participant, as the subscriber, takes out the M first embedding vectors from the first channel corresponding to the current batch, and based on the second embedding vector and the M first embedding vectors taken out, trains the inference model to obtain the first gradient information, the second gradient information, and the third gradient information. Next, the first participant, as the publisher again, puts the first gradient information into the second channel corresponding to the current batch, and updates the second feature extraction model using the second gradient information and updates the inference model using the third gradient information. After the second participant puts the first embedding vector into the corresponding first channel, it can, as the subscriber, determine whether there is first gradient information in the N second channels. If there is a non-empty second channel among the N second channels, it takes out the earliest stored first gradient information from the non-empty second channel and updates the first feature extraction model using the earliest stored first gradient information. If there is no first gradient information in all N second channels, that is, all N second channels are empty, it does not need to wait for the gradient information of the first training samples in the current batch fed back by the first participant and can continue to train the first feature extraction model. Thus, through the publish-subscribe architecture and the channel caching mechanism, the training processes of multiple participants are decoupled, and asynchronous training of the model is achieved, thereby eliminating the training delay caused by synchronous dependencies in vertical federated learning and improving the model training efficiency. In addition, the embedding vectors and gradient information of the training samples in different batches are put into the channels corresponding to the corresponding batches, which can ensure that the sample identifiers are aligned at any time, thereby ensuring the accuracy of model training.

[0011] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation part. Brief Description of the Drawings

[0012] Combined with the drawings and referring to the following specific implementation manners, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale. In the drawings:

[0013] Figure 1 It is a schematic diagram of vertical federated learning based on split learning in the related art.

[0014] Figure 2 It is a schematic diagram of an overview of synchronous dependency analysis in the related art.

[0015] Figure 3 It is a flowchart of a model training method based on vertical federated learning applied to any unlabeled second participant according to an exemplary embodiment.

[0016] Figure 4 It is an overview diagram of a vertical federated learning system according to an exemplary embodiment.

[0017] Figure 5 It is a flowchart of a vertical federated learning system according to an exemplary embodiment.

[0018] Figure 6 It is a flowchart of a model training method based on vertical federated learning applied to a labeled first participant according to an exemplary embodiment.

[0019] Figure 7 It is a block diagram of a model training device based on vertical federated learning applied to any unlabeled second participant according to an exemplary embodiment.

[0020] Figure 8 It is a block diagram of a model training device based on vertical federated learning applied to a labeled first participant according to an exemplary embodiment.

[0021] Figure 9 It is a schematic diagram of the structure of an electronic device according to an exemplary embodiment. Detailed implementation manners

[0022] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0023] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0024] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0025] It should be noted that the concepts such as "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0026] It should be noted that the modification of "one" and "multiple" mentioned in this disclosure is illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0027] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0028] It can be understood that before using the technical solutions disclosed in the embodiments of this disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0029] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of this disclosure according to the prompt message.

[0030] As an optional but non-limiting implementation manner, when responding to receiving an active request from a user, the manner of sending a prompt message to the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0031] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manners of this disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manners of this disclosure.

[0032] Meanwhile, it can be understood that the data involved in the present technical solution (including but not limited to the data itself, data acquisition or use) should comply with the requirements of corresponding laws, regulations and relevant provisions.

[0033] As discussed in the background art, in order to reduce the computational burden and communication overhead during model training, at present, model training is usually carried out based on Split Learning-based VFL. However, this model training method usually adopts synchronous training, which may affect the model training efficiency and result in a large training delay.

[0034] Specifically, Split Learning reduces the computational burden and data exposure risk by splitting the model to be trained (e.g., a deep neural network) into multiple sub-models, so that each participating party only trains its own responsible sub-model. As Figure 1 shown, the model to be trained is split into a bottom model and a top model. Before model training, Organization A and Organization B need to align the private data samples of each other. For example, the intersection of common sample identifiers can be obtained based on the Private Set Intersection (PSI) technology. Then, model training is carried out based on the aligned private data.

[0035] Among them, the model to be trained can be a classification model, which can be used in scenarios such as image classification, text classification, and group classification. The model to be trained can also be a prediction model, which can be applied to scenarios such as trend prediction.

[0036] Specifically, as Figure 1As shown, Organization B can train the bottom model it holds based on the aligned private data to obtain an embedding vector (which belongs to the hidden representation). This process can be called the forward propagation of Organization B's bottom model, and the embedding vector is sent to Organization A after being noise-added through a secure protocol (e.g., differential privacy protocol); Organization A trains the bottom model it holds based on the aligned private data to obtain an embedding vector. This process can be called the forward propagation of Organization A's bottom model. After that, Organization A trains the top model based on the locally generated embedding vector and the noise-added embedding vector received from Organization B. This process can be called the forward propagation of Organization A's top model, and based on the label Y corresponding to the private data and the output result of the top model, the model loss is calculated, and based on the model loss, the gradient information of Organization A's aligned private data, the gradient information of Organization B's aligned private data, and the gradient information of the embedding vector are calculated; Next, Organization A updates the bottom model it holds based on the gradient information of its own aligned private data. This process can be called the backpropagation of Organization A's bottom model, and updates the top model based on the gradient information of the embedding vector. This process can be called the backpropagation of Organization A's top model; At the same time, Organization A sends the gradient information of Organization B's aligned private data to Organization B after noise-adding through a secure protocol, and Organization B can update the bottom model it holds based on the received noise-added gradient information. This process can be called the backpropagation of Organization B's bottom model.

[0037] As Figure 2 shown, Organization A and Organization B perform model training synchronously, and there are synchronous dependencies between different processes. For example, the backpropagation process of Organization A's top model needs to wait until the forward propagation of Organization B's bottom model is completed before it can be executed; the backpropagation process of Organization B's bottom model needs to wait until the backpropagation of Organization A's top model is completed before it can be executed. The existence of synchronous dependencies between different processes will affect the model training efficiency and result in a large training delay.

[0038] In view of this, the present disclosure provides a model training method, medium, device, and product based on vertical federated learning.

[0039] Figure 3 is a flowchart of a model training method based on vertical federated learning applied to any unlabeled second participant shown according to an exemplary embodiment. As Figure 3 shown, the model training method based on vertical federated learning applied to any unlabeled second participant may include the following S101 to S106.

[0040] In S101, obtain the first training samples of the current batch.

[0041] In the present disclosure, multiple participating parties involved in model training include a first participating party with labels and M second participating parties without labels, where M ≥ 1, that is, the number of participating parties without labels involved in model training can be one or more. Each of the M second participating parties holds first private sample data, and the first participating party holds second private sample data and the labels corresponding to the second private sample data. The multiple participating parties can perform model training based on VFL. Before model training, private sample data alignment is required. For example, the intersection of common sample identifiers can be obtained based on PSI. Then, the first private sample data corresponding to the intersection of sample identifiers is used as the first sample set, and the second private sample data corresponding to the intersection of sample identifiers is used as the second sample set. Each of the first sample set and the second sample set is divided according to a preset batch size to obtain N batches respectively, that is, the data held by each of the multiple participating parties is divided into N batches after sample identifier alignment. The M second participating parties and the first participating party can perform batch training on the model. Here, the batch size is the number of samples included in each batch. The training samples in the same batch of the first sample set and the second sample set have the same sample identifier, and N ≥ 1. The first training sample of the current batch is one of the above N batches.

[0042] The model to be trained can be split into a feature extraction model (bottom model) and an inference model (top model). Each of the M second participating parties without labels is used to train its own first feature extraction model according to the first sample set, and the first participating party is used to train its own second feature extraction model and inference model according to the second sample set and the corresponding labels. Here, the structures of the first feature extraction model and the second feature extraction model are the same, and their initial weights can be the same or different.

[0043] In S102, the first training sample is input into the first feature extraction model to obtain a first embedding vector.

[0044] In S103, the first embedding vector is placed into the first channel corresponding to the current batch. The N batches correspond one-to-one with N first channels (also called embedding channels), and the N batches also correspond one-to-one with N second channels (also called gradient channels). The second channel is used to store the first gradient information of the first training sample of the corresponding batch. The first gradient information is obtained by the first participating party training the inference model based on the second embedding vector and the M first embedding vectors taken from the first channel corresponding to the corresponding batch.

[0045] In the present disclosure, the second embedding vector is obtained by the first participating party inputting the corresponding batch of second training samples into the second feature extraction model. Taking the embedding vector from the first channel corresponding to the corresponding batch means removing the embedding vector from the first channel corresponding to the corresponding batch. After removal, the first channel corresponding to the corresponding batch no longer contains the removed embedding vector.

[0046] To ensure that the sample identifiers are aligned at any time during the model training process, N first channels and N second channels that correspond one-to-one with N batches can be established. Among them, the first channel is used to store the embedding vectors of the first training samples of the corresponding batch, and the second channel is used to store the gradient information of the first training samples of the corresponding batch. At the same time, buffer mechanisms can be designed for the first channel and the second channel, and a preset number of buffers can be configured. For example, the first channel can be configured with 5 buffers, which can store up to 5 embedding vectors at most, and the second channel can be configured with 5 buffers, which can store up to 5 gradient information at most. If the buffers in the first channel and the second channel are full, the principle of first-in first-out can be adopted to discard the embedding vectors or gradients.

[0047] Exemplarily, M = 1, both the first sample set and the second sample set include 600 training samples, and the batch size is 200. Then, the first sample set and the second sample set can be divided into three batches, including batch A, batch B, and batch C. At this time, three first channels and three second channels can be established, that is, the first channel A1 and the second channel A2 corresponding to batch A, the first channel B1 and the second channel B2 corresponding to batch B, and the first channel C1 and the first channel C2 corresponding to batch C. Among them, the first channel A1 is used to store the embedding vectors of the first training samples of batch A of the second participating party, the first channel B1 is used to store the embedding vectors of the first training samples of batch B of the second participating party, the first channel C1 is used to store the embedding vectors of the first training samples of batch C of the second participating party, the second channel A2 is used to store the gradient information of the first training samples of batch A of the second participating party, the second channel B2 is used to store the gradient information of the first training samples of batch B of the second participating party, and the second channel C2 is used to store the gradient information of the first training samples of batch C of the second participating party.

[0048] Exemplarily, M = 2. The first sample set and the second sample sets held by two second participating parties each include 600 training samples, and the batch size is 200. Then, the first sample set and the two second sample sets can be divided into three batches, including batch A, batch B, and batch C. At this time, three first channels and three second channels can be established, namely the first channel A1 and the second channel A2 corresponding to batch A, the first channel B1 and the second channel B2 corresponding to batch B, and the first channel C1 and the first channel C2 corresponding to batch C. Among them, the first channel A1 is used to store the embedding vectors of the first training samples of batch A of the two second participating parties, the first channel B1 is used to store the embedding vectors of the first training samples of batch B of the two second participating parties, the first channel C1 is used to store the embedding vectors of the first training samples of batch C of the two second participating parties, the second channel A2 is used to store the gradient information of the first training samples of batch A of the two second participating parties, the second channel B2 is used to store the gradient information of the first training samples of batch B of the two second participating parties, and the second channel C2 is used to store the gradient information of the first training samples of batch C of the second participating party.

[0049] The second participating party can use the first training samples of the current batch to train the local first feature extraction model, obtain the first embedding vectors of the first training samples of the current batch, and put the first embedding vectors into the first channel corresponding to the current batch. For example, if the current batch is batch B, after the second participating party obtains the first embedding vectors of the first training samples of batch B, it can put them into the first channel B1.

[0050] The first participating party can use the second training samples of the current batch to train the local second feature extraction model, that is, input the second training samples of the current batch into the second feature extraction model to obtain the second embedding vectors of the second training samples of the current batch; then, the first participating party takes out M first embedding vectors from the first channel corresponding to the current batch; next, the first participating party trains the local inference model based on the second embedding vectors and the M first embedding vectors taken out, that is, splices the first embedding vectors and the second embedding vectors and inputs them into the inference model to obtain an output result. Then, according to the output result and the label corresponding to the second training samples of the current batch, calculate the model loss, and calculate the first gradient information, the second gradient information, and the third gradient information based on the model loss, where the first gradient information is the gradient information of the first training samples of the current batch and is used to update the first feature extraction model, the second gradient information is the gradient information of the second training samples of the current batch and is used to update the second feature extraction model, and the third gradient information is the gradient information of the embedding vectors and is used to update the inference model; next, the first participating party puts the first gradient information into the second channel corresponding to the current batch, and uses the second gradient information to update the second feature extraction model and uses the third gradient information to update the inference model.

[0051] For example, if the current batch is Batch C, the first participant can use the second training samples of Batch C to train the local second feature extraction model to obtain the second embedding vectors of the second training samples of Batch C. Then, the first participant takes out M first embedding vectors from the first channel C1 corresponding to Batch C. Next, based on the second embedding vectors and the M first embedding vectors taken out, the first participant trains the local inference model to obtain the first gradient information, the second gradient information, and the third gradient information. Next, the first participant puts the first gradient information into the second channel C2 corresponding to Batch C, and uses the second gradient information to update the second feature extraction model, and uses the third gradient information to update the inference model.

[0052] In S104, it is determined whether all N channels are empty.

[0053] After the second participant puts the first embedding vectors into the corresponding first channels, it can attempt to perform backpropagation of the first feature extraction model. At this time, it can be determined whether all N second channels are empty. If all N second channels are empty, it means that the first participant has not generated new gradient information for the time being. At this time, the backpropagation of the first feature extraction model can be temporarily not performed, but the forward propagation of the first feature extraction model can be continued, that is, the training samples of the next batch are obtained, and the same process is executed, that is, the training samples of the next batch are used as the first training samples of the current batch, and S102 to S106 are continued to be executed until all the N batches are executed and then ended. Among them, when all the N batches are executed, it means that one training round has been completed. In this way, multiple training rounds can be repeated until the model converges. If there are non-empty second channels among the N second channels, it means that the first participant has generated new gradient information. At this time, the backpropagation of the first feature extraction model can be performed, that is, the following S105 and S106 are executed, and then the above S104 can be returned.

[0054] In S105, the earliest stored first gradient information is taken out from the non-empty second channel.

[0055] In S106, the first feature extraction model is updated using the earliest stored first gradient information.

[0056] In the above technical solution, before the M unlabeled second parties and the labeled first party perform model training based on vertical federated learning, they first align the sample identifiers of the data they hold respectively, divide the aligned sample sets into N batches respectively, and establish N first channels and N second channels corresponding one by one to the N batches. Among them, the first channels are used to store the embedding vectors of the first training samples in the corresponding batches, and the second channels are used to store the gradient information of the first training samples in the corresponding batches. When performing batch model training, the second party can use the first training samples in the current batch to train the first feature extraction model to obtain the first embedding vector. After that, the second party, as the publisher, puts the first embedding vector into the first channel corresponding to the current batch. At the same time, the first party can use the second training samples in the current batch to train the second feature extraction model to obtain the second embedding vector. After that, the first party, as the subscriber, takes out the M first embedding vectors from the first channel corresponding to the current batch, and based on the second embedding vector and the M first embedding vectors taken out, trains the inference model to obtain the first gradient information, the second gradient information, and the third gradient information. Next, the first party, as the publisher again, puts the first gradient information into the second channel corresponding to the current batch, and updates the second feature extraction model using the second gradient information and updates the inference model using the third gradient information. After the second party puts the first embedding vector into the corresponding first channel, it can, as the subscriber, determine whether there is first gradient information in the N second channels. If there is a non-empty second channel among the N second channels, it takes out the first gradient information stored earliest from the non-empty second channel and updates the first feature extraction model using the earliest stored first gradient information. If there is no first gradient information in all N second channels, that is, all N second channels are empty, it does not need to wait for the gradient information of the first training samples in the current batch fed back by the first party and can continue to train the first feature extraction model. Thus, through the publish-subscribe architecture and the channel caching mechanism, the training processes of multiple parties are decoupled, and asynchronous training of the model is achieved, thereby eliminating the training delay caused by synchronous dependencies in vertical federated learning and improving the model training efficiency. In addition, the embedding vectors and gradient information of the training samples in different batches will be put into the channels corresponding to the corresponding batches, which can ensure that the sample identifiers are aligned at any time, thereby ensuring the accuracy of model training.

[0057] In addition, to avoid data leakage of the second party caused by the cracking of the embedding information, the first channel can integrate a differential privacy protocol to add noise to the embedding information put into the first channel. At the same time, to avoid data leakage of the first party caused by the cracking of the gradient information, the second channel can also integrate a differential privacy protocol to add noise to the gradient information put into the second channel.

[0058] To further improve the training efficiency of the model, M second participating parties and the first participating party can adopt a Parameter Server (PS) architecture to achieve efficient data parallelism and improve the training efficiency of the model. Specifically, each of the second participating parties may respectively include ω p first worker nodes and a first parameter server, where ω p ≥ 1, and each of the ω p first worker nodes corresponds to at least one batch among N batches. Moreover, each of the ω p first worker nodes is respectively deployed with a first feature extraction model.

[0059] Exemplarily, as Figure 4 shown, M = 1, and the second participating party includes two first worker nodes, namely Worker11 and Worker12, that is, ω p = 2. The first sample set is divided into two batches, namely Batch A and Batch B. The batch identifier (Batch ID) of Batch A is k, and the Batch ID of Batch B is l. Among them, Worker11 corresponds to Batch A (i.e., Batch ID k) and is used to train the first feature extraction model deployed on Worker11 using the first training samples of Batch A; Worker12 corresponds to Batch B (i.e., Batch ID l) and is used to train the first feature extraction model deployed on Worker12 using the first training samples of Batch B.

[0060] As another example, M = 2, that is, the multiple participating parties involved in model training include two second participating parties, namely the second participating party L1 and the second participating party L2. Among them, the second participating party L1 includes two first working nodes, namely Worker11 and Worker12, and the second participating party L2 includes two first working nodes, namely Worker13 and Worker14. The first sample set is divided into five batches, namely batch A, batch B, batch C, batch D, and batch F. Among them, both Worker11 and Worker13 correspond to batch A, batch B, and batch F. Worker11 is used to train the first feature extraction model deployed on Worker11 using the first training samples of these three batches held by the second participating party L1, and Worker13 is used to train the first feature extraction model deployed on Worker13 using the first training samples of these three batches held by the second participating party L2; both Worker12 and Worker14 correspond to batch C and batch D. Worker12 is used to train the first feature extraction model deployed on Worker12 using the first training samples of these two batches held by the second participating party L1, and Worker14 is used to train the first feature extraction model deployed on Worker12 using the first training samples of these two batches held by the second participating party L2.

[0061] Specifically, ω p Each of the ω first working nodes is used to: obtain the first training samples of the current batch, where the current batch is one of the first batches, and the first batches include at least one batch corresponding to the first working node; input the first training samples of the current batch into the first feature extraction model deployed on the first working node to obtain a first embedding vector; put the first embedding vector into the first channel corresponding to the current batch; if there is a non-empty second channel in the second channels corresponding to the first batches, take out the earliest stored first gradient information from the non-empty second channels corresponding to the first batches; update the first feature extraction model deployed on the first working node using the earliest stored first gradient information; if all the second channels corresponding to the first batches are empty, obtain the training samples of the next batch and execute the same process until the first batches are all executed and then end.

[0062] In the present disclosure, when all the second channels corresponding to the first batches are empty, the training samples of the next batch of the current batch in the first batches can be used as the first training samples of the current batch, and the same process can be continued until the first batches are all executed and then end. At this time, when all batches in the first batches are executed, it indicates that the first working node has completed one training round it is responsible for. In this way, the first working node can repeat multiple training rounds until the model converges.

[0063] After updating the first feature extraction model deployed on the first worker node using the earliest stored first gradient information, it is possible to continue to determine whether all the second channels corresponding to the first batch are empty; if all the second channels corresponding to the first batch are empty, obtain the training samples of the next batch and execute the same process until the first batch is completed; if there are non-empty second channels among the second channels corresponding to the first batch, take out the earliest stored first gradient information from the non-empty second channels corresponding to the first batch, and then use the earliest stored first gradient information to update the first feature extraction model deployed on the first worker node.

[0064] Exemplarily, M = 1, as Figure 4 shown, Worker11 corresponds to batch A (i.e., Batch ID k). If the current batch is batch A, then Worker11 is used to: input the first training samples of batch A into the first feature extraction model deployed on Worker11 to obtain the first embedding vector; put the first embedding vector into the first channel (i.e., the embedding channel) corresponding to batch A (i.e., Batch ID k); if the second channel (i.e., the gradient channel) corresponding to Batch ID k is empty, the next batch is still batch A, and the step of obtaining the first training samples of batch A is performed.

[0065] Another exemplarily, the first worker node corresponds to batch A, batch B, and batch F. If the current batch is batch B, then the first worker node is used to input the first training samples of batch B into the first feature extraction model deployed on the first worker node to obtain the first embedding vector; put the first embedding vector into the first channel (i.e., the embedding channel) corresponding to batch B; if the second channels corresponding to batch A, batch B, and batch F are all empty, it is possible to obtain the training samples of the next batch and execute the same process, that is, take batch F as the current batch until the first batch is completed.

[0066] In addition, each of the above ω p first worker nodes can also be used to: send the first model parameters of the updated first feature extraction model to the first parameter server; the first parameter server is used to aggregate the first model parameters sent by the ω p first worker nodes to obtain the second model parameters, and send the second model parameters to each first worker node; each of the ω p first worker nodes is also used to update the model parameters of the first feature extraction model deployed on the first worker node to the second model parameters.

[0067] Exemplarily, as Figure 4As shown, Worker11 corresponds to Batch A (i.e., Batch ID k). If the current batch is Batch A, after Worker11 puts the first embedding vector into the first channel corresponding to Batch A (i.e., Batch ID k), it can also be used for: if the second channel corresponding to Batch ID k is not empty, take out the earliest stored first gradient information from the second channel corresponding to Batch ID k; use the earliest stored first gradient information to update the first feature extraction model deployed on Worker11; send the first model parameters of the updated first feature extraction model to the first parameter server PS1.

[0068] The first parameter server PS1 is used to aggregate the first model parameters sent by Worker11 and Worker12 to obtain second model parameters, and send the second model parameters to Worker11 and Worker12; Worker11 is also used to update the model parameters of the first feature extraction model deployed on Worker11 to the second model parameters; Worker12 is also used to update the model parameters of the first feature extraction model deployed on Worker12 to the second model parameters.

[0069] In addition, to reduce the number of communications, as Figure 4 shown, within the PS architecture, a semi-asynchronous update mechanism can be adopted, that is, the worker nodes synchronously interact with the parameter server every ΔT training rounds to update the local model. Specifically, each of the ω p first worker nodes is used to send the updated first model parameters in the most recent ΔT training rounds to the first parameter server every ΔT training rounds, where ΔT ≥ 1; the first parameter server is used to aggregate the updated first model parameters in the most recent ΔT training rounds sent by the ω p first worker nodes to obtain second model parameters, and send the second model parameters to each first worker node; in this way, each first worker node can update the model parameters of the first feature extraction model deployed on the first worker node to the second model parameters. ΔT can be a preset value.

[0070] When M second parties and a first party perform asynchronous training of a model, it may lead to the complication of model convergence or hinder model convergence. To address this challenge, a dynamic semi-asynchronous update mechanism can be provided within each party to achieve efficient iterative training and thus fast convergence. This mechanism can dynamically adjust the asynchronous update frequency according to the real-time feedback during the training process, balancing the need for faster computation and the stability required for model convergence. To implement the dynamic semi-asynchronous update mechanism in the PS architecture of VFL, the synchronization interval ΔT is designed to decrease as the model accuracy approaches the target accuracy. For example, ΔT has a positive correlation with the number of training rounds. In the initial stage of model training, when the model accuracy is far from the target accuracy, the synchronization interval is large, enabling the model to achieve stable learning. As the model accuracy improves, the synchronization interval decreases and the synchronization frequency increases to fine-tune the model and ensure faster convergence.

[0071] Exemplarily, where ΔT t is the synchronization interval corresponding to the training round t; ΔT0 is the initial synchronization interval, which is a preset value; tanh(·) is the hyperbolic tangent function.

[0072] In the present disclosure, parameters such as the above batch size, the first quantity of the first working nodes in each of the M second parties, and the second quantity of the second working points in the first party can all be preset values. In VFL, there are often significant differences in computing resources and data feature dimensions among parties. There is heterogeneity in data or computing resources among different parties, resulting in inconsistent computing speeds among parties, which can affect the training throughput and convergence time. Among them, the number of the first working nodes in each of the M second parties remains the same.

[0073] To solve this problem, a set of optimal parameters can be determined according to the system configuration files of multiple parties (including model information and hardware capability information), such as the batch size, the first quantity of the first working nodes in the second parties, the second quantity of the second working points in the first party, and the number of central processing unit (CPU) cores allocated to each working point. By determining these parameters according to the specific resources and constraints of multiple parties, the goal is to balance the computing load, reduce latency, improve the overall efficiency of the VFL system in the publish-subscribe architecture, while maintaining privacy compliance, so as to balance the computing speeds of multiple parties while eliminating the training latency caused by synchronization dependencies in VFL, thereby maximizing the utilization rate of computing resources, and further reducing the training cost of VFL and promoting the efficient utilization of computing resources of each party.

[0074] Among them, the VFL system based on the publish-subscribe architecture may include M second parties and a first party, where multiple parties all adopt the PS architecture. As Figure 5 shown, the VFL system process may include a profiling phase, a planning phase, and a training phase (not shown in Figure 5 ).

[0075] In the profiling phase, each party uses the validation dataset (non-private data) and the information of the models (i.e., model information) and hardware capability information held by each party ( Figure 5 taking Party A and Party B in as examples). Specifically, the hardware capability information may include communication bandwidth, memory limit, and the model information includes information such as the size E of the embedding vector, the size G of the gradient information, and the parameter size (i.e., the values of the hyperparameters used later).

[0076] In the planning phase, the task publisher or coordinator can construct an optimization problem according to the system configuration files of multiple parties (i.e., system profiling files), and this optimization problem aims to model the single training duration of multiple parties; if M = 1, then directly according to the single training duration of multiple parties, use the dynamic programming algorithm to solve a set of optimal hyperparameter configurations, and according to the hyperparameter configurations, perform batch sharding of the training samples of multiple parties, worker node allocation, and channel allocation; if M > 1, then first find the second party with the longest single training duration among the M unlabeled second parties (for the sake of distinction, hereinafter referred to as the third party), and then, according to the single training duration of the first party and the single training duration of the third party, use the dynamic programming algorithm to solve a set of optimal hyperparameter configurations, and according to the hyperparameter configurations, perform batch sharding of the training samples of multiple parties, worker node allocation, and channel allocation, where the number of the first working points in each second party is the same. The duration of a single batch training of the second party may include the forward propagation duration of the first feature extraction model of the second party, the forward propagation duration of the first feature extraction model of the second party, and the duration for the second party to send the embedding vector to the first party. When M = 1, this one second party is the third party.

[0077] In the training phase, after obtaining the hyperparameter configuration, the VFL system starts to perform model training.

[0078] Specifically, when the above method is applied to the third party, the above method may further include the following steps (a1) to (a4).

[0079] Step (a1): Construct an objective function with the minimum of the first duration as the optimization goal.

[0080] Among them, the independent variables of the objective function include the batch size, the first quantity of the first working nodes in the second participating party, and the second quantity of the second working points in the first participating party. The first duration is the maximum of the second duration of a single-batch training of the second participating party (i.e., the third participating party) for executing the above method and the third duration of a single-batch training of the first participating party. The second duration is determined based on the batch size and the first quantity, and the third duration is determined based on the batch size and the second quantity.

[0081] Step (a2): Construct the first memory constraint of the third participating party according to the basic memory consumption of the third participating party, and construct the second memory constraint of the first participating party according to the basic memory consumption of the first participating party.

[0082] Among them, the basic memory consumption is the memory size occupied by the participating party to maintain basic functions.

[0083] Exemplarily, the first memory constraint of the third participating party can be constructed according to the basic memory consumption of the third participating party through the following equation (1):

[0084] M P (B) = M P0 + ρ P B χ (1)

[0085] Among them, B is the batch size; M P (B) is the first memory constraint, which is a function of the memory consumption of the third participating party with respect to the batch size B; M P0 is the basic memory consumption of the third participating party, obtained through the profiling phase; ρ P and χ are both hyperparameters.

[0086] Exemplarily, the second memory constraint of the first participating party can be constructed according to the basic memory consumption of the first participating party through the following equation (2):

[0087] M A (B) = M A0 + ρ A B χ (2)

[0088] Among them, M A (B) is the second memory constraint, which is a function of the memory consumption of the first participating party with respect to the batch size B; M A0 is the basic memory consumption of the first participating party, obtained through the profiling phase; ρ A is a hyperparameter.

[0089] Step (a3): Construct the constraint limit of the batch size according to the first memory constraint, the second memory constraint, the basic memory consumption of the third participating party, and the basic memory consumption of the first participating party, as the constraint condition of the objective function.

[0090] Exemplarily, according to the first memory constraint, the second memory constraint, the base memory consumption of the third party, and the base memory consumption of the first party, the constraint limit of the batch size can be constructed through the following equation (3):

[0091]

[0092] where Z is the set of candidate batch sizes; B max is the maximum value of the batch size; there are different M A (B) for different B is the maximum value of M A (B); there are different M P (B) for different B is the maximum value of M P (B).

[0093] Step (a4): Solve the objective function under the constraint conditions through the dynamic programming algorithm to obtain the optimal solution of the independent variable.

[0094] In the present disclosure, the second duration of the single-batch training of the third party may include the duration of the forward propagation of the first feature extraction model of the third party, the duration of the forward propagation of the first feature extraction model of the third party, and the duration for the third party to send the embedding vector to the first party.

[0095] Exemplarily, the duration of the forward propagation of the first feature extraction model of the third party can be calculated through the following equation (4)

[0096]

[0097] where λ p , γ p are both hyperparameters; c p,j is the number of CPU cores assigned to the j-th first working node of the third party, j = 1, 2,..., ω p , P1 ≤ ω p ≤ Q1, P1 is the lower limit value of ω p , Q1 is the upper limit value of ω p , and C p is the total number of CPU cores of the third party.

[0098] Exemplarily, the duration of the backward propagation of the first feature extraction model of the third party can be calculated through the following equation (5)

[0099]

[0100] Among them, β p are all hyperparameters.

[0101] Exemplarily, the duration T for the third party to send the embedding vector to the first party can be calculated by the following equation (6) emb :

[0102]

[0103] Among them, B b is the network bandwidth between the third party and the first party.

[0104] It should be noted that the calculation method of the duration of a single training of the other second parties among the M second parties except the third party is similar to the calculation method of the duration of a single training of the third party, and the details are not described in this disclosure.

[0105] The third duration of the single-batch training of the first party may include the forward propagation duration of the second feature extraction model of the first party, the backward propagation duration of the second feature extraction model of the first party, the total forward and backward propagation duration of the inference model of the first party, and the duration for the first party to send gradient information to the third party.

[0106] Exemplarily, the forward propagation duration of the second feature extraction model of the first party can be calculated by the following equation (7)

[0107]

[0108] Among them, λ a and γ a are all hyperparameters; c a,i is the number of CPU cores allocated to the i-th second working node of the first party, i = 1, 2,..., ω a , ω a is the number of second working nodes in the first party, and P ≤ ω a ≤ Q, P is the lower limit value of ω a and Q is the upper limit value of ω a and C a is the total number of CPU cores of the first party.

[0109] Exemplarily, the backward propagation duration of the second feature extraction model of the first party can be calculated by the following equation (8)

[0110]

[0111] Among them, β a are all hyperparameters.

[0112] Exemplarily, the total duration of the forward propagation and the backward propagation of the inference model of the first participant can be calculated by the following equation (9)

[0113]

[0114] where λ a ', β a ', γ' a are all hyperparameters.

[0115] Exemplarily, the duration T for the first participant to send gradient information to the third participant can be calculated by the following equation (10) grad :

[0116]

[0117] Thus, the objective function is

[0118] Next, the specific implementation manner of solving the objective function under the constraint conditions through the dynamic programming algorithm in the above step (a4) to obtain the optimal solution of the independent variable will be described in detail. Specifically, it can be achieved through the following steps (a41) to (a46):

[0119] Step (a41): According to the constraint limit of the batch size, determine the candidate values of the batch size, and according to the value range of the first quantity of the first worker node, that is, P1 ≤ ω p ≤ Q1, determine the candidate values of the first quantity, and according to the value range of the second quantity of the second worker node, that is, P ≤ ω a ≤ Q, determine the candidate values of the second quantity.

[0120] Step (a42): Define the dynamic programming array dp[i][j][k], where i is the index of the candidate value of the batch size, j is the index of the candidate value of the first quantity, k is the index of the candidate value of the second quantity, and dp[i][j][k] represents the optimal solution of the objective function when the batch size takes the i-th candidate value, the first quantity takes the j-th candidate value, and the second quantity takes the k-th candidate value.

[0121] Step (a43): Define the state transition equation according to the objective function:

[0122]

[0123] Step (a44): Set the initial value of the dynamic programming table to infinity.

[0124] Step (a45): Fill the dynamic programming table in a preset order (usually from bottom to top). For each state dp[i][j][k], calculate its value according to the state transition equation. If the calculated value is less than the value in the dynamic programming table, update the dynamic programming table with the calculated value; otherwise, continue to fill the dynamic programming table in the preset order until the dynamic programming table is filled up.

[0125] Step (a46): Determine the latest value of the dynamic programming table as the global optimal solution of the objective function. Thus, the specific values of the optimal parameter combination can be obtained.

[0126] To reduce the occurrence of model training errors, as Figure 4 shown, a waiting deadline mechanism can be set for each of the M unlabeled second parties. Among them, T ddl is the waiting deadline duration of the second party, that is, the first preset duration. Specifically, the above method for model training based on vertical federated learning applied to any second party may further include the following steps:

[0127] After the first preset duration after putting the first embedding vector into the first channel, determine whether there is the first gradient information of the first training sample of the current batch in the second channel corresponding to the current batch;

[0128] If there is no first gradient information of the first training sample of the current batch in the second channel corresponding to the current batch, send a first instruction to the first party.

[0129] Among them, the first instruction is used to instruct the first party to regenerate the first gradient information of the first training sample of the current batch. In response to receiving the first instruction, the first party can regenerate the first gradient information of the first training sample of the current batch. The first preset duration can be a preset value.

[0130] To further reduce the occurrence of model training errors, in addition to setting a waiting deadline mechanism at the second participating party, a waiting deadline mechanism can also be set at the first participating party. Specifically, after a second preset duration from obtaining the second embedding vector, the first participating party determines whether the number of first embedding vectors in the first channel corresponding to the current batch reaches M, that is, determines whether all M second participating parties have put the first embedding vectors of the training samples of the current batch they hold into the first channel; if the number of first embedding vectors in the first channel corresponding to the current batch does not reach M, a second instruction is sent to the fourth participating party, where the second instruction is used to instruct the fourth participating party to regenerate the first embedding vector, and the fourth participating party includes the second participating parties among the M unlabeled second participating parties that have not put the first embedding vectors of the training samples of the current batch they hold into the first channel. At this time, the above method for model training based on vertical federated learning applied to any second participating party may further include the following steps:

[0131] In response to receiving the second instruction sent by the first participating party, re - execute the steps from inputting the first training sample into the first feature extraction model to obtain the first embedding vector to putting the first embedding vector into the first channel corresponding to the current batch.

[0132] After multiple participating parties complete model training based on VFL, the multiple participating parties can perform joint inference based on their respective sub - models. Exemplarily, M = 1, and the multiple participating parties include a first participating party and a second participating party. Among them, the first participating party can complete joint inference based on the trained second feature extraction model and the inference model, and the second participating party is based on the trained first feature extraction model.

[0133] Specifically, each of the M second participating parties can input its own features into the locally trained first feature extraction model to obtain a third embedding vector, and send the third embedding vector to the first participating party; the first participating party inputs its own features into the local second feature extraction model to obtain a fourth embedding vector. Then, the third embedding vectors and the fourth embedding vectors sent by each second participating party are concatenated to obtain a concatenated vector, and the concatenated vector is input into the locally trained inference model to obtain an inference result.

[0134] Figure 6 is a flowchart of a method for model training based on vertical federated learning applied to a labeled first participating party shown according to an exemplary embodiment. As Figure 6 shown, the method for model training based on vertical federated learning applied to a labeled first participating party may include the following S201 - S205.

[0135] In S201, obtain the second training samples of the current batch.

[0136] Among them, the second training sample of the current batch is one of the N batches.

[0137] In S202, the second training sample is input into the second feature extraction model to obtain a second embedding vector.

[0138] In S203, M first embedding vectors are taken from the first channel corresponding to the current batch. The M first embedding vectors are obtained by each of the M second participating parties without labels inputting the first training sample of the current batch into the first feature extraction model and putting them into the first channel corresponding to the current batch. The N batches are in one-to-one correspondence with the N first channels, and the N batches are in one-to-one correspondence with the N second channels.

[0139] In S204, based on the second embedding vector and the M first embedding vectors, the inference model is trained to obtain first gradient information, second gradient information, and third gradient information.

[0140] In S205, the first gradient information is put into the second channel corresponding to the current batch, and the second feature extraction model is updated using the second gradient information, and the inference model is updated using the third gradient information.

[0141] In the above technical solution, before the M unlabeled second participating parties and the labeled first participating party perform model training based on vertical federated learning, they first align the sample identifiers of the data they hold respectively, divide the aligned sample sets into N batches respectively, and establish N first channels and N second channels corresponding one-to-one to the N batches. Among them, the first channel is used to store the embedding vectors of the first training samples in the corresponding batch, and the second channel is used to store the gradient information of the first training samples in the corresponding batch. When performing batch model training, the second participating party can use the first training samples in the current batch to train the first feature extraction model to obtain the first embedding vector. After that, the second participating party, as the publisher, puts the first embedding vector into the first channel corresponding to the current batch. At the same time, the first participating party can use the second training samples in the current batch to train the second feature extraction model to obtain the second embedding vector. After that, the first participating party, as the subscriber, takes out M first embedding vectors from the first channel corresponding to the current batch, and based on the second embedding vector and the M first embedding vectors taken out, trains the inference model to obtain the first gradient information, the second gradient information, and the third gradient information. Next, the first participating party, as the publisher again, puts the first gradient information into the second channel corresponding to the current batch, and uses the second gradient information to update the second feature extraction model and the third gradient information to update the inference model. After the second participating party puts the first embedding vector into the corresponding first channel, it can, as the subscriber, determine whether there is first gradient information in the N second channels. If there is a non-empty second channel among the N second channels, it takes out the earliest stored first gradient information from the non-empty second channel and uses the earliest stored first gradient information to update the first feature extraction model. If there is no first gradient information in all N second channels, that is, all N second channels are empty, it does not need to wait for the gradient information of the first training samples in the current batch fed back by the first participating party and can continue to train the first feature extraction model. Thus, through the publish-subscribe architecture and the channel caching mechanism, the training processes of multiple participating parties are decoupled, and asynchronous training of the model is realized, thereby eliminating the training delay caused by synchronous dependencies in vertical federated learning and improving the model training efficiency. In addition, the embedding vectors and gradient information of the training samples in different batches are put into the channels corresponding to the corresponding batches, which can ensure that the sample identifiers are aligned at any time, thereby ensuring the accuracy of model training.

[0142] Optionally, the first participating party includes ω a second working nodes, ω a ≥1, and each of the ω a second working nodes is corresponding to at least one batch among the N batches; ω aEach of the second working nodes is used to: obtain the second training samples of the current batch, where the current batch is one of the second batches, and the second batch includes at least one batch corresponding to the second working node; input the second training samples into the second feature extraction model deployed on the second working node to obtain second embedding vectors; retrieve M first embedding vectors from the first channel corresponding to the current batch; based on the second embedding vectors and the M first embedding vectors, train the inference model deployed on the second working node to obtain first gradient information, second gradient information, and third gradient information; place the first gradient information into the second channel corresponding to the current batch, and use the second gradient information to update the second feature extraction model deployed on the second working node, and use the third gradient information to update the inference model deployed on the second working node.

[0143] Exemplarily, as Figure 4 shown, the first party includes three first working nodes, namely Worker21, Worker22, and Worker23, i.e., ω a = 3, the second sample set is divided into two batches, namely batch A and batch B. The batch identifier (Batch ID) of batch A is k, and the Batch ID of batch B is l. Among them, both Worker22 and Worker23 correspond to batch A (i.e., Batch ID k) and are used to train the second feature extraction model and the inference model deployed on themselves using the second training samples of batch A; Worker21 corresponds to batch B (i.e., Batch ID l) and is used to train the second feature extraction model and the inference model deployed on Worker21 using the second training samples of batch B.

[0144] Optionally, the first party further includes a second parameter server; ω a Each of the second working nodes is further used to send the third model parameters of the updated second feature extraction model and the fourth model parameters of the updated inference model to the second parameter server; the second parameter server is further used to aggregate the third model parameters sent by ω a the second working nodes to obtain fifth model parameters, aggregate the fourth model parameters sent by ω a the second working nodes to obtain sixth model parameters, and send the fifth model parameters and the sixth model parameters to each second working node; ω a Each of the second working nodes is further used to update the model parameters of the second feature extraction model deployed on the second working node to the fifth model parameters, and update the model parameters of the inference model deployed on the second working node to the sixth model parameters.

[0145] Optionally, ωa Each of the second worker nodes is used to send the updated third model parameters and the updated fourth model parameters in the most recent ΔT training rounds to the second parameter server every ΔT training rounds, where ΔT ≥ 1; the second parameter server is used to aggregate the updated third model parameters in the most recent ΔT training rounds sent by ω a second worker nodes to obtain the fifth model parameters, and aggregate the updated fourth model parameters in the most recent ΔT training rounds sent by ω a second worker nodes to obtain the sixth model parameters, and send the fifth model parameters and the sixth model parameters to each second worker node.

[0146] Optionally, ΔT has a positive correlation with the number of training rounds.

[0147] Optionally, the model training method based on vertical federated learning applied to the labeled first party further includes: constructing an objective function with the minimum first duration as the optimization target, where the independent variables of the objective function include the batch size, the first number of the first working points in the second party, and the second number of the second worker nodes, the first duration is the maximum of the second duration of a single batch training of the third party and the third duration of a single batch training of the first party, the third party is the second party with the longest single batch training duration among the M unlabeled second parties, the second duration is determined based on the batch size and the first number, and the third duration is determined based on the batch size and the second number; constructing a first memory constraint for the third party according to the basic memory consumption of the third party, and constructing a second memory constraint for the first party according to the basic memory consumption of the first party, where the basic memory consumption is the memory size occupied by the party to maintain basic functions; constructing a constraint limit for the batch size according to the first memory constraint, the second memory constraint, the basic memory consumption of the third party, and the basic memory consumption of the first party, as a constraint condition of the objective function; and solving the objective function under the constraint condition through a dynamic programming algorithm to obtain the optimal solution of the independent variables.

[0148] Optionally, the model training method based on vertical federated learning applied to the labeled first party further includes: after a second preset duration after obtaining the second embedding vector, determining whether the number of the second embedding vectors in the first channel corresponding to the current batch reaches M; if the number of the second embedding vectors in the first channel corresponding to the current batch does not reach M, sending a second instruction to the fourth party, where the second instruction is used to instruct the fourth party to regenerate the second embedding vector, and the fourth party includes the second parties among the M unlabeled second parties that have not put the second embedding vector into the first channel.

[0149] Optionally, the method for training a model based on vertical federated learning applied to the labeled first participant further includes: in response to receiving a first instruction, re-executing the steps of inputting the second training sample into the second feature extraction model to obtain a second embedding vector to putting the first gradient information into a second channel corresponding to the current batch, and using the second gradient information to update the second feature extraction model, and using the third gradient information to update the inference model, where the first instruction is used to instruct the first participant to regenerate the first gradient information of the first training sample of the current batch.

[0150] The specific implementation manners of the steps in the method for training a model based on vertical federated learning applied to the labeled first participant according to the embodiments of the present disclosure have been described in detail in the method for training a model based on vertical federated learning applied to any unlabeled second participant according to the embodiments of the present disclosure, and will not be repeated here.

[0151] Figure 7 It is a block diagram of a model training device based on vertical federated learning applied to any unlabeled second participant shown according to an exemplary embodiment. The multiple participants participating in the model training include a labeled first participant and M unlabeled second participants, where M≥1. The data held by the multiple participants are respectively divided into N batches after sample identification alignment, where N≥1. The N batches correspond one-to-one to N first channels and the N batches correspond one-to-one to N second channels. As Figure 7As shown, the model training device 300 based on vertical federated learning is applied to any one of the M second participating parties without labels. The device 300 includes: a first acquisition module 301, configured to acquire a first training sample of the current batch, where the first training sample of the current batch is one of the N batches; a first input module 302, configured to input the first training sample into a first feature extraction model to obtain a first embedding vector; a storage module 303, configured to put the first embedding vector into a first channel corresponding to the current batch; a first extraction module 304, configured to, if there is a non-empty second channel among the N second channels, extract the earliest stored first gradient information from the non-empty second channel; and update the first feature extraction model by using the earliest stored first gradient information; where the second channel is used to store the first gradient information of the first training sample of the corresponding batch, and the first gradient information is obtained by the first participating party training an inference model based on a second embedding vector and M first embedding vectors extracted from the first channel corresponding to the corresponding batch, and the second embedding vector is obtained by the first participating party inputting the second training sample of the corresponding batch into a second feature extraction model; a first trigger module 305, configured to, if all the N second channels are empty, acquire the training sample of the next batch and execute the same process until all the N batches are executed and then end.

[0152] In the above technical solution, before the M unlabeled second parties and the labeled first party perform model training based on vertical federated learning, they first align the sample identifiers of the data they hold respectively, and divide the aligned sample sets into N batches respectively, and establish N first channels and N second channels corresponding to the N batches one by one. Among them, the first channel is used to store the embedding vectors of the first training samples in the corresponding batch, and the second channel is used to store the gradient information of the first training samples in the corresponding batch. When performing batch model training, the second party can use the first training samples in the current batch to train the first feature extraction model to obtain the first embedding vector. Then, the second party, as the publisher, puts the first embedding vector into the first channel corresponding to the current batch. At the same time, the first party can use the second training samples in the current batch to train the second feature extraction model to obtain the second embedding vector. Then, the first party, as the subscriber, takes out M first embedding vectors from the first channel corresponding to the current batch, and based on the second embedding vector and the M first embedding vectors taken out, trains the inference model to obtain the first gradient information, the second gradient information, and the third gradient information. Next, the first party, as the publisher again, puts the first gradient information into the second channel corresponding to the current batch, and uses the second gradient information to update the second feature extraction model and the third gradient information to update the inference model. After the second party puts the first embedding vector into the corresponding first channel, it can, as the subscriber, determine whether there is first gradient information in the N second channels. If there is a non-empty second channel among the N second channels, it takes out the earliest stored first gradient information from the non-empty second channel and uses the earliest stored first gradient information to update the first feature extraction model. If there is no first gradient information in all N second channels, that is, all N second channels are empty, it does not need to wait for the gradient information of the first training samples in the current batch fed back by the first party and can continue to train the first feature extraction model. Thus, through the publish-subscribe architecture and the channel caching mechanism, the training processes of multiple parties are decoupled, and asynchronous training of the model is realized, thereby eliminating the training delay caused by synchronous dependence in vertical federated learning and improving the model training efficiency. In addition, the embedding vectors and gradient information of the training samples in different batches are put into the channels corresponding to the corresponding batches, which can ensure that the sample identifiers are aligned at any time, thereby ensuring the accuracy of model training.

[0153] Optionally, the second party includes ω p first working nodes, ω p ≥1, and each of the ω p first working nodes corresponds to at least one of the N batches; the ω pEach of the first working nodes is used to: obtain the first training samples of the current batch, where the current batch is one of the first batches, and the first batches include the at least one batch corresponding to the first working node; input the first training samples into the first feature extraction model deployed on the first working node to obtain the first embedding vector; put the first embedding vector into the first channel corresponding to the current batch; if there is a non-empty second channel in the second channels corresponding to the first batches, take out the earliest stored first gradient information from the non-empty second channels corresponding to the first batches; update the first feature extraction model deployed on the first working node by using the earliest stored first gradient information; if all the second channels corresponding to the first batches are empty, obtain the training samples of the next batch and execute the same process until the first batches are completed and then end.

[0154] Optionally, the second party further includes a first parameter server; the ω p Each of the first working nodes is further used to: send the first model parameters of the updated first feature extraction model to the first parameter server; the first parameter server is used to aggregate the first model parameters sent by the ω p first working nodes to obtain second model parameters, and send the second model parameters to each of the first working nodes; the ω p Each of the first working nodes is further used to update the model parameters of the first feature extraction model deployed on the first working node to the second model parameters.

[0155] Optionally, the ω p Each of the first working nodes is used to send the updated first model parameters in the most recent ΔT training rounds to the first parameter server every ΔT training rounds, where ΔT≥1; the first parameter server is used to aggregate the updated first model parameters in the most recent ΔT training rounds sent by the ω p first working nodes to obtain second model parameters, and send the second model parameters to each of the first working nodes.

[0156] Optionally, ΔT has a positive correlation with the number of training rounds.

[0157] Optionally, when the model training device 300 based on vertical federated learning is applied to a third party, the device 300 further includes: a first construction module, configured to construct an objective function with the minimum first duration as the optimization target, where the third party is the second party with the longest single-batch training duration among the M unlabeled second parties, the independent variables of the objective function include the batch size, the first number of first working nodes among the second parties, and the second number of second working points among the first parties, the first duration is the maximum of the second duration of the third party's single-batch training and the third duration of the first party's single-batch training, the second duration is determined based on the batch size and the first number, and the third duration is determined based on the batch size and the second number; a second construction module, configured to construct a first memory constraint for the third party according to the basic memory consumption of the third party, and construct a second memory constraint for the first party according to the basic memory consumption of the first party, where the basic memory consumption is the memory size occupied by the party to maintain basic functions; a third construction module, configured to construct a constraint limit for the batch size according to the first memory constraint, the second memory constraint, the basic memory consumption of the third party, and the basic memory consumption of the first party, as a constraint condition of the objective function; a first solution module, configured to solve the objective function under the constraint condition by a dynamic programming algorithm to obtain an optimal solution of the independent variables.

[0158] Optionally, the model training device 300 based on vertical federated learning applied to any unlabeled second party further includes: a first determination module, configured to determine whether there is first gradient information of the first training sample of the current batch in the second channel corresponding to the current batch after a first preset duration from putting the first embedding vector into the first channel; a first sending module, configured to send a first instruction to the first party if there is no first gradient information of the first training sample of the current batch in the second channel corresponding to the current batch, where the first instruction is used to instruct the first party to regenerate the first gradient information of the first training sample of the current batch.

[0159] Optionally, the model training device 300 based on vertical federated learning applied to any unlabeled second party further includes: a second trigger module, configured to re-execute the step of inputting the first training sample into the first feature extraction model to obtain a first embedding vector to the step of putting the first embedding vector into the first channel corresponding to the current batch in response to receiving a second instruction sent by the first party, where the second instruction is used to instruct the second party to regenerate the first embedding vector.

[0160] Figure 8 It is a block diagram of a model training device based on vertical federated learning applied to a labeled first participant shown according to an exemplary embodiment. The multiple participants participating in model training include a labeled first participant and M unlabeled second participants, where M≥1. The data held by each of the multiple participants is divided into N batches after sample identification alignment, where N≥1. The N batches correspond one-to-one to N first channels and the N batches correspond one-to-one to N second channels; as Figure 8 shown, the model training device 400 based on vertical federated learning applied to the labeled first participant includes: a second acquisition module 401, configured to acquire a second training sample of a current batch, where the second training sample of the current batch is one of the N batches; a second input module 402, configured to input the second training sample into a second feature extraction model to obtain a second embedding vector; a second extraction module 403, configured to extract M first embedding vectors from a first channel corresponding to the current batch; where the M first embedding vectors are obtained by each of the M unlabeled second participants inputting the first training sample of the current batch into a local first feature extraction model and putting them into the first channel corresponding to the current batch; a training module 404, configured to train an inference model based on the second embedding vector and the M first embedding vectors to obtain first gradient information, second gradient information, and third gradient information; a first update module 405, configured to put the first gradient information into the second channel corresponding to the current batch, and update the second feature extraction model using the second gradient information, and update the inference model using the third gradient information.

[0161] In the above technical solution, before the M unlabeled second parties and the labeled first party perform model training based on vertical federated learning, they first align the sample identifiers of the data they hold, divide the aligned sample sets into N batches respectively, and establish N first channels and N second channels corresponding to the N batches one by one. Among them, the first channel is used to store the embedding vectors of the first training samples in the corresponding batch, and the second channel is used to store the gradient information of the first training samples in the corresponding batch. When performing batch model training, the second party can use the first training samples in the current batch to train the first feature extraction model to obtain the first embedding vector. After that, the second party, as the publisher, puts the first embedding vector into the first channel corresponding to the current batch. At the same time, the first party can use the second training samples in the current batch to train the second feature extraction model to obtain the second embedding vector. After that, the first party, as the subscriber, takes out the M first embedding vectors from the first channel corresponding to the current batch, and based on the second embedding vector and the M first embedding vectors taken out, trains the inference model to obtain the first gradient information, the second gradient information, and the third gradient information. Next, the first party, as the publisher again, puts the first gradient information into the second channel corresponding to the current batch, and uses the second gradient information to update the second feature extraction model and the third gradient information to update the inference model. After the second party puts the first embedding vector into the corresponding first channel, it can, as the subscriber, determine whether there is first gradient information in the N second channels. If there is a non-empty second channel among the N second channels, it takes out the earliest stored first gradient information from the non-empty second channel and uses the earliest stored first gradient information to update the first feature extraction model. If there is no first gradient information in all N second channels, that is, all N second channels are empty, it does not need to wait for the gradient information of the first training samples in the current batch fed back by the first party and can continue to train the first feature extraction model. Thus, through the publish-subscribe architecture and the channel caching mechanism, the training processes of multiple parties are decoupled, and asynchronous training of the model is realized, thereby eliminating the training delay caused by synchronous dependence in vertical federated learning and improving the model training efficiency. In addition, the embedding vectors and gradient information of the training samples in different batches are put into the channels corresponding to the corresponding batches, which can ensure that the sample identifiers are aligned at any time, thereby ensuring the accuracy of model training.

[0162] Optionally, the first party includes ω a second working nodes, ω a ≥1, and each of the ω a second working nodes is corresponding to at least one batch among the N batches; the ω aEach of the ω second working nodes is used to: obtain the second training samples of the current batch, where the current batch is one of the second batches, and the second batches include the at least one batch corresponding to the second working node; input the second training samples into the second feature extraction model deployed on the second working node to obtain the second embedding vectors; retrieve the M first embedding vectors from the first channel corresponding to the current batch; based on the second embedding vectors and the M first embedding vectors, train the inference model deployed on the second working node to obtain the first gradient information, the second gradient information, and the third gradient information; place the first gradient information into the second channel corresponding to the current batch, and use the second gradient information to update the second feature extraction model deployed on the second working node, and use the third gradient information to update the inference model deployed on the second working node.

[0163] Optionally, the first party further includes a second parameter server; the ω a Each of the ω second working nodes is further used to send the third model parameters of the updated second feature extraction model and the fourth model parameters of the updated inference model to the second parameter server; the second parameter server is further used to aggregate the third model parameters sent by the ω a second working nodes to obtain fifth model parameters, aggregate the fourth model parameters sent by the ω a second working nodes to obtain sixth model parameters, and send the fifth model parameters and the sixth model parameters to each of the second working nodes; the ω a Each of the ω second working nodes is further used to update the model parameters of the second feature extraction model deployed on the second working node to the fifth model parameters, and update the model parameters of the inference model deployed on the second working node to the sixth model parameters.

[0164] Optionally, each of the ω a second working nodes is used to send the updated third model parameters and the updated fourth model parameters in the most recent ΔT training rounds to the second parameter server every ΔT training rounds, where ΔT ≥ 1; the second parameter server is used to aggregate the updated third model parameters in the most recent ΔT training rounds sent by the ω a second working nodes to obtain fifth model parameters, and for the ω aAggregate the updated fourth model parameters in the most recent ΔT training rounds sent by the second worker nodes to obtain sixth model parameters, and send the fifth model parameters and the sixth model parameters to each of the second worker nodes.

[0165] Optionally, ΔT has a positive correlation with the number of training rounds.

[0166] Optionally, the model training device 400 for longitudinal federated learning applied to the first participating party with labels further includes: a fourth construction module, configured to construct an objective function with the minimum first duration as the optimization target, where the independent variables of the objective function include the batch size, the first number of the first working points in the second participating party, and the second number of the second worker nodes, the first duration is the maximum of the second duration of a single-batch training of the third participating party and the third duration of a single-batch training of the first participating party, the third participating party is the second participating party with the longest single-batch training duration among the M second participating parties without labels, the second duration is determined based on the batch size and the first number, and the third duration is determined based on the batch size and the second number; a fifth construction module, configured to construct a first memory constraint for the third participating party according to the basic memory consumption of the third participating party, and construct a second memory constraint for the first participating party according to the basic memory consumption of the first participating party, where the basic memory consumption is the memory size occupied by the participating party to maintain basic functions; a sixth construction module, configured to construct a constraint limit for the batch size according to the first memory constraint, the second memory constraint, the basic memory consumption of the third participating party, and the basic memory consumption of the first participating party, as a constraint condition of the objective function; a second solution module, configured to solve the objective function under the constraint condition through a dynamic programming algorithm to obtain the optimal solution of the independent variables.

[0167] Optionally, the model training device 400 for longitudinal federated learning applied to the first participating party with labels further includes: a second determination module, configured to determine whether the number of the first embedding vectors in the first channel corresponding to the current batch reaches M after a second preset duration from obtaining the second embedding vector; a second sending module, configured to, if the number of the first embedding vectors in the first channel corresponding to the current batch does not reach M, send a second instruction to the fourth participating party, where the second instruction is used to instruct the fourth participating party to regenerate the first embedding vector, and the fourth participating party includes the second participating parties among the M second participating parties without labels that have not put the first embedding vector into the first channel.

[0168] Optionally, the model training apparatus 400 for longitudinal federated learning applied to the labeled first participating party further includes: a third trigger module, configured to, in response to receiving a first instruction, re-execute the step of inputting the second training sample into the second feature extraction model to obtain a second embedding vector to the step of placing the first gradient information into the second channel corresponding to the current batch, and using the second gradient information to update the second feature extraction model, and using the third gradient information to update the inference model, where the first instruction is used to instruct the first participating party to regenerate the first gradient information of the first training sample of the current batch.

[0169] In addition, the present disclosure also provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processing device, the steps of the above-mentioned model training method for longitudinal federated learning applied to any unlabeled second participating party are implemented, or the steps of the above-mentioned model training method for longitudinal federated learning applied to a labeled first participating party are implemented.

[0170] The present disclosure also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned model training method for longitudinal federated learning applied to any unlabeled second participating party are implemented, or the steps of the above-mentioned model training method for longitudinal federated learning applied to a labeled first participating party are implemented.

[0171] Next, refer to Figure 9 , which shows a schematic structural diagram of an electronic device (such as a terminal device or a server) 600 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0172] As Figure 9As shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in the read-only memory (ROM) 602 or a program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0173] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 9 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0174] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0175] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0176] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (for example, a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (for example, the Internet), and end-to-end networks (for example, ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0177] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately and not be assembled into the electronic device.

[0178] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain a first training sample of the current batch, where the first training sample of the current batch is one of N batches, and the multiple parties participating in the model training include a first party with labels and M second parties without labels, where M≥1. The data held by the multiple parties are respectively divided into N batches after sample identification alignment, N≥1, and the N batches correspond one-to-one to N first channels and also correspond one-to-one to N second channels; input the first training sample into a first feature extraction model to obtain a first embedding vector; place the first embedding vector into the first channel corresponding to the current batch; if there is a non-empty second channel among the N second channels, take out the earliest stored first gradient information from the non-empty second channel; update the first feature extraction model using the earliest stored first gradient information; where the second channel is used to store the first gradient information of the first training sample of the corresponding batch, and the first gradient information is obtained by the first party training an inference model based on a second embedding vector and M first embedding vectors taken from the first channel corresponding to the corresponding batch, and the second embedding vector is obtained by the first party inputting the second training sample of the corresponding batch into a second feature extraction model; if all N second channels are empty, obtain the training sample of the next batch and execute the same process until all N batches are completed and then end.

[0179] Alternatively, the above computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: obtain the second training samples of the current batch, where the second training samples of the current batch are one of N batches, and the multiple participants participating in the model training include a first participant with labels and M second participants without labels, where M ≥ 1. The data held by the multiple participants are respectively divided into N batches after sample identification alignment, N ≥ 1, and the N batches correspond one-to-one to N first channels and also correspond one-to-one to N second channels; input the second training samples into a second feature extraction model to obtain second embedding vectors; retrieve M first embedding vectors from the first channel corresponding to the current batch; where the M first embedding vectors are obtained by each of the M second participants without labels inputting the first training samples of the current batch into a local first feature extraction model and putting them into the first channel corresponding to the current batch; based on the second embedding vectors and the M first embedding vectors, train an inference model to obtain first gradient information, second gradient information, and third gradient information; put the first gradient information into the second channel corresponding to the current batch, and use the second gradient information to update the second feature extraction model, and use the third gradient information to update the inference model.

[0180] Computer program code for carrying out operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0181] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0182] The modules described in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the module itself in some cases. For example, the first acquisition module can also be described as "the module for acquiring the first training samples of the current batch".

[0183] The functions described above herein can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.

[0184] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0185] According to one or more embodiments of the present disclosure, Example 1 provides a model training method based on vertical federated learning. The multiple parties participating in the model training include a labeled first party and M unlabeled second parties, where M≥1. The data held by each of the multiple parties is divided into N batches after sample identification alignment, where N≥1. The N batches correspond one-to-one to N first channels and also correspond one-to-one to N second channels. The method is applied to any one of the M unlabeled second parties, and the method includes: obtaining a first training sample of the current batch, where the first training sample of the current batch is one of the N batches; inputting the first training sample into a first feature extraction model to obtain a first embedding vector; putting the first embedding vector into the first channel corresponding to the current batch; if there is a non-empty second channel among the N second channels, taking out the earliest stored first gradient information from the non-empty second channel; updating the first feature extraction model using the earliest stored first gradient information; where the second channel is used to store the first gradient information of the first training sample of the corresponding batch, and the first gradient information is obtained by the first party training an inference model based on a second embedding vector and M first embedding vectors taken out from the first channel corresponding to the corresponding batch, and the second embedding vector is obtained by the first party inputting the second training sample of the corresponding batch into a second feature extraction model; if all N second channels are empty, obtaining the training sample of the next batch and performing the same process until all N batches are completed and then ending.

[0186] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, where the second party includes ω p first working nodes, ω p ≥1, and each of the ω p first working nodes corresponds to at least one of the N batches; the ω pEach of the first working nodes is used to: obtain the first training samples of the current batch, where the current batch is one of the first batches, and the first batches include the at least one batch corresponding to the first working node; input the first training samples into the first feature extraction model deployed on the first working node to obtain the first embedding vector; put the first embedding vector into the first channel corresponding to the current batch; if there is a non-empty second channel among the second channels corresponding to the first batches, take out the earliest stored first gradient information from the non-empty second channels corresponding to the first batches; update the first feature extraction model deployed on the first working node by using the earliest stored first gradient information; if all the second channels corresponding to the first batches are empty, obtain the training samples of the next batch and execute the same process until the first batches are completed.

[0187] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 2, and the second participating party further includes a first parameter server; the ω p Each of the first working nodes is further used to: send the first model parameters of the updated first feature extraction model to the first parameter server; the first parameter server is used to aggregate the first model parameters sent by the ω p first working nodes to obtain second model parameters, and send the second model parameters to each of the first working nodes; the ω p Each of the first working nodes is further used to update the model parameters of the first feature extraction model deployed on the first working node to the second model parameters.

[0188] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 3, the ω p Each of the first working nodes is used to send the updated first model parameters in the most recent ΔT training rounds to the first parameter server every ΔT training rounds, where ΔT≥1; the first parameter server is used to aggregate the updated first model parameters in the most recent ΔT training rounds sent by the ω p first working nodes to obtain second model parameters, and send the second model parameters to each of the first working nodes.

[0189] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 4, and ΔT has a positive correlation with the number of training rounds.

[0190] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 2. When the method is applied to a third party, the method further includes: constructing an objective function with the minimum of the first duration as the optimization goal, where the third party is the second party with the longest single-batch training duration among the M unlabeled second parties, the independent variables of the objective function include the batch size, the first number of the first working nodes among the second parties, and the second number of the second working points among the first parties, the first duration is the maximum value between the second duration of the third party's single-batch training and the third duration of the first party's single-batch training, the second duration is determined based on the batch size and the first number, and the third duration is determined based on the batch size and the second number; constructing a first memory constraint for the third party according to the basic memory consumption of the third party, and constructing a second memory constraint for the first party according to the basic memory consumption of the first party, where the basic memory consumption is the memory size occupied by the party to maintain basic functions; constructing a constraint limit for the batch size according to the first memory constraint, the second memory constraint, the basic memory consumption of the third party, and the basic memory consumption of the first party, as the constraint condition of the objective function; and solving the objective function under the constraint condition through a dynamic programming algorithm to obtain the optimal solution of the independent variables.

[0191] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 1. The method further includes: after a first preset duration after putting the first embedding vector into the first channel, determining whether there is first gradient information of the first training sample of the current batch in the second channel corresponding to the current batch; if there is no first gradient information of the first training sample of the current batch in the second channel corresponding to the current batch, sending a first instruction to the first party, where the first instruction is used to instruct the first party to regenerate the first gradient information of the first training sample of the current batch.

[0192] According to one or more embodiments of the present disclosure, Example 8 provides the method of Example 1. The method further includes: in response to receiving a second instruction sent by the first party, re-executing the step of inputting the first training sample into the first feature extraction model to obtain the first embedding vector to the step of putting the first embedding vector into the first channel corresponding to the current batch, where the second instruction is used to instruct the second party to regenerate the first embedding vector.

[0193] According to one or more embodiments of the present disclosure, Example 9 provides a model training method based on vertical federated learning. The multiple parties participating in the model training include a first party with labels and M second parties without labels, where M≥1. After the data held by the multiple parties is aligned with sample identifiers, it is respectively divided into N batches, where N≥1. The N batches correspond one-to-one to N first channels and also correspond one-to-one to N second channels. The method is applied to the first party, and the method includes: obtaining the second training samples of the current batch, where the second training samples of the current batch are one of the N batches; inputting the second training samples into a second feature extraction model to obtain second embedding vectors; taking out M first embedding vectors from the first channel corresponding to the current batch; where the M first embedding vectors are obtained by each of the M second parties without labels inputting the first training samples of the current batch into their local first feature extraction models and putting them into the first channel corresponding to the current batch; training an inference model based on the second embedding vectors and the M first embedding vectors to obtain first gradient information, second gradient information, and third gradient information; putting the first gradient information into the second channel corresponding to the current batch, and using the second gradient information to update the second feature extraction model, and using the third gradient information to update the inference model.

[0194] According to one or more embodiments of the present disclosure, Example 10 provides the method of Example 9, where the first party includes ω a second working nodes, where ω a ≥1, and each of the ω a second working nodes corresponds to at least one of the N batches; the ω aEach of the ω second worker nodes is used to: obtain the second training samples of the current batch, where the current batch is one of the second batches, and the second batches include the at least one batch corresponding to the second worker node; input the second training samples into the second feature extraction model deployed on the second worker node to obtain the second embedding vectors; retrieve the M first embedding vectors from the first channel corresponding to the current batch; based on the second embedding vectors and the M first embedding vectors, train the inference model deployed on the second worker node to obtain the first gradient information, the second gradient information, and the third gradient information; place the first gradient information into the second channel corresponding to the current batch, and update the second feature extraction model deployed on the second worker node using the second gradient information, and update the inference model deployed on the second worker node using the third gradient information.

[0195] According to one or more embodiments of the present disclosure, Example 11 provides the method of Example 10, and the first participating party further includes a second parameter server; the ω a Each of the ω second worker nodes is further used to send the third model parameters of the updated second feature extraction model and the fourth model parameters of the updated inference model to the second parameter server; the second parameter server is further used to aggregate the third model parameters sent by the ω a second worker nodes to obtain fifth model parameters, aggregate the fourth model parameters sent by the ω a second worker nodes to obtain sixth model parameters, and send the fifth model parameters and the sixth model parameters to each of the second worker nodes; the ω a Each of the ω second worker nodes is further used to update the model parameters of the second feature extraction model deployed on the second worker node to the fifth model parameters, and update the model parameters of the inference model deployed on the second worker node to the sixth model parameters.

[0196] According to one or more embodiments of the present disclosure, Example 12 provides the method of Example 11, the ω a Each of the ω second worker nodes is used to send the updated third model parameters and the updated fourth model parameters in the most recent ΔT training rounds to the second parameter server every ΔT training rounds, where ΔT≥1; the second parameter server is used to aggregate the updated third model parameters in the most recent ΔT training rounds sent by the ω a second worker nodes to obtain fifth model parameters, for the ωa aggregate the updated fourth model parameters in the most recent ΔT training rounds sent by a second working node to obtain sixth model parameters, and send the fifth model parameters and the sixth model parameters to each of the second working nodes.

[0197] According to one or more embodiments of the present disclosure, Example 13 provides the method of Example 12, where ΔT has a positive correlation with the number of training rounds.

[0198] According to one or more embodiments of the present disclosure, Example 14 provides the method of Example 10, and the method further includes:

[0199] Construct an objective function with the minimum of the first duration as the optimization goal, where the independent variables of the objective function include the batch size, the first number of the first working points in the second party, and the second number of the second working nodes. The first duration is the maximum of the second duration of a single batch training of the third party and the third duration of a single batch training of the first party. The third party is the second party with the longest single batch training duration among the M unlabeled second parties. The second duration is determined based on the batch size and the first number, and the third duration is determined based on the batch size and the second number. Construct a first memory constraint for the third party according to the basic memory consumption of the third party, and construct a second memory constraint for the first party according to the basic memory consumption of the first party, where the basic memory consumption is the memory size occupied by the party to maintain basic functions. Construct a constraint limit for the batch size according to the first memory constraint, the second memory constraint, the basic memory consumption of the third party, and the basic memory consumption of the first party as the constraint condition of the objective function. Solve the objective function under the constraint condition through a dynamic programming algorithm to obtain the optimal solution of the independent variables.

[0200] According to one or more embodiments of the present disclosure, Example 15 provides the method of Example 9, and the method further includes: after a second preset duration after obtaining the second embedding vector, determine whether the number of the first embedding vectors in the first channel corresponding to the current batch reaches M; if the number of the first embedding vectors in the first channel corresponding to the current batch does not reach M, send a second instruction to the fourth party, where the second instruction is used to instruct the fourth party to regenerate the first embedding vector, and the fourth party includes the second parties among the M unlabeled second parties that have not put the first embedding vector into the first channel.

[0201] According to one or more embodiments of the present disclosure, Example 16 provides the method of Example 9, the method further comprising: in response to receiving a first instruction, re-executing the step of inputting the second training sample into the second feature extraction model to obtain a second embedding vector to the step of putting the first gradient information into the second channel corresponding to the current batch and updating the second feature extraction model by using the second gradient information, and updating the inference model by using the third gradient information, wherein the first instruction is used to instruct the first party to regenerate the first gradient information of the first training sample of the current batch.

[0202] According to one or more embodiments of the present disclosure, Example 17 provides a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processing device, the steps of the method according to any one of Examples 1-16 are implemented.

[0203] According to one or more embodiments of the present disclosure, Example 18 provides an electronic device, comprising: a storage device having a computer program stored thereon; a processing device for executing the computer program in the storage device to implement the steps of the method according to any one of Examples 1-16.

[0204] According to one or more embodiments of the present disclosure, Example 19 provides a computer program product, comprising a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of Examples 1-16 are implemented.

[0205] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principle. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0206] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0207] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.

Claims

1. A model training method based on vertical federated learning, characterized in that: The multiple participants participating in the model training include a first participant with a label and M second participants without a label, wherein M≥1, and the data held by the multiple participants are respectively divided into N batches after sample identification alignment, N≥1, and the N batches correspond one-to-one to the N first channels, and the N batches correspond one-to-one to the N second channels; The method is applied to any one of the M unlabeled second participants, and the method includes: Obtaining a first training sample of a current batch, wherein the first training sample of the current batch is one of the N batches; Inputting the first training sample into a first feature extraction model to obtain a first embedding vector; placing the first embedding vector into a first channel corresponding to the current batch; If there is a non-empty second channel among the N second channels, take out the earliest stored first gradient information from the non-empty second channel; and use the earliest stored first gradient information to update the first feature extraction model; wherein the second channel is used to store the first gradient information of the first training sample of the corresponding batch, the first gradient information is obtained by the first participant training the inference model based on the second embedding vector and the M first embedding vectors taken out from the first channel corresponding to the corresponding batch, and the second embedding vector is obtained by the first participant inputting the second training sample of the corresponding batch into the second feature extraction model; If the N second channels are all empty, the next batch of training samples is obtained to execute the same process until the N batches are all executed.

2. The method according to claim 1, characterized in that: The second participant includes p The first working node, ω p ≥1, the ω p Each of the first working nodes corresponds to at least one batch of the N batches; The ω p Each of the first working nodes is used for: Acquire the first training sample of the current batch, wherein the current batch is one of the first batches, and the first batch includes the at least one batch corresponding to the first working node; Inputting the first training sample into the first feature extraction model deployed on the first working node to obtain the first embedding vector; Putting the first embedding vector into the first channel corresponding to the current batch; If there is a non-empty second channel in the second channels corresponding to the first batch, taking out the earliest stored first gradient information from the non-empty second channels corresponding to the first batch; Using the earliest stored first gradient information, updating the first feature extraction model deployed on the first working node; If the second channels corresponding to the first batch are all empty, the next batch of training samples is obtained to execute the same process until the first batch is completed.

3. The method according to claim 2, characterized in that The second participant also includes a first parameter server; The ω p Each of the first working nodes is further configured to: Sending the updated first model parameters of the first feature extraction model to the first parameter server; The first parameter server is used to p aggregating the first model parameters sent by the first working nodes to obtain second model parameters, and sending the second model parameters to each of the first working nodes; The ω p Each of the first working nodes is also used to update the model parameters of the first feature extraction model deployed on the first working node to the second model parameters.

4. The method according to claim 3, characterized in that The ω p Each of the first working nodes is used to send the first model parameters updated in the most recent ΔT training rounds to the first parameter server every ΔT training rounds, ΔT≥1; The first parameter server is used to p The first model parameters updated in the most recent ΔT training rounds sent by the first working nodes are aggregated to obtain second model parameters, and the second model parameters are sent to each of the first working nodes.

5. The method according to claim 4, characterized in that ΔT is positively correlated with the training rounds.

6. The method according to claim 2, characterized in that When the method is applied to a third party, the method further includes: Taking minimizing the first duration as the optimization goal, constructing an objective function, wherein the third participant is the second participant with the longest single batch training duration among the M unlabeled second participants, the independent variables of the objective function include the batch size, the first number of first working nodes in the second participant, and the second number of second working points in the first participant, the first duration is the maximum value of the second duration of the single batch training of the third participant and the third duration of the single batch training of the first participant, the second duration is determined based on the batch size and the first number, and the third duration is determined based on the batch size and the second number; According to the basic memory consumption of the third party, a first memory constraint of the third party is constructed, and according to the basic memory consumption of the first party, a second memory constraint of the first party is constructed, wherein the basic memory consumption is the memory size occupied by the party to maintain basic functions; According to the first memory constraint, the second memory constraint, the basic memory consumption of the third party and the basic memory consumption of the first party, construct a constraint limit of the batch size as a constraint condition of the objective function; The objective function is solved under the constraint conditions by a dynamic programming algorithm to obtain the optimal solution of the independent variable.

7. The method according to claim 1, characterized in that The method further comprises: After a first preset time length after the first embedding vector is placed in the first channel, determining whether first gradient information of a first training sample of the current batch exists in the second channel corresponding to the current batch; If the first gradient information of the first training sample of the current batch does not exist in the second channel corresponding to the current batch, a first instruction is sent to the first participant, wherein the first instruction is used to instruct the first participant to regenerate the first gradient information of the first training sample of the current batch.

8. The method according to claim 1, characterized in that The method further comprises: In response to receiving a second instruction sent by the first participant, re-execute the step of inputting the first training sample into the first feature extraction model to obtain the first embedding vector to the step of placing the first embedding vector into the first channel corresponding to the current batch, wherein the second instruction is used to instruct the second participant to regenerate the first embedding vector.

9. A model training method based on vertical federated learning, characterized in that: The multiple participants participating in the model training include a first participant with a label and M second participants without a label, wherein M≥1, and the data held by the multiple participants are respectively divided into N batches after sample identification alignment, N≥1, and the N batches correspond one-to-one to the N first channels, and the N batches correspond one-to-one to the N second channels; The method is applied to the first participant, and the method comprises: Acquire a second training sample of a current batch, wherein the second training sample of the current batch is one of the N batches; Inputting the second training sample into a second feature extraction model to obtain a second embedding vector; Taking out M first embedding vectors from the first channel corresponding to the current batch; wherein the M first embedding vectors are obtained by each of the M unlabeled second participants respectively inputting the first training sample of the current batch into a local first feature extraction model and put into the first channel corresponding to the current batch; Based on the second embedding vector and the M first embedding vectors, training the inference model to obtain first gradient information, second gradient information, and third gradient information; The first gradient information is placed in the second channel corresponding to the current batch, and the second feature extraction model is updated using the second gradient information, and the inference model is updated using the third gradient information.

10. The method according to claim 9, characterized in that The first participant includes a The second working node, ω a ≥1, the ω a Each second working node of the second working nodes corresponds to at least one batch of the N batches; The ω a Each of the second working nodes is used for: Acquire the second training sample of the current batch, wherein the current batch is one of the second batches, wherein the second current batch includes the at least one batch corresponding to the second working node; Inputting the second training sample into the second feature extraction model deployed on the second working node to obtain the second embedding vector; Extract the M first embedding vectors from the first channel corresponding to the current batch; Based on the second embedding vector and the M first embedding vectors, training the inference model deployed on the second working node to obtain the first gradient information, the second gradient information, and the third gradient information; The first gradient information is placed in the second channel corresponding to the current batch, and the second feature extraction model deployed on the second working node is updated using the second gradient information, and the inference model deployed on the second working node is updated using the third gradient information.

11. The method according to claim 10, characterized in that The first participant also includes a second parameter server; The ω a Each of the second working nodes is further used to send the updated third model parameter of the second feature extraction model and the updated fourth model parameter of the inference model to the second parameter server; The second parameter server is further configured to a The third model parameters sent by the second working nodes are aggregated to obtain the fifth model parameters, and the ω a aggregating the fourth model parameters sent by the second working nodes to obtain sixth model parameters, and sending the fifth model parameters and the sixth model parameters to each of the second working nodes; The ω a Each of the second working nodes is also used to update the model parameters of the second feature extraction model deployed on the second working node to the fifth model parameters, and to update the model parameters of the inference model deployed on the second working node to the sixth model parameters.

12. The method according to claim 11, characterized in that The ω a Each of the second working nodes is used to send the third model parameters updated in the most recent ΔT training rounds and the fourth model parameters updated in the last ΔT training rounds to the second parameter server every ΔT training rounds, ΔT≥1; The second parameter server is used to a Aggregate the third model parameters updated in the most recent ΔT training rounds sent by the second working nodes to obtain the fifth model parameters, and a The fourth model parameters updated in the most recent ΔT training rounds sent by the second working nodes are aggregated to obtain sixth model parameters, and the fifth model parameters and the sixth model parameters are sent to each of the second working nodes.

13. The method according to claim 12, characterized in that ΔT is positively correlated with the training rounds.

14. The method according to claim 10, characterized in that The method further comprises: Taking minimizing the first duration as the optimization objective, constructing an objective function, wherein the independent variables of the objective function include the batch size, the first number of the first working points in the second participant, and the second number of the second working nodes, the first duration is the maximum value of the second duration of the single batch training of the third participant and the third duration of the single batch training of the first participant, the third participant is the second participant with the longest single batch training duration among the M unlabeled second participants, the second duration is determined based on the batch size and the first number, and the third duration is determined based on the batch size and the second number; According to the basic memory consumption of the third party, a first memory constraint of the third party is constructed, and according to the basic memory consumption of the first party, a second memory constraint of the first party is constructed, wherein the basic memory consumption is the memory size occupied by the party to maintain basic functions; According to the first memory constraint, the second memory constraint, the basic memory consumption of the third party and the basic memory consumption of the first party, construct a constraint limit of the batch size as a constraint condition of the objective function; The objective function is solved under the constraint conditions by a dynamic programming algorithm to obtain the optimal solution of the independent variable.

15. The method according to claim 9, characterized in that The method further comprises: After a second preset time after obtaining the second embedding vector, determining whether the number of the first embedding vectors in the first channel corresponding to the current batch reaches M; If the number of the first embedding vectors in the first channel corresponding to the current batch does not reach M, a second instruction is sent to the fourth participant, wherein the second instruction is used to instruct the fourth participant to regenerate the first embedding vector, and the fourth participant includes the second participant among the M unlabeled second participants that has not placed the first embedding vector into the first channel.

16. The method according to claim 9, characterized in that The method further comprises: In response to receiving the first instruction, the step of inputting the second training sample into the second feature extraction model to obtain the second embedding vector is re-executed to the step of placing the first gradient information into the second channel corresponding to the current batch, and updating the second feature extraction model using the second gradient information, and updating the inference model using the third gradient information, wherein the first instruction is used to instruct the first participant to regenerate the first gradient information of the first training sample of the current batch.

17. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 16 are implemented.

18. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 16.

19. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 16 are implemented.

Citation Information

Patent Citations

  • Method, device, equipment and system for obtaining model and readable storage medium

    CN114444708A

  • Federal learning method

    CN115099424A

  • Longitudinal federal learning model training method and device, storage medium and program product

    CN116245202A

  • Asynchronous model training method and system based on federal learning and electronic equipment

    CN118396084A

  • Trusted and decentralized aggregation for federated learning

    US20220374762A1