Self-supervised longitudinal federated learning method and device based on instance similarity and dynamic balance pool

By combining instance similarity and dynamic balance pool modules in vertical federated learning, the problem of insufficient model performance in scarce labels is solved, and better generalization and representation capabilities are achieved, while avoiding regularization conflicts.

CN120218188AActive Publication Date: 2025-06-27杭州半云科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510291354.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Traditional vertical federated learning methods are difficult to effectively utilize the participants’ unlabeled data in the context of scarce labels, and the self-supervised learning method conflicts with the conventional regularization method, affecting the model performance.

Method used

A self-supervised vertical federated learning method based on instance similarity and dynamic balance pooling is adopted, inter-domain knowledge is captured through the pre-training stage using aligned samples and labelless data, and a dynamic balance pool module is used to dynamically balance inter-domain and intra-domain knowledge in the fine-tuning stage to avoid regularization conflicts.

Benefits of technology

The performance of vertical federated learning models in scarce label scenarios has been improved, the generalization and representation capabilities of the model have been enhanced, and the regularization conflict problem has been avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218188A_ABST
    Figure CN120218188A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of federated learning, and discloses a self-supervised longitudinal federated learning method and device based on instance similarity and a dynamic balance pool, which combine a comparison self-supervised method with instance similarity, can help private longitudinal federated learning models of participants to capture more important inter-domain knowledge, and improve the learning efficiency. Therefore, the generalization ability and the representation ability of each participant model are improved, and regularization conflicts are avoided at the same time. Besides, a dynamic balance pool module is provided and is used for performing fine adjustment on the pre-training model in a downstream supervision task and dynamically balancing inter-domain and intra-domain knowledge, so that the performance of the joint longitudinal federal learning model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of federated learning, and particularly relates to a self-supervised vertical federated learning method and device based on instance similarity and a dynamic balance pool. Background Art

[0002] Federated learning (FL), as a promising paradigm, enables multiple independent parties to collaboratively execute machine learning tasks while protecting user privacy data. Vertical federated learning (VFL) is one of the important research fields in FL, and its application scenario is that the data sets owned by the parties are different in the feature space, but may have partial overlap in the sample ID space. Due to its ability to enhance model performance by using complementary data while ensuring data privacy, VFL has shown great application prospects in fields such as finance, advertising, and healthcare, attracting extensive attention from academia and industry.

[0003] Traditional VFL usually uses aligned and labeled data to train the VFL model. However, the process of data labeling is both time-consuming and expensive, and the aligned samples between parties are relatively scarce. Therefore, in practical application scenarios, the parties can only obtain a limited number of aligned and labeled samples, which is not conducive to the training of the VFL model.

[0004] To facilitate the use of a large amount of high-value unlabeled data held by the parties, current VFL frameworks usually integrate contrastive self-supervised learning (CSSL) methods. For example, combining the CSSL method with a cross-perspective method to utilize the labeled and unlabeled samples of each party. However, the CSSL method usually conflicts with conventional regularization methods aimed at improving model performance because the CSSL method aims to capture the internal information of the representation to maximize the discrimination between different samples, which significantly deviates from the objectives of conventional regularization methods. For example, the Kullback-Leibler (KL) divergence is a recognized effective regularization technique. However, existing research has found that introducing the KL divergence into the CSSL method will increase the instability during the training process. The representations with tiny changes learned by the CSSL method can lead to significant fluctuations in the KL divergence, thereby weakening the representation ability of the model and having a negative impact on model performance. Therefore, how to solve the conflict between the CSSL method and conventional regularization methods to improve the performance of the VFL model has become an urgent problem to be solved. Summary of the Invention

[0005] The purpose of the present invention is to provide a self-supervised vertical federated learning method and device based on instance similarity and a dynamic balance pool, which solves the conflict between the CSSL method and conventional regularization methods and improves the performance of the VFL model.

[0006] To achieve the above purpose, the technical solutions adopted by the present invention are as follows:

[0007] In a first aspect, the present invention provides a self-supervised vertical federated learning method based on instance similarity and a dynamic balance pool, which is used for K participants and a server to jointly train a joint vertical federated learning model. The joint vertical federated learning model includes a teacher network, a local network, a dynamic balance pool module, and a classifier. Among the K participants, those with labels are active parties, and the rest are passive parties. The self-supervised vertical federated learning method based on instance similarity and a dynamic balance pool includes a pre-training stage and a fine-tuning stage, where:

[0008] The pre-training stage includes:

[0009] Each participant uses aligned samples to jointly train its own teacher network. During the joint training, the instance similarity is calculated based on the output of the teacher network, and the loss is calculated in combination with the instance similarity;

[0010] Each participant uses its local private samples and its own teacher network to train its own local network;

[0011] Each participant uploads its local network to the server. After the server aggregates all the local networks, a global network is generated and sent back to each participant as a new local network;

[0012] Repeat the pre-training stage until the pre-training is completed;

[0013] The fine-tuning stage includes:

[0014] Each participant inputs the aligned and labeled samples into its own teacher network and local network respectively, and uses the dynamic balance pool module to aggregate the intermediate representations output by the teacher network and the intermediate representations output by the local network to obtain its own final representation;

[0015] The passive party sends the final representation to the active party for aggregation to obtain an aggregated representation;

[0016] The active party uses the classifier to convert the aggregated representation into a predicted value, and trains the dynamic balance pool module and the classifier based on the predicted value and the label;

[0017] Repeat the fine-tuning stage until the fine-tuning is completed, and output the trained joint vertical federated learning model.

[0018] The following also provides several optional methods, which are not additional limitations to the above overall solution, but are only further supplements or optimizations. Without technical or logical contradictions, each optional method can be combined with the above overall solution alone, or multiple optional methods can be combined with each other.

[0019] Preferably, when each participant uses aligned samples to jointly train its own teacher network, it includes:

[0020] The teacher network includes a teacher encoder and a teacher projector. The aligned samples are input into the teacher encoder to be converted into teacher intermediate representations, and then the teacher intermediate representations are input into the teacher projector to be converted into teacher final representations;

[0021] Generate the first positive sample pair based on the teacher final representation, and calculate the representation self-supervised loss according to the first positive sample pair;

[0022] Calculate the instance similarity of each party based on the teacher intermediate representation, generate the second positive sample pair according to the instance similarity, and calculate the instance self-supervised loss according to the second positive sample pair;

[0023] Aggregate the representation self-supervised loss and the instance self-supervised loss as the final co-training loss, and update the teacher network according to the final co-training loss.

[0024] Preferably, the generating the first positive sample pair based on the teacher final representation and calculating the representation self-supervised loss according to the first positive sample pair includes:

[0025] Take the teacher final representation of party k

[0026] Each party uploads the teacher final representation to the server, and the server calculates the average teacher representation and distributes the average teacher representation to each party;

[0027] After each party receives the average teacher representation it matches the teacher final representation converted from the same sample and the average teacher representation as the first positive sample pair, and calculates the self-supervised loss based on the first positive sample pair as the representation self-supervised loss.

[0028] Preferably, the calculating the instance similarity of each party based on the teacher intermediate representation, generating the second positive sample pair according to the instance similarity, and calculating the instance self-supervised loss according to the second positive sample pair includes:

[0029] Take the teacher intermediate representation of party k Construct the instance similarity matrix of party k as follows:

[0030]

[0031] where (·) T is the transpose operation, ω is the temperature hyperparameter, and ||·||2 is the L2 norm;

[0032] Obtain the instance similarity S of participant k k as follows:

[0033]

[0034] In the formula, Remove_diag(·) represents the operation of eliminating the diagonal elements in the matrix;

[0035] Each participant uploads the instance similarity S k to the server, and the server calculates the target instance similarity and distributes the target instance similarity to each participant;

[0036] After each participant receives the target instance similarity it will match the instance similarity S k obtained by converting the same sample and the target instance similarity as the second positive sample pair, and calculate the self-supervised loss based on the second positive sample pair as the instance self-supervised loss.

[0037] Preferably, each participant uses the local private samples and their respective teacher networks to train their respective local networks, including:

[0038] Regarding the local network as the local online network, and obtaining the local target network based on the local online network by using the contrastive self-supervised learning method;

[0039] Adopt the data augmentation strategy to convert the local private samples into the first augmented view and the second augmented view;

[0040] Input the first augmented view into the local online network, and input the second augmented view into the local target network to generate the third positive sample pair;

[0041] Input the second augmented view into the local online network, and input the first augmented view into the local target network to generate the fourth positive sample pair;

[0042] Calculate the self-supervised loss based on the third positive sample pair and the fourth positive sample pair as the symmetric loss;

[0043] Input the first augmented view into the local online network, and input the first augmented view into the teacher network to generate the fifth positive sample pair;

[0044] Input the second augmented view into the local online network, and input the second augmented view into the teacher network to generate the sixth positive sample pair;

[0045] Calculate the self-supervised loss based on the fifth positive sample pair and the sixth positive sample pair as the contrastive knowledge transfer loss;

[0046] Update the local online network by combining the symmetric loss and the contrastive knowledge transfer loss.

[0047] Preferably, each participant uploads the local network to the server, and the server aggregates the local networks to generate a global network, including:

[0048] The local network includes a local encoder and a local projector. Each participant divides the local encoder into a local bottom encoder and a local top encoder;

[0049] Each participant uploads the local top encoder and the local projector to the server;

[0050] The server aggregates the local top encoders to obtain a global top encoder, aggregates the local projectors to obtain a global projector, and uses the global top encoder and the global projector as the global network.

[0051] Preferably, the local encoder includes two fully connected layers, and the last layer of the local encoder is divided as the local top encoder, and the remaining layers are used as the local bottom encoder;

[0052] Alternatively, the local encoder includes an embedding layer and two fully connected layers, and the last layer of the local encoder is divided as the local top encoder, and the remaining layers are used as the local bottom encoder;

[0053] Alternatively, the local encoder is a ResNet-18 network, and the last three modules of the local encoder are divided as the local top encoder, and the remaining layers are used as the local bottom encoder.

[0054] Preferably, using the dynamic balance pool module to aggregate the intermediate representations output by the teacher network and the intermediate representations output by the local network to obtain their respective final representations, including:

[0055] Take the intermediate representation output by the teacher network of participant k and the intermediate representation output by the local network of participant k Calculate the forgetting factor O k as follows:

[0056]

[0057] In the formula, represents the concatenation operation along the feature axis, MLP(·) is a multi-layer perceptron, and σ(·) is the sigmoid function;

[0058] Calculate the update factor U k as follows:

[0059]

[0060] Extract the implicit information Ik As follows:

[0061]

[0062] Wherein, φ(·) is an activation function, FC(·) is a fully connected layer, and ⊙ is a Hadamard product;

[0063] Then the final representation is calculated as follows:

[0064]

[0065] Wherein, R k is the final representation obtained by the participating party k.

[0066] In a second aspect, a self-supervised vertical federated learning device based on instance similarity and dynamic balance pool according to the present invention, the self-supervised vertical federated learning device based on instance similarity and dynamic balance pool includes K participating parties and a server, and the K participating parties and a server cooperate to train a joint vertical federated learning model. The joint vertical federated learning model includes a teacher network, a local network, a dynamic balance pool module and a classifier. Among the K participating parties, those with labels are used as the active parties, and the rest are used as the passive parties. The K participating parties and a server perform the following operations:

[0067] In the pre-training stage:

[0068] Each participating party uses the aligned samples to jointly train its own teacher network, calculates the instance similarity based on the output of the teacher network during the joint training, and calculates the loss in combination with the instance similarity;

[0069] Each participating party uses its local private samples and its own teacher network to train its own local network;

[0070] Each participating party uploads its local network to the server, and the server aggregates the local networks to generate a global network and distributes the global network to each participating party as a new local network;

[0071] Each participating party and the server repeat the pre-training stage until the pre-training ends;

[0072] In the fine-tuning stage:

[0073] Each participating party inputs the aligned and labeled samples into its own teacher network and local network respectively, and uses the dynamic balance pool module to aggregate the intermediate representations output by the teacher network and the intermediate representations output by the local network to obtain its own final representation;

[0074] The passive party sends the final representation to the active party for aggregation to obtain an aggregated representation;

[0075] The initiator uses a classifier to convert the aggregated representation into a predicted value, and trains the dynamic balance pool module and the classifier according to the predicted value and the label;

[0076] Each participating party and the server repeatedly execute the fine-tuning stage until the fine-tuning ends, and output the trained joint vertical federated learning model.

[0077] A self-supervised vertical federated learning method and device based on instance similarity and dynamic balance pool provided by the present invention. In the pre-training stage, compared with the traditional CSSL method, the present invention combines the CSSL method with instance similarity, which can help the private VFL models of each participating party capture more important inter-domain knowledge, thereby improving the generalization ability and representation ability of the models of each participating party, and avoiding regularization conflicts at the same time. In addition, the present invention proposes a dynamic balance pool (abbreviated as DBP) module, which is used to fine-tune the pre-trained model in the downstream supervised task and dynamically balance inter-domain and intra-domain knowledge, thereby improving the performance of the joint VFL model. A large number of experiments conducted on image datasets and tabular datasets show that the present invention has achieved leading performance in scenarios with scarce labels and does not produce regularization conflict problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 It is a schematic diagram for sample discrimination of the present invention;

[0079] Figure 2 It is a flowchart of the self-supervised vertical federated learning method based on instance similarity and dynamic balance pool of the present invention;

[0080] Figure 3 It is a flowchart of the pre-training stage of the present invention;

[0081] Figure 4 It is a flowchart of the downstream supervised task of the present invention;

[0082] Figure 5 It is a schematic structural diagram of the DBP module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0083] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0084] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.

[0085] Similar to the conventional VFL setting, in this embodiment, there are K participants (indexed by k) and a server collaborating to train a joint VFL model. Participant k holds a local private sample dataset where is the complete dataset partitioned vertically. Assume that only Participant 1 has labels, so it is selected as the active party, while the other participants only have features and thus are all passive parties. The participants link their datasets by aligning samples with the same ID, thereby obtaining the complete dataset However, in a real scenario, the complete dataset only contains extremely limited aligned and labeled samples where is the part held by Participant k, and is the corresponding true label held by the active party. The aligned but unlabeled samples are denoted as where is the part held by Participant k. In addition represents the unaligned private samples of Participant k. Figure 1 shows an example of data partitioning for two participants, where Participant 1 holds the private samples this time where Participant 2 holds the private samples this time

[0086] Different from most traditional VFL models that only use limited aligned and labeled data X al , this embodiment uses all aligned and unaligned data to pre-train the local private models of each participant. In addition, this embodiment requires a fine-tuning stage later to utilize a small amount of aligned and labeled samples X al to perform downstream supervised tasks. This embodiment aims to improve the performance of the joint VFL model in downstream supervised tasks. Therefore, the performance of this embodiment can be evaluated by downstream supervised tasks.

[0087] Such as Figure 2As shown in the figure, this embodiment proposes a self-supervised vertical federated learning method based on instance similarity and dynamic balance pool (abbreviated as ISSVFL) for K participants and a server to jointly train a joint vertical federated learning model. The joint vertical federated learning model includes a teacher network, a local network, a dynamic balance pool module, and a classifier. Among the K participants, those with labels are regarded as active parties, and the rest are regarded as passive parties. The core idea of ISSVFL is to achieve representation-level knowledge transfer by integrating the CSSL method and instance similarity in the VFL environment, thereby improving the inter-domain generalization ability and representation ability of the private VFL models of each participant, while avoiding the problem of regularization conflict. In addition, ISSVFL proposes a DBP module for dynamically balancing inter-domain knowledge and intra-domain knowledge. Specifically, ISSVFL includes two stages: the pre-training stage and the fine-tuning stage.

[0088] (1) Pre-training stage: As Figure 3 shown, in this stage, all available samples are used to pre-train the private VFL models of each participant in the scenario of scarce labels, so that the private models can generate representations of inter-domain knowledge and intra-domain knowledge. The private VEL models of each participant are the teacher network and the local network parts in the joint vertical federated learning model.

[0089] This embodiment introduces the CSSL method in the pre-training stage. The CSSL method is based on the instance discrimination task, regarding each instance and its transformed body as an independent category. The CSSL method directly compares representations rather than real labels, aiming to generate representations that can effectively distinguish different instances. ISSVFL integrates three representative CSSL methods into the VFL environment respectively, namely MoCo (Momentum Contrast), BYOL (Bootstrap Your Own Latent), and SimSiam (Simple Siamese). For MoCo and BYOL, the target network is the momentum version of the online network, while the target network of SimSiam is equivalent to the online network. The target network applies a gradient stop operation. And, BYOL and SimSiam additionally add a predictor on the top layer of the online network. In addition, after matching the positive sample pairs in MoCo, the other samples in the same batch are regarded as negative samples, and the positive sample pairs and negative samples participate in the self-supervised loss calculation at the same time. For the sake of clear representation, Figure 3 the predictors of BYOL and SimSiam are omitted in

[0090] (1.1) Each participant uses aligned samples to jointly train their respective teacher networks. During the joint training, the instance similarity is calculated based on the output of the teacher network, and the loss is calculated in combination with the instance similarity. The teacher network includes a teacher encoder and a teacher projector. If BYOL and SimSiam are adopted, it also includes a teacher predictor.

[0091] For the aligned samples of each party, the features held by each party are equivalent to the vertical segmentation of the entire feature set and can be used as complementary views. The complementary views contain rich inter-domain knowledge and naturally form positive samples. Therefore, each party uses the aligned samples to co-train their respective teacher networks through the CSSL method, enabling the teacher networks to learn more inter-domain knowledge.

[0092] Specifically, given a batch of aligned samples The vertically segmented part held by party k is denoted as For party k, its teacher encoder Converts the input Into an intermediate representation (i.e., the teacher intermediate representation), and then the teacher projector Converts the intermediate representation Into a final representation (i.e., the teacher final representation). When using BYOL or SimSiam, the teacher predictor Converts the final Into a predicted representation

[0093] To control the communication overhead and computational overhead, each party uploads the final representation To the server, and the server calculates the average teacher representation Then the average teacher representation Is sent down to each party. During this process, all representations participating in communication apply the gradient stop operation. Since each positive sample pair consists of complementary views of the same sample, each party can easily match the corresponding first positive sample pair After receiving (or ), Is An element of Is An element of Is An element of (abbreviation: representation self-supervised loss) to update their respective teacher networks (i.e., Or ), as shown in formula (1).

[0094]

[0095] Where Is the self-supervised loss using MoCo, BYOL or SimSiam.

[0096] The intermediate representation implies general structural information, which helps improve the generalization ability of the VFL model. In this embodiment, a representation-level knowledge transfer method based on the similarity of aligned sample instances (abbreviated as FCISL) is proposed, aiming to capture important inter-domain knowledge contained in the intermediate representation by constructing an instance similarity matrix.

[0097] Specifically, given the same batch of aligned samples The teacher encoder of Party k maps to an intermediate representation of d dimensions For each Party k, the instance similarity matrix constructed by its teacher network is shown in Equation (2).

[0098]

[0099] where (·) T is the transpose operation, and ω is the temperature hyperparameter. Then, FCISL eliminates the diagonal elements in k i.e., the similarity values of each sample to itself. The final instance similarity S calculated by the teacher network of Party k

[0100]

[0101] is shown in Equation (3).

[0102] Considering the computational and communication burdens, each party uploads the instance similarity S k to the server, and the server calculates the target instance similarity and distributes the target instance similarity to each party. During this process, all instance similarities apply the gradient stop operation. Similar to FCSSL, after Party k receives , it matches the corresponding second positive sample pairs in S k and Then, each party updates its teacher network by aligning the target instance similarity through CSSL. The FCISL loss of Party k abbreviated as the instance self-supervised loss) is shown in Equation (4). is shown in Equation (4).

[0103]

[0104] As shown in formula (4), the FCISL method aims to achieve representation-level knowledge transfer in VFL, enabling the model to capture richer inter-domain knowledge from the representations of different parties, thereby enhancing the generalization ability and representation ability of the teacher network while avoiding regularization conflicts.

[0105] As shown in formula (5), aggregating the in formula (1) and the in formula (4), Party k can optimize its teacher network by minimizing the final loss in collaborative training

[0106]

[0107] where μ is a hyperparameter that weighs the and

[0108] (1.2) Each party uses its local private samples and its respective teacher network to train its local network.

[0109] In the local training of ISSVFL, each party uses its local private samples to pre-train its local network to learn intra-domain knowledge through the local network. At the same time, ISSVFL adopts a contrastive knowledge transfer (CKT) module to prevent the local network from overfitting to its local private data distribution.

[0110] For Party k, given a batch of local private samples adopt a data augmentation strategy to transform into the corresponding first augmented view and the second augmented view Based on the same data augmentation strategy, different outputs are obtained according to the same input. In this embodiment, the local network is used as the local online network, and a contrastive self-supervised learning method is adopted to obtain the local target network based on the local online network.

[0111] Input the first augmented view into the local online network, and input the second augmented view into the local target network to generate the third positive sample pair. Specifically: the local online encoder f l k transforms the first augmented view V1 k into an intermediate representation while the local target encoder transforms the second augmented view into an intermediate representation Then, the local online projector transforms the intermediate representation into an online representation ​​Local target projector Convert the intermediate representation into the target representation When using BYOL or SimSiam, the local online predictor further converts into the prediction representation Since each positive sample is derived from the same sample, Party k can directly match (or ) and the corresponding positive sample pairs in to obtain the third positive sample pair (or ), is an element of is an element of is an element of

[0112] Input the second augmented view into the local online network and the first augmented view into the local target network to generate the fourth positive sample pair. Specifically: The symmetric calculation path adopted in this embodiment exchanges the first augmented view V1 k and the second augmented view as the input of the local network to obtain (or and Party k can also directly match (or ) and the corresponding positive sample pairs in to obtain the fourth positive sample pair (or ), is an element of is an element of is an element of

[0113] Calculate the self-supervised loss as the symmetric loss based on the third positive sample pair and the fourth positive sample pair. As shown in formula (6), Party k updates its local online network (i.e., f l k ,, or ) by minimizing the symmetric loss

[0114]

[0115] By minimizing Positive sample pairs attract each other, while negative sample pairs repel each other. For each party, its local target network is the momentum version of the local online network, and the momentum update mode depends on the specific CSSL method.

[0116] Input the first augmented view into the local online network and the teacher network to generate the fifth positive sample pair; input the second augmented view into the local online network and the teacher network to generate the sixth positive sample pair. Specifically, ISSVFL further adopts the CKT module, combines CSSL with knowledge transfer, and transfers the cross-domain knowledge in the teacher network to the local network while keeping the teacher network frozen. Specifically, for V1 converted from k and calculate the representations and respectively from the frozen teacher network. Since each positive sample pair originates from the same augmented view, party k can directly match (or ) and with the corresponding positive samples in (or ) to obtain the sixth positive sample pair. Similarly, party k can directly match (or ) and with the corresponding positive samples in (or ) to obtain the fifth positive sample pair. Therefore, the CKT loss of party k is as shown in formula (7).

[0117]

[0118] where λ is used to control the strength of CKT.

[0119] Update the local online network by combining the contrastive self-supervised learning loss and the contrastive knowledge transfer loss. CKT works in concert with the local SSL loss in formula (6) to balance the knowledge from other parties and the local area, thereby improving the inter-domain generalization performance and intra-domain discriminative performance of the local network. Therefore, as shown in formula (8), party k performs local training by minimizing the total local loss to update the local online network, i.e., the local network.

[0120]

[0121] (1.3) Each party uploads its local network to the server. After aggregating all local networks, the server generates a global network and distributes the global network to each party as a new local network.

[0122] The model aggregation algorithm is crucial for VFL model training and can integrate all the knowledge of local networks while protecting data privacy. ISSVFL adopts the commonly used aggregation mechanism FedAvg, enabling the aggregated global encoder to generate a set of common representations that are convenient for all participating parties to share.

[0123] Since the features held by each participating party in the VFL environment are heterogeneous, participating party k uploads its local network to the server for aggregation. Specifically, participating party k first divides the local encoder f in its local network l k into a participating-party-specific local bottom encoder and a local top encoder Then it uploads them to the server. The server is responsible for aggregating the uploaded local top encoders, local projectors, and local predictors to obtain the global top encoder the global projector and the global predictor where n k is the number of samples owned by participating party k, and N is the total number of samples owned by all participating parties. Finally, the server sends back to each participating party as the new local network.

[0124] Among them, the division of the local encoder is an existing method, and this embodiment does not limit it. For example, if the local encoder contains two fully connected layers, the last layer of the local encoder is divided as the local top encoder, and the remaining layers are used as the local bottom encoder; or, if the local encoder contains an embedding layer and two fully connected layers, the last layer of the local encoder is divided as the local top encoder, and the remaining layers are used as the local bottom encoder; or, if the local encoder is a ResNet-18 network, the last three modules of the local encoder are divided as the local top encoder, and the remaining layers are used as the local bottom encoder.

[0125] (1.4) Re-execute steps (1.1) - (1.3) for the next round of training until the pre-training ends. This embodiment sets the pre-training end condition as the iteration reaching a preset number of times.

[0126] (2) Fine-tuning stage: After the pre-training stage, participating party k obtains the teacher encoder f that has completed pre-training T k and the local online encoder (i.e., the local encoder) f l k. Further, a joint vertical federated learning model is obtained based on the teacher encoders and local encoders of each participant. The joint vertical federated learning model includes a teacher network, a local network, a dynamic balance pool module, and a classifier. Each participant has a teacher network, a local network, and a dynamic balance pool module, and the initiator also has a classifier. The teacher network and the local network in each participant are arranged in parallel, and respectively receive the input aligned and labeled samples, and input the outputs to the dynamic balance pool module. The downstream supervision task is jointly carried out by all participants using a small number of aligned and labeled samples (i.e., X al ) to fine-tune the classifier held by the actively participating party to adapt to a specific downstream supervision task. The downstream supervision task process is as shown in Figure 4 .

[0127] (2.1) Each participant respectively inputs the aligned and labeled samples into its own teacher network and local network, and uses the dynamic balance pool module to aggregate the intermediate representations output by the teacher network and the intermediate representations output by the local network to obtain its own final representation.

[0128] The teacher encoder of participant k converts the input aligned and labeled samples of participant k into an intermediate representation The local online encoder f of participant k l k converts into an intermediate representation To provide a representation containing richer general information for the downstream supervision task, ISSVFL proposes a novel dynamic balance pool (DBP) module for dynamically balancing inter-domain and intra-domain knowledge in the VFL environment.

[0129] It should be noted that in this embodiment, the outputs of the same network are represented by the same symbol, regardless of whether their inputs are the same. For example, the output of the teacher encoder for the aligned samples is represented as The output of the teacher encoder for the aligned and labeled samples is also represented as to facilitate understanding that they are all the outputs of the teacher encoder . In other embodiments, the outputs of the same network for different inputs can be represented by different symbols, which is not limited in this embodiment.

[0130] As shown in Figure 5 , the DBP module converges and to provide the final representation R for the downstream supervision task k .

[0131] Specifically, as shown in formula (9), in the DBP module, the forgetting factor O k is calculated by the forgetting pool.

[0132]

[0133] Among them, represents the concatenation operation along the feature axis, σ(·) is the sigmoid function, and MLP(·) is the multi-layer perceptron.

[0134] Similarly, as shown in formula (10), the update factor U k is calculated by the update pool.

[0135]

[0136] The forgetting factor O k and the update factor U k are calculated in the same way, but the parameters of the multi-layer perceptron are different.

[0137] Then, as shown in formula (11), the implicit information I k is calculated by the implicit information extractor.

[0138]

[0139] Among them, φ(·) is the activation function, FC(·) is the fully connected layer, and ⊙ is the Hadamard product.

[0140] As shown in formula (12), the final representation R k obtained by the participant k is calculated by the balance pool of the DBP module.

[0141]

[0142] In the DBP module, the forgetting factor O k decides whether to ignore the knowledge from the teacher encoder. The update factor U k decides how much knowledge in the implicit information is retained in the final representation R k .

[0143] (2.2) The passive party sends the final representation to the active party for aggregation to obtain the aggregated representation; the active party uses the classifier to convert the aggregated representation into a prediction value, and trains the dynamic balance pool module and the classifier according to the prediction value and the label.

[0144] Referring to the traditional VFL method, the passive party then sends the final representation R j (j ∈ [2, K]) to the active party for aggregation. As shown in formula (13), the only classifier g c held by the active party converts the aggregated representation into the final prediction value

[0145]

[0146] Among them, represents the aggregated representation after aggregation.

[0147] Finally, as shown in formula (14), by minimizing the global objective function fine-tuning can be carried out collaboratively to train the dynamic balance pool module and the classifier.

[0148]

[0149] where l CE is the cross-entropy loss, is the corresponding true label held by the active party. In this embodiment, the first participating party is taken as an example of the active party. The final performance of ISSVFL is evaluated after fine-tuning. If there is only one active party among the participating parties, all passive parties will send the final representation to the active party for aggregation; if there are multiple active parties among the participating parties, the passive parties will send their respective final representations to each active party respectively, and each active party will perform aggregation separately and calculate the loss respectively. Finally, the losses calculated by all active parties are aggregated proportionally as the final loss, so as to achieve collaborative fine-tuning.

[0150] (2.3) Repeat steps (2.1) - (2.2) to perform fine-tuning until the fine-tuning ends.

[0151] In order to demonstrate the advantages of the ISSVFL method of the present invention, the following experiments are carried out.

[0152] (1) Experimental settings.

[0153] The performance of ISSVFL was evaluated on image datasets and tabular datasets in the experiment.

[0154] (1.1) Datasets and Models: To evaluate the performance of ISSVFL, experiments were conducted on four datasets: NUS-WIDE, Avazu, Breast Histopathological Images (abbreviated as BHI), and ModelNet. The first two are tabular datasets, and the last two are image datasets. For the NUS-WIDE dataset, in this experiment, data from 10 categories were selected from 81 available categories for the multi-class classification task. To simulate the VFL environment, one party holds image features and the other holds text features. For the Avazu dataset, in this experiment, 14 categorical features and 8 continuous features were randomly assigned to two participating parties to predict click-through rates. In addition, the categorical features were transformed into 32-dimensional embeddings. Considering the computational burden, 800,000 samples were randomly selected as the training set and 200,000 samples as the test set in this experiment. For the BHI dataset, in this experiment, two different images of the same patient were assigned to two parties for the binary classification task. For the ModelNet dataset, 15 categories were selected for the multi-class classification task in this experiment. Each category contains multiple 3D objects, and 12 pictures were generated for each object to represent different perspectives. To simulate the VFL environment, 12 perspectives were evenly assigned to four parties, and each party received 3 adjacent perspectives. To increase the complexity of the dataset, in this experiment, one picture was randomly selected from each participating party to construct the VFL samples for each object.

[0155] The model architectures of each participating party on each dataset are shown in Table 1. All classifiers g held by the active participating party c are composed of a single fully connected layer (abbreviated as FC). For all datasets, all projectors are composed of three FCs, and the dimension of each FC is 512. All predictors are composed of two FCs, and the dimensions of the FCs are set to 128 and 512 in sequence.

[0156] Table 1 Model Architectures for Different Datasets

[0157] Dataset <![CDATA[Local encoder and teacher encoder (f l and f T )]]> <![CDATA[Local top encoder f lt > NUS-WIDE 2FCs <![CDATA[f l the last layer]]> Avazu 1 Embedding Layer + 2FCs <![CDATA[f l the last layer]]> BHI ResNet-18 <![CDATA[f l the last three modules]]> ModelNet ResNet-18 <![CDATA[f l the last three modules]]>

[0158] (1.2) Baseline Methods: To comprehensively demonstrate the performance of the proposed ISSVFL in this invention, ISSVFL was compared with the following advanced methods in the VFL environment with scarce labels.

[0159] A. FedCVT is a semi-supervised learning method that enhances the performance of the VFL model in scenarios with limited aligned and labeled samples by augmenting the training dataset through representation estimation and pseudo-label prediction. FedCVT is designed specifically for the VFL scenario of two participating parties.

[0160] B. FedHSSL is a self-supervised learning framework that can utilize all available samples of each participant and achieve SOTA performance in scenarios with scarce labels. FedHSSL integrates three representative CSSL methods, including BYOL, MoCo, and SimSiam, for pre-training the VFL models of participants. According to the different integrated CSSL methods, FedHSSL is extended into three different baseline models, namely FedHSSL-BYOL, FedHSSL-MoCo, and FedHSSL-Simsiam.

[0161] For fair comparison, the proposed ISSVFL and all baseline methods use the same number of aligned and labeled samples in both the pre-training and fine-tuning stages. In the downstream supervised task, the number of aligned and labeled samples for fine-tuning ranges from 200 to 1000, and the average value of five runs is taken as the experimental result.

[0162] (1.3) Evaluation metrics: For the NUS-WIDE dataset and the ModelNet dataset, the Top-1 accuracy of the classifier held by the active party is used in this experiment to evaluate the final performance. For the Avazu dataset, the area under the ROC curve (AUC) is used as the evaluation metric in this experiment, and the smaller the AUC value, the worse the performance. For the BHI dataset, the F1-Score is used as the evaluation metric in this experiment.

[0163] (1.4) Data augmentation: For the NUS-WIDE dataset, 30% of the features are distorted in this experiment by replacing the original values with random values. For the Avazu dataset, the continuous features are processed in the same way as the NUS-WIDE dataset, and the categorical features are replaced by additional untrained embedding vectors. For the BHI dataset and the ModelNet dataset, they are randomly resized and then cropped to the input size, directly normalized. When the input size is 32, Gaussian blur is not performed; if the size is greater than 32, Gaussian blur is performed with a 50% probability.

[0164] (1.5) Implementation details: In ISSVFL, each participating party uses all its local samples for local training. At the same time, in this experiment, 40% of the dataset samples are used as alignment samples for collaborative training. Collaborative training and local training are carried out alternately, and each training method can be carried out for multiple cycles to minimize the communication cost. In this experiment, the number of pre-training rounds for all datasets is set to 40. The local encoders of ISSVFL and all baselines use the same model architecture for all datasets. In the fine-tuning stage, this experiment explored the performance of the joint VFL model after fine-tuning at learning rates of 0.025 and 0.01 respectively, and selected the best results for presentation. Each MLP of the DBP module consists of two FCs. For the Avazu dataset, the DBP module is deprecated in this experiment. For all datasets, the hyperparameter λ is set to 0.5. For the NUS-WIDE and Avazu datasets, the hyperparameters μ and ω are set to 2.0 and 0.03 respectively, and for the BHI and ModelNet datasets, the hyperparameters μ and ω are fixed at 0.5 and 0.01 respectively.

[0165] (2) Experimental results.

[0166] In this experiment, the performance of ISSVFL integrated with BYOL, MoCo, and SimSiam (i.e., the baseline models ISSVFL-BYOL, ISSVFL-MoCo, and ISSVFL-SimSiam) and the baseline methods on four datasets was compared. The results are shown in Table 2.

[0167] Table 2 Performance comparison of different methods on four datasets

[0168]

[0169]

[0170] The two-party scenario in the table represents the scenario of two participating parties, and the four-party scenario represents the scenario of four participating parties.

[0171] Research shows that compared with FedCVT, all ISSVFL-based methods generally show significant improvements in performance on all datasets. For example, when using 200 aligned and labeled samples, the best-performing ISSVFL method improves the performance on the NUSWIDE dataset by 13.3%, on the Avazu dataset by 5.0%, and on the BHI dataset by 10.0% compared with FedCVT. In the case of extremely limited aligned and labeled samples, the missing features inferred by FedCVT are more vulnerable to noise interference, increasing the difficulty of accurately estimating the missing features and reducing the overall effect of semi-supervised methods.

[0172] The FedHSSL-based method exhibits robust performance in scenarios with scarce labels. However, FedHSSL simply applies the CSSL method and is insufficient to fully exploit the potential of the SSL framework in VFL. To address the above deficiencies, the proposed FCISL method in this invention further combines CSSL with instance-level similarity, solves the regularization conflict problem, and captures important inter-domain knowledge. In addition, the proposed DBP module in this invention can dynamically balance cross-domain and intra-domain knowledge, thereby generating more meaningful representations to further improve the performance of the joint VFL model. Experimental results show that when using 200 aligned and labeled samples, ISSVFL-SimSiam has a 2.5% performance improvement over FedHSSL-SimSiam on the NUS-WIDE dataset and a 1.6% performance improvement on the Avazu dataset. Similarly, the performance of ISSVFL-BYOL and ISSVFL-MoCo on the NUS-WIDE and Avazu datasets is also superior to that of FedHSSL-BYOL and FedHSSL-MoCo respectively.

[0173] The study found that when using more aligned and labeled samples in the fine-tuning stage, the performance of ISSVFL on the BHI dataset improved significantly. It can be seen that the FCISL method of ISSVFL improves the performance of the joint VFL model on image datasets while avoiding regularization conflicts with the CSSL method. In addition, the performance of ISSVFL on the ModelNet dataset is also excellent, verifying its adaptability to multi-party scenarios, that is, as the number of participating parties increases, ISSVFL can still effectively address the problem of scarce labels in VFL.

[0174] In most cases, when using more aligned and labeled samples, ISSVFL improves performance faster than the baseline method, indicating that it can capture inter-domain knowledge in the model representations of different participating parties and can transfer knowledge to the local network more efficiently.

[0175] (3) Ablation experiment.

[0176] To evaluate the contribution of each module of ISSVFL to the overall performance, this experiment compared different variants created after removing modules. Specifically, "-" indicates that no module was removed, "w / o C" indicates that the co-training step in the pre-training phase (i.e., step 1.1) was removed, "w / o F" indicates that the FCISL method in the pre-training phase was removed, "w / o D" indicates that the DBP module in the fine-tuning phase was removed, and "w / oDF" indicates that both the DBP module and the FCISL method were removed. This experiment conducted ablation experiments on ISSVFL-BYOL, ISSVFL-MoCo, and ISSVFL-SimSiam respectively, and used 200 aligned and labeled samples for fine-tuning. Table 3 lists the average results of each variant in five experiments. To minimize the impact of randomness, this experiment designated 30% of the dataset samples as aligned samples for collaborative training.

[0177] Table 3 Performance comparison of various variants on different datasets

[0178]

[0179] The results show that the method based on the complete ISSVFL always exhibits the highest accuracy, indicating that removing any one module will lead to a performance decline. In particular, it should be noted that removing the co-training step results in a significant performance decline on all datasets, indicating that the co-training step has a very important impact on the model performance. For each CSSL method, removing the DBP and FCISL methods both lead to a performance decline of about 2% on the NUS-WIDE dataset, thus verifying the effectiveness of the two modules.

[0180] Except for a slight decline in the performance of ISSVFL-MoCo on the Avazu dataset, the FCISL method significantly improves the final performance. It can be seen that in most ablation experiments, the FCISL method captures the implicit inter-domain knowledge from the model representations of different parties, thus enhancing the representation ability of the local model. The learnable DBP module also plays a positive role in improving the final performance. For each CSSL method, the variant with the DBP module removed has a slight performance decline on all datasets, indicating that the DBP module can dynamically balance inter-domain knowledge and intra-domain knowledge and provide representations rich in general information for downstream supervised tasks.

[0181] In another embodiment, a self-supervised vertical federated learning device based on instance similarity and dynamic balance pool is provided. The self-supervised vertical federated learning device based on instance similarity and dynamic balance pool includes K parties and a server. The K parties and a server jointly train a joint vertical federated learning model. The joint vertical federated learning model includes a teacher network, a local network, a dynamic balance pool module, and a classifier. Among the K parties, those with labels are active parties, and the rest are passive parties. The K parties and a server perform the following operations:

[0182] In the pre-training stage:

[0183] Each party uses the aligned samples to jointly train its own teacher network. During the joint training, the instance similarity is calculated based on the output of the teacher network, and the loss is calculated in combination with the instance similarity;

[0184] Each party uses its local private samples and its own teacher network to train its own local network;

[0185] Each party uploads its local network to the server. After the server aggregates the local networks, a global network is generated, and the global network is sent to each party as a new local network;

[0186] Each party and the server repeat the pre-training stage until the pre-training ends;

[0187] In the fine-tuning stage:

[0188] Each party respectively inputs the aligned and labeled samples into its own teacher network and local network, and uses the dynamic balance pool module to aggregate the intermediate representations output by the teacher network and the intermediate representations output by the local network to obtain its own final representation;

[0189] The passive party sends the final representation to the active party for aggregation to obtain an aggregated representation;

[0190] The active party uses the classifier to convert the aggregated representation into a predicted value, and trains the dynamic balance pool module and the classifier according to the predicted value and the label;

[0191] Each party and the server repeat the fine-tuning stage until the fine-tuning ends.

[0192] For the specific limitations of the self-supervised vertical federated learning device based on instance similarity and dynamic balance pool, reference can be made to the limitations of the self-supervised vertical federated learning method based on instance similarity and dynamic balance pool in the above text, which will not be elaborated here.

[0193] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0194] The above-described embodiments only express several implementation manners of the present invention, and the description is relatively specific and detailed. However, it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.

Claims

1. A self-supervised longitudinal federated learning method based on instance similarity and dynamic balancing pool, for K participants and a server to collaboratively train a joint longitudinal federated learning model, characterized in that: The joint longitudinal federated learning model includes a teacher network, a local network, a dynamic balancing pool module and a classifier. Among the K participants, the ones with labels are active parties, and the rest are passive parties. The self-supervised longitudinal federated learning method based on instance similarity and dynamic balancing pool includes a pre-training stage and a fine-tuning stage, wherein: The pre-training stage includes: Each participant uses the aligned samples to collaboratively train their own teacher network. During the collaborative training, the instance similarity is calculated based on the output of the teacher network, and the loss is calculated in combination with the instance similarity. Each participant uses local private samples and their own teacher network to train their own local network; Each participant uploads the local network to the server, which aggregates the local networks to generate a global network and sends the global network to each participant as a new local network. Repeat the pre-training phase until the pre-training is completed; The fine-tuning phase includes: Each participant inputs the aligned and labeled samples into their respective teacher network and local network, and uses the dynamic balancing pool module to aggregate the intermediate representations output by the teacher network and the intermediate representations output by the local network to obtain their respective final representations; The passive party sends the final representation to the active party for aggregation to obtain the aggregated representation; The active party uses the classifier to convert the aggregate representation into a prediction value, and trains the dynamic balancing pool module and classifier based on the prediction value and label; The fine-tuning phase is repeated until the fine-tuning is completed, and the trained joint longitudinal federated learning model is output.

2. The self-supervised longitudinal federated learning method based on instance similarity and dynamic balancing pool according to claim 1, characterized in that: The participants use the aligned samples to jointly train their respective teacher networks, including: The teacher network includes a teacher encoder and a teacher projector, the aligned sample is input into the teacher encoder to be converted into a teacher intermediate representation, and then the teacher intermediate representation is input into the teacher projector to be converted into a teacher final representation; Generate a first positive sample pair based on the teacher's final representation, and calculate the representation self-supervision loss based on the first positive sample pair; Calculate the instance similarity of each participant based on the teacher's intermediate representation, generate a second positive sample pair based on the instance similarity, and calculate the instance self-supervision loss based on the second positive sample pair; The representation self-supervision loss and instance self-supervision loss are aggregated as the co-training final loss, and the teacher network is updated according to the co-training final loss.

3. The self-supervised longitudinal federated learning method based on instance similarity and dynamic balancing pool according to claim 2, characterized in that: The generating a first positive sample pair based on the teacher's final representation, and calculating the representation self-supervision loss according to the first positive sample pair, includes: Take the final representation of the teacher of participant k Each participant will ultimately represent the teacher Upload to the server, and the server calculates the average teacher representation And the average teacher representation Distribute to all participants; Each participant receives the average teacher representation After that, the final representation of the teacher obtained based on the same sample conversion and the average teacher representation The match is taken as the first positive sample pair, and the self-supervised loss is calculated based on the first positive sample pair as the representation self-supervised loss.

4. The self-supervised longitudinal federated learning method based on instance similarity and dynamic balancing pool according to claim 2, characterized in that: The method of calculating instance similarity of each participant based on the teacher intermediate representation, generating a second positive sample pair according to the instance similarity, and calculating the instance self-supervision loss according to the second positive sample pair includes: Take the teacher's intermediate representation of participant k Construct the instance similarity matrix of participant k as follows: In the formula, (·) T is the transposition operation, ω is the temperature hyperparameter, and ||·||2 is the L2 norm; Get the instance similarity S of participant k k as follows: Where Remove_diag(·) represents the operation of removing diagonal elements in the matrix; Each participant will assign instance similarity S k Upload to the server, and the server calculates the similarity of the target instance And the target instance similarity Distribute to all participants; Each participant receives the target instance similarity Then, the instance similarity S obtained based on the same sample conversion is k Similarity to the target instance The match is taken as the second positive sample pair, and the self-supervised loss is calculated based on the second positive sample pair as the instance self-supervised loss.

5. The self-supervised longitudinal federated learning method based on instance similarity and dynamic balancing pool according to claim 1, characterized in that: Each participant uses a local private sample and a respective teacher network to train a respective local network, including: The local network is used as the local online network, and a contrastive self-supervised learning method is used to obtain the local target network based on the local online network; A data augmentation strategy is adopted to transform the local private sample into a first enhanced view and a second enhanced view; Inputting the first enhanced view into the local online network, and inputting the second enhanced view into the local target network, to generate a third positive sample pair; Inputting the second enhanced view into the local online network and the first enhanced view into the local target network to generate a fourth positive sample pair; Compute the self-supervised loss as a symmetric loss based on the third positive sample pair and the fourth positive sample pair; Inputting the first enhanced view into the local online network, and inputting the first enhanced view into the teacher network to generate a fifth positive sample pair; Inputting the second enhanced view into the local online network, and inputting the second enhanced view into the teacher network to generate a sixth positive sample pair; Calculate the self-supervision loss based on the fifth and sixth positive sample pairs as the comparative knowledge transfer loss; Combining symmetric loss and contrastive knowledge transfer loss to update the local online network.

6. The self-supervised longitudinal federated learning method based on instance similarity and dynamic balancing pool according to claim 1, characterized in that: Each participant uploads the local network to the server, and the server aggregates the local networks to generate a global network, including: The local network includes a local encoder and a local projector, and each participant divides the local encoder into a local bottom encoder and a local top encoder; Each participant uploads the local top encoder and local projector to the server; The server aggregates local top encoders to obtain a global top encoder, aggregates local projectors to obtain a global projector, and uses the global top encoder and the global projector as a global network.

7. The self-supervised longitudinal federated learning method based on instance similarity and dynamic balancing pool according to claim 6, characterized in that: The local encoder includes two fully connected layers, and the last layer of the local encoder is divided as a local top encoder and the remaining layers are divided as local bottom encoders; Alternatively, the local encoder includes one embedding layer and two fully connected layers, and the last layer of the local encoder is divided as a local top encoder, and the remaining layers are divided as local bottom encoders; Alternatively, the local encoder is a ResNet-18 network, and the last three modules of the local encoder are divided as a local top encoder, and the remaining layers are divided as a local bottom encoder.

8. The self-supervised longitudinal federated learning method based on instance similarity and dynamic balancing pool according to claim 1, characterized in that: The method of using the dynamic balancing pool module to aggregate the intermediate representation output by the teacher network and the intermediate representation output by the local network to obtain their respective final representations includes: Take the intermediate representation of the output of the participant k teacher network and the intermediate representation of the local network output of participant k Calculate the forgetting factor O k as follows: In the formula, represents the concatenation operation along the feature axis, MLP(·) is a multi-layer perceptron, and σ(·) is a sigmoid function; Calculate the update factor U k as follows: Extracting implicit information k as follows: Where φ(·) is the activation function, FC(·) is the fully connected layer, and ⊙ is the Hadamard product; The final representation is calculated as follows: In the formula, R k is the final representation obtained by participant k.

9. A self-supervised longitudinal federated learning device based on instance similarity and dynamic balancing pool, characterized in that: The self-supervised longitudinal federated learning device based on instance similarity and dynamic balancing pool includes K participants and a server. The K participants and the server jointly train a joint longitudinal federated learning model. The joint longitudinal federated learning model includes a teacher network, a local network, a dynamic balancing pool module and a classifier. The K participants with labels are active parties, and the rest are passive parties. The K participants and the server perform the following operations: During the pre-training phase: Each participant uses the aligned samples to collaboratively train their own teacher network. During the collaborative training, the instance similarity is calculated based on the output of the teacher network, and the loss is calculated in combination with the instance similarity. Each participant uses local private samples and their own teacher network to train their own local network; Each participant uploads the local network to the server, which aggregates the local networks to generate a global network and sends the global network to each participant as a new local network. Each participant and the server repeats the pre-training phase until the pre-training is completed; During the fine-tuning phase: Each participant inputs the aligned and labeled samples into their respective teacher network and local network, and uses the dynamic balancing pool module to aggregate the intermediate representations output by the teacher network and the intermediate representations output by the local network to obtain their respective final representations; The passive party sends the final representation to the active party for aggregation to obtain the aggregated representation; The active party uses the classifier to convert the aggregate representation into a prediction value, and trains the dynamic balancing pool module and classifier based on the prediction value and label; Each participant and the server repeatedly executes the fine-tuning phase until the fine-tuning is completed and outputs the trained joint longitudinal federated learning model.

Citation Information

Patent Citations

  • Federal learning method and system based on self-supervised learning

    CN116306969A

  • Longitudinal federal learning method and system for balancing survey data difference of parties

    CN116341688A

  • Federal semi-supervised medical image diagnosis method based on structure alignment and false label self-correction

    CN116958656A