Distributed collaborative learning method and device, equipment and medium

Through the cross-domain, cross-modal and cross-task personalized knowledge sharing methods of generative adversarial networks, the difficulty of knowledge sharing caused by multiple heterogeneity in distributed collaborative learning is solved, efficient personalized knowledge transfer and task adaptation are achieved, and learning flexibility and adaptability are improved.

CN120338049APending Publication Date: 2025-07-18TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510290632.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing distributed collaborative learning methods are difficult to achieve efficient personalized knowledge sharing among multiple heterogeneous clients, especially in the problems of gradient conflicts and knowledge requirements differences caused by domain, modal and task heterogeneity have not been effectively resolved.

Method used

The method of generating adversarial network is adopted, and the cross-domain feature representation transformation, cross-domain reconstruction and cross-task adaptation are achieved through the adversarial optimization of generators and discriminators, the differences in domain distribution are eliminated, the modal missing features are completed, and high-quality pseudo-samples are generated to expand the task data distribution.

Benefits of technology

Under the premise of privacy protection, efficient collaborative learning is achieved, which significantly improves the client's performance in diversified tasks, meets personalized needs, and improves the flexibility and adaptability of distributed collaborative learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338049A_ABST
    Figure CN120338049A_ABST
Patent Text Reader

Abstract

The distributed collaborative learning method and device, the equipment and the medium provided by the embodiment of the invention have the following beneficial effects: through cross-domain transformation feature representation between the clients, cross-modal reconstruction feature representation between the clients and cross-task adaptation between the clients, the cross-domain transformation feature representation between the clients, the cross-modal reconstruction feature representation between the clients and the cross-task adaptation between the clients are realized; the construction of the three customized generative adversarial networks adopts a training strategy that cross-domain transformation is firstly trained to ensure alignment of domain features, then cross-modal reconstruction is trained to complement missing modals, and finally cross-task adaptation is trained to optimize the personalized performance of a local task model. Cross-domain representation and cross-modal reconstruction representation are gradually transmitted between clients, and an efficient cross-client knowledge migration mechanism is formed. According to the method, field, modal and task differences between clients are effectively eliminated, and efficient personalized knowledge extraction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to a distributed collaborative learning method, device, equipment and medium. Background Art

[0002] Distributed Cooperative Learning (abbreviation: DCL, full name: Distributed Cooperative Learning) plays a key role in artificial intelligence systems, such as cross-device knowledge sharing, privacy protection and multi-task collaboration. However, distributed cooperative learning faces challenges such as data heterogeneity, modality missing, domain differences, communication limitations and privacy security. Therefore, it has attracted more and more researchers' attention and become an important research direction in the fields of artificial intelligence and edge computing.

[0003] As a popular DCL technology, federated learning realizes cross-client knowledge transfer by aggregating model parameters trained based on their respective local data on a central server, and avoids sharing local data. With the diversification of online services and user-generated content, customers from different fields usually only focus on specific modalities and tasks. The latest research attempts to extend federated learning to modality-task-agnostic scenarios, aiming to share knowledge among clients with personalized neural structures in multiple heterogeneous aspects such as domain, modality and task. A direct way is to divide the model into many sub-blocks according to the attributes such as the roles and structures of different neural layers, and then perform parameter integration among the common sub-blocks of different clients.

[0004] However, this type of method still has certain limitations on the local model structure, and in order to share knowledge as fully as possible, complex disentanglement strategies need to be designed to find more common sub-blocks. Another method divides the local model of each client into two components, one for learning domain-modal-task-agnostic global shared knowledge, which is regularly sent to the central server during the training process to achieve knowledge transfer through parameter integration. The other focuses on capturing local personalized domain-modal-task-related information to optimize local tasks, and the parameters of this part of the component are only updated locally and will not be shared with the central server. This type of method can better balance global knowledge sharing and local personalized learning, and allows each client to build a more flexible personalized neural structure, avoiding complex disentanglement calculations related to separating shared and private model components.

[0005] Although the above methods show a certain effect on knowledge sharing among multiple heterogeneous clients, however, they inherit the specific parameter integration mode of federated learning in the way of knowledge sharing. Due to the statistical data heterogeneity among different clients in federated learning, such as label distribution skew, gradient conflict problems may occur during parameter aggregation, leading to negative knowledge transfer and even worse results than individual clients learning independently. In the scenario of multiple heterogeneities, the significant non-statistical heterogeneities such as domain, modality, and task among clients further exacerbate the gradient conflict during parameter aggregation, making the negative knowledge transfer problem more serious. Moreover, heterogeneous clients also mean differences in their needs for external knowledge. In the knowledge sharing mode of parameter integration, the single global model after integration is distributed to all clients, making it difficult for clients to obtain personalized global knowledge that meets their own needs. Therefore, it is an urgent need in the current academic and industrial circles.

[0006] As can be seen from the above description, how to construct a new DCL framework to achieve efficient personalized knowledge sharing among multiple heterogeneous clients is a technical problem that those skilled in the art need to solve urgently. Summary of the Invention

[0007] In order to overcome the deficiency that the existing distributed collaborative learning methods cannot effectively achieve personalized knowledge sharing among multiple heterogeneous clients, the present invention proposes a distributed collaborative learning method, device, equipment and medium.

[0008] To achieve the above object, according to the first aspect of the present invention, an embodiment of the present invention provides a distributed collaborative learning method, which includes the following steps:

[0009] When the client j and the client i have data of different domains but the same modality, the generator of the client j generates a cross-domain feature representation mapped from the local original features of the client j and sends it to the discriminator of the client i. The discriminator of the client i calculates the cross-domain loss of the generator of the client j from the cross-domain feature representation mapped to the client i, and calculates the cross-domain loss of the discriminator of the client i from the cross-domain feature representation mapped to the client i and the local original feature representation of the client i. The cross-domain loss of the generator of the client j is sent to the generator of the client j for optimization, and the cross-domain loss of the discriminator of the client i is used to optimize the discriminator of the client i; the generator of the client i generates a cross-domain feature representation mapped to the client j from the local original feature representation of the client i and sends it to the discriminator of the client j. The discriminator of the client j calculates the cross-domain loss of the generator of the client i from the cross-domain feature representation mapped to the client j, and calculates the cross-domain loss of the discriminator of the client j from the cross-domain feature representation mapped to the client j and the local original feature representation of the client j. The cross-domain loss of the generator of the client i is sent to the generator of the client i for optimization, and the cross-domain loss of the discriminator of the client j is used to optimize the discriminator of the client j; where the domain represents the business domain and the modality represents the data type;

[0010] When the client i lacks data of the modality relative to the client j, the client i generates a reconstructed feature representation from the cross-domain feature representation mapped to the client i sent by the client j and sends it to the discriminator of the client j. The discriminator of the client j calculates the cross-modal loss of the generator of the client i from the reconstructed feature representation, and calculates the cross-modal loss of the discriminator of the client j from the reconstructed feature representation and the local corresponding missing modality feature representation of the client j. The cross-modal loss of the generator of the client i is sent to the generator of the client i for optimization, and the cross-modal loss of the discriminator of the client j is used to optimize the discriminator of the client j;

[0011] When the client j and the client i have data of different tasks, the client i receives the external feature representation from the client j, and the external feature representation includes the cross-domain feature representation and / or the reconstructed feature representation. The external feature representation and the local feature representation of the client i are average pooled to obtain the unlabeled sample feature representation and the labeled sample feature representation. The generator of the client i generates the pseudo-sample feature representation. The predictor of the client i performs prediction training and optimization according to the labeled sample feature representation. The discriminator of the client i performs classification training and optimization on the unlabeled sample feature representation, the labeled sample feature representation, and the pseudo-sample feature representation. The generator of the client i is optimized by the pseudo-sample feature representation and the labeled sample feature representation; the task represents the model training objective;

[0012] Among them, both i and j are natural numbers.

[0013] Optionally, when the client j and the client i have data in different domains but the same modality, the generator of the client j generates a cross-domain feature representation mapped to the client i from the local original data of the client j, and the formula is as follows:

[0014]

[0015] Wherein, represents the cross-domain feature representation of the generator of the client j mapped to the client i output, represents the local original feature representation of the client j, represents the label corresponding to the domain of the client i, represents the trainable parameters of the generator of the client j, m ∈ M i ∩M j represents the set of the same modalities shared by the client i and the client j;

[0016] The discriminator of the client i calculates the cross-domain loss of the generator of the client j from the cross-domain feature representation mapped to the client i, and the loss function formula is as follows:

[0017]

[0018] Wherein, represents the cross-domain loss of the generator of the client j, is the discriminator of the client i, are the trainable parameters of the discriminator of the client i;

[0019] The cross-domain loss of the discriminator of the client i is calculated from the cross-domain feature representation mapped to the client i and the local original feature representation of the client i, and the loss function formula is as follows:

[0020]

[0021] Wherein, is the discriminator of the client i of the cross-domain loss, is the local original feature representation of the client i.

[0022] Optionally, the client i generates a reconstructed feature representation from the cross-domain feature representation mapped to the client i sent by the client j, and the formula is as follows:

[0023]

[0024] Wherein, represents the reconstructed feature representation generated at time t, represents the sequence of historical generated reconstructed feature representations from time 0 to t - 1, Characterize the cross-domain feature representation mapped to client i sent by client j, Characterize the generator of client i with trainable parameters, m ∈ M i ∩M j Characterize the set of the same modalities shared by client i and client j, n ∈ M i \M j Characterize the set of modalities that client j has but client i lacks;

[0025] The discriminator of client j calculates the cross-modal loss of the generator of client i from the reconstructed feature representation, and the loss function formula is as follows:

[0026]

[0027] where, Characterize the cross-modal loss of the generator of client i l represents the sequence length of the reconstructed feature representation, Characterize the discriminator of client j, Characterize the trainable parameters of the discriminator of client j ;

[0028] The cross-modal loss of the discriminator of client j is calculated from the reconstructed feature representation and the missing modality feature representation corresponding to client j locally, and the loss function formula is as follows:

[0029]

[0030] where, Characterize the cross-modal loss of the discriminator of client j Characterize the missing modality feature representation corresponding to client j locally at time t.

[0031] Optionally, the generator of client i generates a pseudo-sample feature representation, and the formula is as follows:

[0032]

[0033] where, Characterize the pseudo-sample feature representation generated by the generator of client i Characterize the initialization noise representation randomly sampled from the standard normal distribution N(0, 1), Characterize the generator with trainable parameters;

[0034] The discriminator of client i classifies and trains and optimizes the unlabeled sample feature representation, the labeled sample feature representation, and the pseudo-sample feature representation, and the loss function formula is as follows:​​

[0035]

[0036] Among them, the discriminator characterizing client i cross-task loss of is the feature representation of the labeled samples of client i, is the feature representation of the unlabeled samples, is the feature representation of the pseudo-samples, is the discriminator of client i trainable parameters of;

[0037] The predictor of the client i is predicted, trained and optimized according to the feature representation of the labeled samples, and the loss function formula is as follows:

[0038]

[0039] Among them, characterizes the cross-task loss of the predictor T of client i i of is the labeled category of the labeled samples, is the predicted category of the unlabeled samples, is the predicted category of the pseudo-samples, is the predictor T of client i i trainable parameters of;

[0040] The generator of the client i is optimized by the feature representation of the pseudo-samples and the labeled sample representation, and the loss function formula is as follows:

[0041]

[0042] Among them, characterizes the cross-task loss of the generator of client i

[0043] According to the second aspect of the present invention, an embodiment of the present invention further provides a distributed collaborative learning device, including:

[0044] The cross - domain transformation module is used to control that when the client j and the client i have data of different domains but the same modality, the generator of client j generates a cross - domain feature representation mapped from the local original feature representation of client j and sends it to the discriminator of client i. The discriminator of client i calculates the cross - domain loss of the generator of client j from the cross - domain feature representation mapped to client i, and calculates the cross - domain loss of the discriminator of client i from the cross - domain feature representation mapped to client i and the local original feature representation of client i. The cross - domain loss of the generator of client j is sent to the generator of client j for optimization, and the cross - domain loss of the discriminator of client i is used to optimize the discriminator of client i; The generator of client i generates a cross - domain feature representation mapped to client j from the local original feature representation of client i and sends it to the discriminator of client j. The discriminator of client j calculates the cross - domain loss of the generator of client i from the cross - domain feature representation mapped to client j, and calculates the cross - domain loss of the discriminator of client j from the cross - domain feature representation mapped to client j and the local original feature representation of client j. The cross - domain loss of the generator of client i is sent to the generator of client i for optimization, and the cross - domain loss of the discriminator of client j is used to optimize the discriminator of client j; where the domain represents the business domain, and the modality represents the data type;

[0045] The cross - modality reconstruction module is used to control that when client i lacks modality data relative to client j, client i generates a reconstructed feature representation from the cross - domain feature representation mapped to client i sent by client j and sends it to the discriminator of client j. The discriminator of client j calculates the cross - modality loss of the generator of client i from the reconstructed feature representation, and calculates the cross - modality loss of the discriminator of client j from the reconstructed feature representation and the local corresponding missing - modality feature representation of client j. The cross - modality loss of the generator of client i is sent to the generator of client i for optimization, and the cross - modality loss of the discriminator of client j is used to optimize the discriminator of client j;

[0046] The cross - task adaptation module is used to control that when client j and client i have data of different tasks, client i receives the external feature representation from client j, where the external feature representation includes the cross - domain feature representation and / or the reconstructed feature representation, averages the external feature representation and the local feature representation of client i to obtain the unlabeled sample feature representation and the labeled sample feature representation. The generator of client i generates a pseudo - sample feature representation. The predictor of client i performs prediction training and optimization based on the labeled sample feature representation. The discriminator of client i performs classification training and optimization on the unlabeled sample feature representation, the labeled sample feature representation, and the pseudo - sample feature representation. The generator of client i is optimized by the pseudo - sample feature representation and the labeled sample feature representation; the task represents the model training objective; where i and j are both natural numbers.

[0047] Optionally, the cross-domain transformation module is used to control that when the client j and the client i have data in different domains but the same modality, the generator of the client j generates a cross-domain feature representation mapped to the client i from the local original data of the client j. The formula is as follows:

[0048]

[0049] Wherein, represents the generator of the client j The output cross-domain feature representation mapped to the client i, represents the local original feature representation of the client j, represents the label corresponding to the domain of the client i, represents the trainable parameters of the generator of the client j, m ∈ M i ∩M j represents the set of the same modality shared by the client i and the client j;

[0050] The discriminator of the client i calculates the cross-domain loss of the generator of the client j from the cross-domain feature representation mapped to the client i. The loss function formula is as follows:

[0051]

[0052] Wherein, represents the cross-domain loss of the generator of the client j, is the discriminator of the client i, are the trainable parameters of the discriminator of the client i;

[0053] The cross-domain loss of the discriminator of the client i is calculated from the cross-domain feature representation mapped to the client i and the local original feature representation of the client i. The loss function formula is as follows:

[0054]

[0055] Wherein, is the discriminator of the client i The cross-domain loss, is the local original feature representation of the client i.

[0056] Optionally, the cross-modal reconstruction module is used to control the client i to generate a reconstructed feature representation from the cross-domain feature representation mapped to the client i sent by the client j. The formula is as follows:

[0057]

[0058] Wherein, represents the reconstructed feature representation generated at time t, Characterize the historical generated reconstruction feature representation sequence from time 0 to t-1, Characterize the cross-domain feature representation mapped from client j to client i, Characterize the generator of client i The trainable parameters, m ∈ M i ∩M j Characterize the set of the same modalities shared by client i and client j, n ∈ M i \M j Characterize the set of modalities that client j has but client i lacks;

[0059] The discriminator of the client j calculates the cross-modal loss of the generator of client i from the reconstructed feature representation, and the loss function formula is as follows:

[0060]

[0061] Among them, Characterize the generator of client i The cross-modal loss, l represents the sequence length of the reconstructed feature representation, Characterize the discriminator of client j, Characterize the discriminator of client j The trainable parameters;

[0062] The cross-modal loss of the discriminator of client j is calculated from the reconstructed feature representation and the missing modality feature representation corresponding to client j locally, and the loss function formula is as follows:

[0063]

[0064] Among them, Characterize the discriminator of client j The cross-modal loss, Characterize the missing modality feature representation corresponding to client j locally at time t.

[0065] Optionally, the cross-task adaptation module is used to control the generator of client i to generate pseudo-sample feature representations, and the formula is as follows:

[0066]

[0067] Among them, Characterize the generator of client i The generated pseudo-sample feature representation, Characterize the initialized noise representation randomly sampled from the standard normal distribution N(0,1), Characterize the generator The trainable parameters;

[0068] The discriminator of the client i classifies and trains the feature representations of unlabeled samples, labeled samples, and pseudo-samples and optimizes them. The loss function formula is as follows:

[0069]

[0070] where, represents the cross-task loss of the discriminator of client i , is the feature representation of labeled samples of client i, is the feature representation of unlabeled samples, is the feature representation of pseudo-samples, is the discriminator of client i 's trainable parameters;

[0071] The predictor of the client i performs prediction training and optimization based on the feature representation of labeled samples. The loss function formula is as follows:

[0072]

[0073] where, represents the cross-task loss of the predictor T of client i i , is the labeled category of labeled samples, is the predicted category of unlabeled samples, is the predicted category of pseudo-samples, is the predictor T of client i i 's trainable parameters;

[0074] The generator of the client i is optimized by the feature representation of pseudo-samples and the representation of labeled samples. The loss function formula is as follows:

[0075]

[0076] where, represents the cross-task loss of the generator of client i .

[0077] According to the third aspect of the present invention, an embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the distributed collaborative learning method in any one of the above embodiments are implemented.

[0078] According to a fourth aspect of the present invention, an embodiment of the present invention further provides a storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the steps of the distributed collaborative learning method in any of the above embodiments.

[0079] As described above, a distributed collaborative learning method, device, equipment and medium provided by an embodiment of the present invention have the following beneficial effects: In view of the problem of difficult knowledge sharing caused by multiple heterogeneities of domain, modality and task in distributed collaborative learning, the embodiment of the present invention proposes three customized generative adversarial network models. Through the adversarial optimization of the generator and the discriminator, the domain distribution difference is eliminated, the missing modality features are complemented, and high-quality pseudo samples are generated to expand the task data distribution, so as to achieve efficient collaborative learning under the premise of privacy protection. Moreover, by generating pseudo sample feature representations highly relevant to the local task characteristics and combining high-confidence pseudo labels for personalized model optimization, the performance of the client in diverse tasks is significantly improved, which can fully meet personalized needs and effectively improve the flexibility of distributed collaborative learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 is a schematic flowchart of a distributed collaborative learning method provided by an embodiment of the present invention;

[0081] Figure 2 is a schematic flowchart of a cross-domain transformation provided by an embodiment of the present invention;

[0082] Figure 3 is a schematic flowchart of a cross-modal reconstruction provided by an embodiment of the present invention;

[0083] Figure 4 is a schematic flowchart of a cross-task adaptation provided by an embodiment of the present invention;

[0084] Figure 5 is a schematic structural diagram of a distributed collaborative learning device provided by an embodiment of the present invention;

[0085] Figure 6 is a schematic hardware structure diagram of an electronic device for executing the distributed collaborative learning method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] To enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0087] The distributed collaborative learning method provided by the embodiments of the present invention is based on a generative adversarial network. This method addresses the problem of multiple heterogeneities such as domain, modality, and task among clients, introduces a customized generative adversarial network, and through the adversarial mechanism of the generator and discriminator, realizes efficient cross-domain, cross-modal, and cross-task personalized knowledge sharing. This method can maintain good adaptability and robustness in complex heterogeneous environments such as limited data privacy, domain differences, modality absence, and task offset, providing an innovative solution for distributed collaborative learning.

[0088] Please refer to Figures 1 to 6 . It should be noted that the illustrations provided in this embodiment only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the illustrations, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0089] See Figure 1 , which is a schematic flowchart of a distributed collaborative learning method provided by the embodiments of the present invention. As Figure 1 shown, the embodiments of the present invention illustrate the process of the distributed collaborative learning method.

[0090] Step S101: Cross-domain transformation of feature representations between clients.

[0091] When clients j and i have data in different domains but the same modality, the generator of client j generates a cross-domain feature representation mapped to client i from the local original feature representation of client j and concurrently sends it to the discriminator of client i. The discriminator of client i calculates the cross-domain loss of the generator of client j from the cross-domain feature representation mapped to client i, and calculates the cross-domain loss of the discriminator of client i from the cross-domain feature representation mapped to client i and the local original feature representation of client i. The cross-domain loss of the generator of client j is sent to the generator of client j for optimization, and the cross-domain loss of the discriminator of client i is used to optimize the discriminator of client i.

[0092] In the embodiments of the present invention, the "domain" represents the business domain, and the "modal" represents the data type. There are domain differences among clients. In the embodiments of the present invention, data of the same modal in different domains are converted to solve the problem of domain distribution differences.

[0093] It should be noted that, for the convenience of describing the steps of the distributed collaborative learning method in the embodiments of the present invention, both i and j are natural numbers, representing different clients. Data interaction can be carried out between clients in one-to-one, one-to-many, many-to-one, many-to-many and other ways to implement the steps of the distributed collaborative learning method.

[0094] In an exemplary embodiment, the modal may include one or more of text, audio, and image, and the domain may include the business customer service domain, the social media domain, etc. Specifically, Client 1 (corresponding to j = 1) and Client 2 (corresponding to i = 2) jointly have text and audio data, that is, the same modal of the two clients is text and audio. However, the data of Client 1 is in the business customer service domain. The text is mainly the consultation questions and replies of customers, and the audio is the phone recordings of customers and customer service staff, usually containing formal language and professional terms; while the data of Client 2 is in the social media domain. The text is mainly the comments and posts of users, and the audio is daily conversations or video dubbing, usually with strong colloquialism, slang, and emotional expressions; that is, the two clients have data in different domains. In the embodiments of the present invention, Client 1 and the client belong to the business customer service domain and the social media domain respectively. To eliminate the obstacle of domain differences between Client 1 and Client 2 to knowledge transfer, in the embodiments of the present invention, a generator and a discriminator for cross-domain conversion are deployed on each client. Among them, the generator converts the original data in the local specific source domain into a de-identified feature vector representation in the target domain. Subsequently, the representations after cross-domain conversion are exchanged between clients, and this exchange depends on the shared modal between the two clients. For example, after the text and audio data of Client 1 and Client 2 are cross-domain converted, the representations generated by each will be sent to the other. Subsequently, each client calculates the losses of the discriminator and the generator using the local specific domain representation and the received external representation, and feeds back the loss of the generator to the sender. This process guides the generator of the sender to generate an external representation aligned with the local specific domain distribution.

[0095] In specific implementation, refer to Figure 2 , which is a schematic diagram of a cross-domain transformation process provided by the embodiments of the present invention. As shown in the figure, the embodiments of the present invention show the process of cross-domain transformation between Client 1 and Client 2. In the embodiments of the present invention, the generator is responsible for mapping the input feature representation from one source domain to another target domain to generate a feature representation consistent with the characteristics of the target domain; the discriminator is used to distinguish the generated cross-domain feature representation from the real feature representation of the target domain. Specifically, asFigure 2 As shown, Client 1 and Client 2 each contain their own generator and discriminator. Taking Client 1 as an example, the generator of Client 1 uses the local original feature representation of local text (characterized by symbol t) or audio (characterized by symbol a) data as input. The local original feature representation comes from the local data of Client 1. Specifically, for text data, the local original feature representation is obtained by using a pre-trained BERT model, and each word is encoded as a 768-dimensional embedding vector representation; for audio data, the local original feature representation is the frame-level Mel-frequency cepstral coefficients (MFCCs) extracted by the librosa library, and each sample contains a 40-dimensional acoustic feature representation; for image (characterized by symbol v) data, the local original feature representation is obtained by using a pre-trained ResNet-101 model, and each frame of image is encoded as a 512-dimensional deep feature vector representation. These local original feature representations contain rich semantic information in the original data.

[0096] Subsequently, under the guidance of the conditional information the generator transforms the local original feature representation of Client 1 into the domain of Client 2, obtaining a cross-domain feature representation which is passed to Client 2 and input into the discriminator of Client 2 together with the local original feature representation of Client 2 where the conditional information is a label used to distinguish different domains and has no semantics itself. Exemplarily, a vector of all 1s can be used to represent Domain 1, and a vector of all 2s can be used to represent Domain 2. Their role is to inform the generator to transform the input feature representation into the corresponding domain. In addition, the discriminator is essentially a classifier, and its goal is to distinguish as much as possible whether the input representation comes from local (i.e., Client 2) or the transformed representation passed from Client 1, and to guide the generator of Client 1 to improve its own domain transformation ability to generate cross-domain feature representations that the discriminator of Client 2 cannot distinguish. Similarly, the generator of Client 2 will transform the local original feature representation into a cross-domain feature representation consistent with the domain of Client 1 and hand over these features to the discriminator of Client 1 for evaluation. Through the adversarial optimization of the generator and the discriminator, the above cross-domain transformation can effectively eliminate the domain distribution differences between clients, laying a foundation for cross-domain knowledge sharing.

[0097] Through the adversarial optimization of the generator and the discriminator, the above cross-domain transformation can effectively eliminate the domain distribution differences between clients, laying a foundation for cross-domain knowledge sharing.

[0098] In summary, generally, client j and client i have data in different domains but of the same modality. For the local original feature representation of client j obtained from the local data of client j and the conditional information of the target domain (i.e., the label corresponding to the domain of client i), the generator of client j will map from the domain distribution of client j to the domain distribution of client i, and its generation process is as follows:

[0099]

[0100] wherein, represents the cross - domain feature representation after cross - domain transformation mapped to client i by the generator of client j , represents the trainable parameters of the generator of client j, m ∈ M i ∩M j represents the same modality shared by client i and client j.

[0101] The generated cross - domain representation is sent to the target client i and, together with the local original feature representation of client i , is input into the discriminator of client i The discriminator updates its parameters through the binary cross - entropy loss function. The cross - domain loss function of the discriminator of client i is defined as follows:

[0102]

[0103] wherein, is the cross - domain loss of the discriminator of client i , is the trainable parameter of the discriminator . The cross - domain loss function of this discriminator of client i is used to update the parameters of the discriminator to achieve optimization.

[0104] Meanwhile, the discriminator also passes the supervision information to the generator of client j The cross - domain loss function of the generator of client j is defined as follows:

[0105]

[0106] wherein, represents the loss of the generator of client j , is the discriminator of client i are the trainable parameters of the client i discriminator, which are returned to the generator of client j for optimizing the generator Through this adversarial learning mechanism, the generator of client j gradually generates an external representation that is closer to the distribution of the target domain (the domain of client i), thereby improving the quality and adaptability of cross-domain representations and providing support for cross-domain knowledge sharing.

[0107] Similarly, the generator of client i generates a cross-domain feature representation mapped from the local original feature representation of client i to client j and sends it to the discriminator of client j. The discriminator of client j calculates the cross-domain loss of the generator of client i from the cross-domain feature representation mapped to client j and the local original feature representation of client j, and calculates the cross-domain loss of the discriminator of client j from the cross-domain feature representation mapped to client j and the local original feature representation of client j. The cross-domain loss of the generator of client i is sent to the generator of client i for optimization, and the cross-domain loss of the discriminator of client j is used to optimize the discriminator of client j. For the specific process, refer to the description of the above embodiments and will not be elaborated here.

[0108] Step S102: Cross-modal reconstruction feature representation between clients.

[0109] Modal missing means that in a client, data of one or more specific modalities are unavailable or not collected, resulting in the missing of information and the incompleteness of the processing process. In response to this situation, the adversarial network is focused on solving the problem of modal missing between clients.

[0110] In an exemplary embodiment, when the client i lacks data of a modality relative to the client j, the client i generates a reconstructed feature representation from the cross-domain feature representation mapped to the client i sent by the client j and sends it to the discriminator of the client j. The discriminator of the client j calculates the cross-modal loss of the generator of the client i from the reconstructed feature representation, and calculates the cross-modal loss of the discriminator of the client j from the reconstructed feature representation and the local corresponding missing modality feature representation of the client j. The cross-modal loss of the generator of the client i is sent to the generator of the client i for optimization, and the cross-modal loss of the discriminator of the client j is used to optimize the discriminator of the client j. For example, the client 2 (i.e., corresponding to j = 2) has image data, while the client 1 (i.e., corresponding to i = 1) lacks it. In this case, the client 1 cannot directly process and analyze the data of the image modality, thus limiting the performance and application scope of its model. Therefore, the embodiments of the present invention deploy a generator and a discriminator to learn the potential relationship between different modality data, so as to be able to reconstruct the missing modality information from the data of the existing modality. A corresponding generator is deployed on the client 1 lacking the image modality. The generator generates a cross-modal reconstructed image representation based on the external feature representation received from the client 2 before, and sends the reconstructed feature representation back to the client 2. The client 2 calculates the loss functions of the discriminator of the client 2 and the generator of the client 1 using the local real image representation and the received reconstructed feature representation, and returns the generator loss of the client 1 to the sender (i.e., the client 1) to guide the sender to generate a more realistic reconstructed representation.

[0111] See Figure 3 , which is a schematic diagram of a cross-modal reconstruction process provided by the embodiments of the present invention. As shown in the figure, the embodiments of the present invention illustrate the process of reconstructing cross-modal feature representations between clients.

[0112] In specific implementation, the generator generates a reconstructed feature representation of the missing modality based on the shared modality feature and the external condition information; the discriminator is used to verify whether the generated reconstructed feature representation is consistent with the real modality feature representation. As Figure 3 shown, the client 1 lacks image data, while the client 2 contains image data. By learning the potential relationship between different modalities on the client 2, the generator on the client 1 is trained to be able to reconstruct the image modality features from text (characterized by the symbol t) and audio (characterized by the symbol a) data. Specifically, the generator of the client 1 and respectively take the cross-domain feature representations and transmitted by the client 2 as the input, and take the condition information and These generated reconstructed feature representations are then passed to Client 2 and input into the discriminator of Client 2 together with the local corresponding missing modality feature representation of Client 2 (i.e., the local real image feature representation). and input into the discriminator of Client 2 together and The above discriminator validates these reconstructed feature representations to distinguish the reconstructed image features from the real image feature representation of Client 2. Similarly, the feedback of the discriminator is passed to the generator of Client 1 through the cross-modal loss and to guide the generator to optimize the generation process, thereby generating reconstructed feature representations that are closer to the real image feature distribution.

[0113] In summary, the embodiments of the present invention can generate semantically aligned complementary modality features through an adversarial learning mechanism, improve the modality feature space of the client, and achieve cross-modal collaborative learning. Generally, when there is a modality difference between Client i and Client j, Client i lacks modality n, while Client j has modality n, then the generator of Client i generates semantically aligned cross-modal reconstructed feature representations based on the external representation from Client j, and its generation process is as follows: generates semantically aligned cross-modal reconstructed feature representations based on the external representation from Client j, and its generation process is as follows:

[0114]

[0115] wherein, represents the cross-modal complementary reconstructed feature representation generated at time t, denotes the generator of Client i is the sequence of historically generated reconstructed feature representations from time 0 to t - 1, is the cross-domain feature representation mapped to Client i sent by Client j, represents the trainable parameters of the generator of Client i where m ∈ M i ∩M j represents the set of the same modalities shared by Client i and Client j, where n ∈ M i \M j represents the set of modalities that Client j has but Client i lacks.

[0116] Subsequently, the generated cross-modal complementary reconstructed feature representation is sent back to Client j and input into the discriminator of Client j together with the local corresponding missing modality feature representation of Client j and input into the discriminator of Client j together The task of the discriminator is to distinguish whether the input reconstructed feature representation is the local corresponding missing modality feature representation of Client j, and its loss function is defined as follows:

[0117]

[0118] Among them, is the cross-modal loss of client j discriminator l represents the sequence length of the reconstructed feature representation, represents the missing modal feature representation corresponding to client j locally at time t, represents the cross-domain feature representation mapped from client j and sent to client i, represents the discriminator of the trainable parameters. At the same time, the discriminator will pass the loss to the generator of client i The generator aims to generate a more realistic reconstructed feature representation for cross-modal completion, making it difficult to be distinguished by the discriminator The loss function is defined as follows:

[0119]

[0120] Among them, represents the cross-modal loss of the generator of client i , represents the discriminator of client j, and \(W_{(D_j)^{(m→n)}}\) represents the trainable parameters of the discriminator \(D_j^{(m→n)}\) of client j.

[0121] In this way, through the adversarial learning mechanism, the generator of client i and the discriminator of client j optimize each other, and the generator gradually generates a reconstructed feature representation for cross-modal completion that is closer to the real modality, thereby making up for the missing modal representation of client i, improving its modal feature space, and providing support for cross-modal collaborative learning.

[0122] Step S103: Cross-task adaptation between clients.

[0123] When clients j and i have data of different tasks, client i receives the external feature representation from client j. The external feature representation includes cross-domain feature representation and / or reconstructed feature representation. The external feature representation and the local feature representation of client i are averaged to obtain an unlabeled sample feature representation and a labeled sample feature representation. The generator of client i generates a pseudo-sample feature representation. The predictor of client i is predicted and trained and optimized based on the labeled sample feature representation. The discriminator of client i classifies and trains and optimizes the unlabeled sample feature representation, the labeled sample feature representation, and the pseudo-sample feature representation. The generator of client i is optimized by the pseudo-sample feature representation and the labeled sample feature representation.

[0124] In specific implementation, the cross-modal reconstructed feature representations are transmitted again between the clients. This transmission is still based on the situation that another client has the locally missing modality. Based on the description of the above embodiments, that is, the reconstructed feature representation of the reconstructed image modality is transmitted from Client 1 (corresponding to j = 1) to Client 2 (corresponding to i = 2). The difference is that in this step, Client 1 generates a cross-modal reconstructed image representation based on its locally transformed text and audio representations and sends it to Client 2. The purpose of this step is to enable Client 2 to supplement the missing modality in the external representation, align the domains and modality distributions of the local and external representations, so that Client 2 can establish a unified model and use both the local and external representations for learning specific tasks. Further, due to the task shift between the clients, the training objectives of the task representation models are different. For example, the objective of Client 1 is to train an emotion recognition model, while the objective of Client 2 is to train a sentiment polarity regression model. Therefore, their local data has different artificial labels. In the embodiments of the present invention, Client j sets a generator, a discriminator, and a predictor for the local specific task, aiming to improve the performance of the predictor on the local task through adversarial training between the generator, the discriminator, and the predictor. Taking Client 2 as an example, the external representations received from other clients are regarded as unlabeled samples, while the local representations are regarded as labeled samples for training the sentiment polarity regression model, and semi-supervised learning is carried out based on these samples.

[0125] The generative adversarial network in the embodiments of the present invention is used to improve the task adaptation ability of the clients. The generator is responsible for generating realistic and large-margin pseudo-samples to expand the task data distribution of the clients; the discriminator is used to distinguish these generated samples from the real samples. In addition, the local predictor not only generates high-confidence pseudo-labels using the generated pseudo-samples, but also adjusts the margin to optimize the model, improve the classification accuracy of real samples, and reduce the classification margin of false samples.

[0126] See Figure 4 , which is a schematic diagram of a cross-task adaptation process provided by the embodiments of the present invention. As shown in the figure, Client 2 can receive the reconstructed feature representations of one or more clients. In the embodiments of the present invention, the external feature representations of Client 1 and Client 3 received by it and are regarded as unlabeled samples. Among them, represents the cross-domain feature representation mapped to Client 2 sent by Client 1 (the symbols t and a in it refer to the descriptions of the above embodiments and correspond to the text object and the audio object respectively), represents the reconstructed feature representation sent by Client 1 (the symbols t, a, and v in it represent the text object, the audio object, and the image object respectively), Characterize the cross - domain feature representation sent by client 3 and mapped to client 2 Characterize the reconstructed feature representation sent by client 3, while the local feature representation of client 2 Is regarded as the labeled samples with artificial annotations for training the sentiment polarity regression model, and semi - supervised learning is carried out based on these samples. First, the feature representation sequences of different modalities are averaged in the column direction, and then they are concatenated together to obtain the unlabeled sample feature representation And the labeled sample feature representation The generator of client 2 then generates the pseudo - sample feature representation Their vector dimensions are the same.

[0127] Subsequently, these feature representations are fed into the predictor T2 and discriminator of client 2 For prediction. T2 performs sentiment polarity regression on the feature vectors of the labeled samples To optimize the prediction performance of the model. Meanwhile, the discriminator Classifies the feature representations of the real samples And the generated pseudo - sample feature representations To determine whether they are real samples. Such a setting aims to improve the generalization ability of the model on unlabeled data, and at the same time, further enhance the training effect of the classifier through the high - quality pseudo - samples generated by the generator The whole process forms a semi - supervised learning framework. In this way, client 2 can effectively utilize the information from other clients and the local labeled data to jointly promote the performance improvement of the model in the sentiment polarity task.

[0128] To sum up, generally, the generator of client i Generates pseudo - sample feature representations related to the real representation from random noise, and its formula is as follows:

[0129]

[0130] Where Characterize the pseudo - sample feature representation generated by the generator of client i μ * ~N(0,1) characterizes the initialization noise representation randomly sampled from the standard normal distribution, Characterize the generator Of the trainable parameters.

[0131] The generated pseudo - sample feature representation Is input into the discriminator In, and the task of the discriminator Is to distinguish whether the input features are generated by the generator Generate and optimize through the following binary cross - entropy loss function:

[0132]

[0133] Among them, Characterize the discriminator of client i The cross - task loss of, Is the feature representation of the labeled samples local to client i, Is the feature representation of the unlabeled samples, Characterize the generator of client i The feature representation of the pseudo - samples generated, while Is the discriminator of client i The trainable parameters of.

[0134] Predictor T i Receives the feature representation of the pseudo - samples From the generator And the feature representation of the real samples (i.e., the feature representation of the labeled samples), and is optimized in combination with the generated high - confidence pseudo - labels. Its goal is to improve the performance of the task model. Predictor T i The loss function of is defined as follows:

[0135]

[0136] Among them, Characterize the cross - task loss of predictor T of client i i Of, Is the labeled category of the labeled samples, And Are the predicted categories of the unlabeled samples and the pseudo - samples respectively, Is the predictor T of client i i The trainable parameters of.

[0137] Generator The optimization goal of is to deceive the discriminator So that the discriminator Is difficult to distinguish between pseudo - samples and labeled samples, and at the same time generate high - confidence predictions through predictor T i The loss function is defined as follows:

[0138]

[0139] Among them, Characterize the cross - task loss of the generator of client i Of, through the generator Discriminator And predictor T iThrough alternating optimization, the cross-task adaptation module can generate high-quality pseudo samples and expand the data distribution, thereby enhancing the client's adaptability and generalization performance in diverse tasks.

[0140] In addition, it should be noted that the generator, discriminator, and predictor involved in the embodiments of the present invention are all deep learning models. In the embodiments of the present invention, no specific limitations are imposed on the deep learning models used, and corresponding deep learning models can be selected according to specific usage scenario requirements.

[0141] To further illustrate the beneficial effects of the present invention, the embodiments of the present invention also provide development tools and frameworks: developed using Python 3.8.18 and PyTorch 2.0.0, each dataset is assigned to an independent client, and the model is deployed on a separate RTX 4090 GPU (24GB). Simulating a distributed environment in the real world, the representation transmission between clients is achieved through GPU communication.

[0142] Datasets: Three datasets from different domains, modalities, and tasks are selected. Assuming each client has a separate dataset to simulate the heterogeneous environment in multi-modal federated learning. Specifically, REST14: a single-modal dataset (only text), consisting of restaurant reviews, with annotation information for sentiment classification tasks. Each review involves one or more aspects (such as food, service, and environment), and each aspect is assigned a sentiment polarity label (negative, neutral, or positive); MELD: a multi-modal dataset (text and audio), consisting of multi-turn conversations, for sentiment recognition. Each conversation consists of multiple consecutive sentences, and each sentence is annotated as one of seven sentiment categories; MOSEI: a three-modal dataset (text, audio, and image), consisting of YouTube video clips, for speaker sentiment intensity regression tasks. Each video clip reflects the speaker's sentiment intensity towards a specific topic, ranging from [-3, +3], where -3 represents strong negative sentiment and +3 represents strong positive sentiment.

[0143] The verification results are shown in the following table:

[0144]

[0145] Data Preprocessing and Feature Extraction: Use the librosa library to extract frame-level acoustic features from the data, including 40-dimensional Mel-frequency cepstral coefficients (MFCCs); adopt a dual-network of Multi-task Cascaded Convolutional Networks (MTCNN) and ResNet-101, pre-trained from the AffectNet dataset, to extract 512-dimensional features related to the speaker's facial expressions and scene context; use the large language model BERT to extract the original text representation of each sample, with each word encoded as a 768-dimensional embedding vector, and fine-tune it through the loss of downstream tasks without an additional Transformer-based text encoder. Except for text, input the low-level features of each modality into a modality-specific encoder, which is constructed by multiple Transformer layers to capture the corresponding original information representation.

[0146] Optimization and Training: All losses are optimized using the Adam optimizer, and the hyperparameter β is set to (0.5, 0.999). The batch size for all datasets is uniformly set to 32. The experimental environment, through high-performance hardware and advanced model structures, ensures the efficient operation of the validation experiment in a distributed, multi-modal learning scenario.

[0147] Comparison Methods: This method, HeMuGAN, is compared with three types of baseline models. Among them, the local model is independently trained on each client and cannot achieve knowledge sharing; the traditional federated model only addresses the issue of statistical heterogeneity and collaborates through parameter aggregation of the same modality, but cannot adapt to non-statistical heterogeneity of modalities and tasks; the heterogeneous federated model solves the problems of multi-modal and task heterogeneity through contrastive learning (such as CreamFL) or meta-learning (such as Meta-HAR and CMMC) to promote cross-client collaboration. In contrast, this method, HeMuGAN, combines a generative adversarial network and a task adaptation mechanism, can effectively handle statistical and non-statistical heterogeneity, and provides a more powerful knowledge sharing ability and adaptability.

[0148] Validation Results: According to the analysis of the validation results in Table 1, this method, HeMuGAN, significantly outperforms existing traditional federated models and heterogeneous federated models, achieving higher performance on all datasets and metrics, with an average improvement of 8.72%, and improvements of 10.76%, 9.92%, and 6.26% on the REST14, MELD, and MOSEI datasets respectively. The traditional federated model cannot solve the problem of multiple heterogeneities, and the performance improvement of the heterogeneous federated model is limited. This method, HeMuGAN, successfully solves the problems of domain shift, modality gap, and task drift through de-identified intermediate representations and generative adversarial networks, achieving efficient knowledge sharing.

[0149] As can be seen from the description of the above embodiments, a distributed collaborative learning method provided by the embodiments of the present invention is as follows: when the client j and the client i have data in different domains and the same modality, the generator of the client j generates a cross-domain feature representation mapped to the client i from the local original feature representation of the client j and sends it to the discriminator of the client i. The discriminator of the client i calculates the cross-domain loss of the generator of the client j from the cross-domain feature representation mapped to the client i, and calculates the cross-domain loss of the discriminator of the client i from the cross-domain feature representation mapped to the client i and the local original feature representation of the client i. The cross-domain loss of the generator of the client j is sent to the generator of the client j for optimization, and the cross-domain loss of the discriminator of the client i is used to optimize the discriminator of the client i; the generator of the client i generates a cross-domain feature representation mapped to the client j from the local original feature representation of the client i and sends it to the discriminator of the client j. The discriminator of the client j calculates the cross-domain loss of the generator of the client i from the cross-domain feature representation mapped to the client j, and calculates the cross-domain loss of the discriminator of the client j from the cross-domain feature representation mapped to the client j and the local original feature representation of the client j. The cross-domain loss of the generator of the client i is sent to the generator of the client i for optimization, and the cross-domain loss of the discriminator of the client j is used to optimize the discriminator of the client j; wherein, the domain represents the business domain, and the modality represents the data type; when the client i lacks modal data relative to the client j, the client i generates a reconstructed feature representation from the cross-domain feature representation mapped to the client i sent by the client j and sends it to the discriminator of the client j. The discriminator of the client j calculates the cross-modal loss of the generator of the client i from the reconstructed feature representation, and calculates the cross-modal loss of the discriminator of the client j from the reconstructed feature representation and the corresponding missing modal feature representation of the client j locally. The cross-modal loss of the generator of the client i is sent to the generator of the client i for optimization, and the cross-modal loss of the discriminator of the client j is used to optimize the discriminator of the client j; when the client j and the client i have data for different tasks, the client i receives an external feature representation from the client j, and the external feature representation includes a cross-domain feature representation and / or a reconstructed feature representation. The external feature representation and the local feature representation of the client i are averaged to obtain an unlabeled sample feature representation and a labeled sample feature representation. The generator of the client i generates a pseudo-sample feature representation. The predictor of the client i performs prediction training and optimization based on the labeled sample feature representation. The discriminator of the client i performs classification training and optimization on the unlabeled sample feature representation, the labeled sample feature representation, and the pseudo-sample feature representation. The generator of the client i is optimized by the pseudo-sample feature representation and the labeled sample feature representation; the task represents the model training objective; wherein, both i and j are natural numbers.In the embodiments of the present invention, aiming at the problem of difficult knowledge sharing caused by multiple heterogeneities in distributed collaborative learning, such as domain, modality, and task, three customized generative adversarial network models are proposed. Through the adversarial optimization between the generator and the discriminator, the domain distribution differences are eliminated, the missing modality features are complemented, and high-quality pseudo-samples are generated to expand the task data distribution, so as to achieve efficient collaborative learning under the premise of privacy protection. Moreover, by generating pseudo-sample feature representations highly relevant to the local task characteristics and combining high-confidence pseudo-labels for personalized model optimization, the performance of the client in diverse tasks is significantly improved, which can fully meet personalized needs and effectively improve the flexibility of distributed collaborative learning.

[0150] Through the description of the above method embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes contributions to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0151] The embodiments of the present invention provide a non-volatile computer storage medium, and the computer storage medium stores computer-executable instructions, and these computer-executable instructions can execute the distributed collaborative learning method in any of the above method embodiments.

[0152] Corresponding to the method embodiments of the distributed collaborative learning provided by the present invention, the present invention also provides a distributed collaborative learning device.

[0153] See Figure 5 , which is a schematic structural diagram of a distributed collaborative learning device provided by the embodiments of the present invention. As shown in the figure, the device includes:

[0154] The cross - domain transformation module 11 is used to control that when the client j and the client i have data of different domains but the same modality, the generator of the client j generates a cross - domain feature representation mapped from the local original feature representation of the client j and sends it to the discriminator of the client i. The discriminator of the client i calculates the cross - domain loss of the generator of the client j from the cross - domain feature representation mapped to the client i, and calculates the cross - domain loss of the discriminator of the client i from the cross - domain feature representation mapped to the client i and the local original feature representation of the client i. The cross - domain loss of the generator of the client j is sent to the generator of the client j for optimization, and the cross - domain loss of the discriminator of the client i is used to optimize the discriminator of the client i; The generator of the client i generates a cross - domain feature representation mapped to the client j from the local original feature representation of the client i and sends it to the discriminator of the client j. The discriminator of the client j calculates the cross - domain loss of the generator of the client i from the cross - domain feature representation mapped to the client j, and calculates the cross - domain loss of the discriminator of the client j from the cross - domain feature representation mapped to the client j and the local original feature representation of the client j. The cross - domain loss of the generator of the client i is sent to the generator of the client i for optimization, and the cross - domain loss of the discriminator of the client j is used to optimize the discriminator of the client j; wherein, the domain represents the business domain, and the modality represents the data type;

[0155] The cross - modality reconstruction module 12 is used to control that when the client i lacks modality data relative to the client j, the client i generates a reconstructed feature representation from the cross - domain feature representation mapped to the client i sent by the client j and sends it to the discriminator of the client j. The discriminator of the client j calculates the cross - modality loss of the generator of the client i from the reconstructed feature representation, and calculates the cross - modality loss of the discriminator of the client j from the reconstructed feature representation and the corresponding missing modality feature representation of the client j locally. The cross - modality loss of the generator of the client i is sent to the generator of the client i for optimization, and the cross - modality loss of the discriminator of the client j is used to optimize the discriminator of the client j;

[0156] The cross-task adaptation module 13 is used to control that when the client j and the client i have data of different tasks, the client i receives the external feature representation from the client j. The external feature representation includes the cross-domain feature representation and / or the reconstruction feature representation. The external feature representation and the local feature representation of the client i are averaged to obtain the unlabeled sample feature representation and the labeled sample feature representation. The generator of the client i generates the pseudo-sample feature representation. The predictor of the client i performs prediction training and optimization based on the labeled sample feature representation. The discriminator of the client i performs classification training and optimization on the unlabeled sample feature representation, the labeled sample feature representation, and the pseudo-sample feature representation. The generator of the client i is optimized by the pseudo-sample feature representation and the labeled sample feature representation; the training objective of the task representation model; where i and j are both natural numbers.

[0157] In an exemplary embodiment, the cross-domain transformation module 11 is used to control that when the client j and the client i have data of different domains but the same modality, the generator of the client j generates the cross-domain feature representation mapped to the client i from the local raw data of the client j. The formula is as follows:

[0158]

[0159] Where represents the generator of the client j The output cross-domain feature representation mapped to the client i, represents the local raw feature representation of the client j, represents the label corresponding to the domain of the client i, represents the trainable parameters of the generator of the client j, m ∈ M i ∩M j represents the common set of the same modality of the client i and the client j;

[0160] The discriminator of the client i calculates the cross-domain loss of the generator of the client j from the cross-domain feature representation mapped to the client i. The loss function formula is as follows:

[0161]

[0162] Where represents the cross-domain loss of the generator of the client j, is the discriminator of the client i, are the trainable parameters of the discriminator of the client i;

[0163] The cross-domain loss of the discriminator of the client i is calculated from the cross-domain feature representation mapped to the client i and the local raw feature representation of the client i. The loss function formula is as follows:

[0164]

[0165] Among them, is the cross-domain loss of the client i discriminator and is the local original feature representation of client i.

[0166] In one exemplary embodiment, the cross-modal reconstruction module 12 is used to control the client i to generate a reconstruction feature representation from the cross-domain feature representation mapped from the client j to the client i. The formula is as follows:

[0167]

[0168] Among them, represents the reconstruction feature representation generated at time t, represents the sequence of historical generated reconstruction feature representations from time 0 to t-1, represents the cross-domain feature representation mapped from the client j to the client i, represents the client i generator with trainable parameters, m ∈ M i ∩M j represents the set of the same modalities shared by the client i and the client j, n ∈ M i \M j represents the set of modalities that the client j has but the client i lacks;

[0169] The discriminator of the client j calculates the cross-modal loss of the client i generator from the reconstruction feature representation. The loss function formula is as follows:

[0170]

[0171] Among them, represents the cross-modal loss of the client i generator l represents the sequence length of the reconstruction feature representation, represents the discriminator of the client j, represents the client j discriminator with trainable parameters;

[0172] The cross-modal loss of the client j discriminator is calculated from the reconstruction feature representation and the missing modality feature representation corresponding to the client j locally. The loss function formula is as follows:

[0173]

[0174] Among them, represents the cross-modal loss of the client j discriminator and Characterize the missing modal feature representation corresponding to the client j locally at time t.

[0175] In one exemplary embodiment, the cross-task adaptation module 13 is used to control the generator of the client i to generate a pseudo-sample feature representation, and the formula is as follows:

[0176]

[0177] Where, Characterize the generator of the client i The generated pseudo-sample feature representation, Characterize the initialization noise representation randomly sampled from the standard normal distribution N(0,1), Characterize the generator Of the trainable parameters;

[0178] The discriminator of the client i classifies and trains and optimizes the unlabeled sample feature representation, the labeled sample feature representation, and the pseudo-sample feature representation, and the loss function formula is as follows:

[0179]

[0180] Where, Characterize the discriminator of the client i Of the cross-task loss, Is the labeled sample feature representation of the client i, Is the unlabeled sample feature representation, Is the pseudo-sample feature representation, Is the discriminator of the client i Of the trainable parameters;

[0181] The predictor of the client i performs prediction training and optimization based on the labeled sample feature representation, and the loss function formula is as follows:

[0182]

[0183] Where, Characterize the cross-task loss of the predictor T of the client i i Of, Is the annotation category of the labeled sample, Is the predicted category of the unlabeled sample, Is the predicted category of the pseudo-sample, Is the predictor T of the client i i Of the trainable parameters;

[0184] The generator of the client i is optimized by the pseudo-sample feature representation and the labeled sample representation, and the loss function formula is as follows:

[0185]

[0186] Among them, the generator characterizing client i of the cross-task loss.

[0187] Figure 6 is a schematic diagram of the hardware structure of an electronic device for executing the distributed collaborative learning method provided by an embodiment of the present invention, as Figure 6 shown, the device includes:

[0188] one or more processors 610 and a memory 620, Figure 5 taking one processor 610 as an example in

[0189] The device for executing the distributed collaborative learning method may further include: an input device 630 and an output device 640.

[0190] The processor 610, the memory 620, the input device 630, and the output device 640 may be connected by a bus or other means, Figure 6 taking connection by bus as an example in

[0191] The memory 620, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the distributed collaborative learning method in the embodiments of the present invention (for example, the cross-domain transformation module 11, the cross-modal reconstruction module 12, and the cross-task adaptation module 13 shown in Figure 5 ). The processor 610 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 620, that is, implements the charging battery state of charge estimation method in the above method embodiments.

[0192] The memory 620 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the distributed collaborative learning device, etc. In addition, the memory 620 may include a high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 620 may optionally include a memory remotely set relative to the processor 610, and these remote memories may be connected to the processing device of the distributed collaborative learning through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0193] The input device 630 can receive input digital or character information and generate key signal inputs related to user settings and function controls of the processing device for distributed collaborative learning. The output device 640 can include display devices such as a display screen.

[0194] The one or more modules are stored in the memory 620 and, when executed by the one or more processors 610, perform the distributed collaborative learning method in any of the above method embodiments.

[0195] The above product can execute the method provided in the embodiments of the present invention and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided in the embodiments of the present invention.

[0196] Of course, Figure 5 and Figure 6 The devices and equipment shown are only exemplary embodiments. In specific implementations, they can be understood as devices or equipment configured on any client, enabling interaction between each client, thereby achieving personalized knowledge sharing among multiple heterogeneous (in terms of domain, modality, and task) clients.

[0197] The electronic devices in the embodiments of the present invention exist in various forms, including but not limited to:

[0198] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones (such as iPhone), multimedia phones, functional phones, and low-end phones, etc.

[0199] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc., such as iPad.

[0200] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players (such as iPod), handheld game consoles, e-books, and smart toys and portable in-vehicle navigation devices.

[0201] (4) Servers: Devices that provide computing services. The composition of a server includes a processor, hard disk, memory, system bus, etc. Servers are similar to general computer architectures, but due to the need to provide highly reliable services, they have higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability.

[0202] (5) Other electronic devices with data interaction functions.

[0203] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0204] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for device or system embodiments, since they are basically similar to method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the partial description of the method embodiments. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0205] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0206] The above description is only the specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A distributed collaborative learning method, characterized in that, Including: When clients j and i have data of different domains but the same modality, the generator of client j generates a cross-domain feature representation mapped to client i from the local original feature representation of client j and sends it to the discriminator of client i. The discriminator of client i calculates the cross-domain loss of the generator of client j from the cross-domain feature representation mapped to client i, and calculates the cross-domain loss of the discriminator of client i from the cross-domain feature representation mapped to client i and the local original feature representation of client i. The cross-domain loss of the generator of client j is sent to the generator of client j for optimization, and the cross-domain loss of the discriminator of client i is used to optimize the discriminator of client i; The generator of client i generates a cross-domain feature representation mapped to client j from the local original feature representation of client i and sends it to the discriminator of client j. The discriminator of client j calculates the cross-domain loss of the generator of client i from the cross-domain feature representation mapped to client j, and calculates the cross-domain loss of the discriminator of client j from the cross-domain feature representation mapped to client j and the local original feature representation of client j. The cross-domain loss of the generator of client i is sent to the generator of client i for optimization, and the cross-domain loss of the discriminator of client j is used to optimize the discriminator of client j; wherein, the domain represents the business domain, and the modality represents the data type; When client i lacks modality data relative to client j, client i generates a reconstructed feature representation from the cross-domain feature representation mapped to client i sent by client j and sends it to the discriminator of client j. The discriminator of client j calculates the cross-modal loss of the generator of client i from the reconstructed feature representation, and calculates the cross-modal loss of the discriminator of client j from the reconstructed feature representation and the local corresponding missing modality feature representation of client j. The cross-modal loss of the generator of client i is sent to the generator of client i for optimization, and the cross-modal loss of the discriminator of client j is used to optimize the discriminator of client j; When clients j and i have data of different tasks, client i receives an external feature representation from client j, where the external feature representation includes a cross-domain feature representation and / or a reconstructed feature representation, averages the external feature representation and the local feature representation of client i to obtain an unlabeled sample feature representation and a labeled sample feature representation. The generator of client i generates a pseudo-sample feature representation. The predictor of client i performs prediction training and optimization based on the labeled sample feature representation. The discriminator of client i performs classification training and optimization on the unlabeled sample feature representation, the labeled sample feature representation, and the pseudo-sample feature representation. The generator of client i is optimized by the pseudo-sample feature representation and the labeled sample feature representation; the task represents the model training objective; Wherein, both i and j are natural numbers.

2. The distributed collaborative learning method according to claim 1, wherein When clients j and i have data of different domains but the same modality, the generator of client j generates a cross-domain feature representation mapped to client i from the local original data of client j. The formula is as follows: Among them, The generator characterizing client j The cross - domain feature representation output and mapped to client i The local original feature representation characterizing client j The label characterizing the domain of corresponding client i The trainable parameters of the generator of client j, m ∈ M i ∩M j Represents the set of the same modalities shared by client i and client j; The discriminator of the client i calculates the cross - domain loss of the generator of the client j from the cross - domain feature representation mapped to the client i. The formula of the loss function is as follows: Among them, represents the cross-domain loss of the client j generator, is the discriminator of client i, are the trainable parameters of the client i discriminator; The cross - domain loss of the discriminator of the client i is calculated from the cross - domain feature representation mapped to the client i and the local original feature representation of the client i. The formula of the loss function is as follows: Among them, is the cross-domain loss of the client i discriminator and is the local original feature representation of client i.

3. The distributed collaborative learning method according to claim 1, wherein The client i generates a reconstructed feature representation from the cross - domain feature representation mapped to the client i sent by the client j. The formula is as follows: wherein, represents the reconstructed feature representation generated at time t, represents the sequence of historical generated reconstructed feature representations from time 0 to t-1, represents the cross-domain feature representation mapped from client j to client i, represents the generator of client i with trainable parameters, m ∈ M i ∩M j represents the set of the same modalities shared by client i and client j, n ∈ M i \M j represents the set of modalities that client j has but client i lacks; The discriminator of the client j calculates the cross - modal loss of the generator of the client i from the reconstructed feature representation. The formula of the loss function is as follows: Among them, characterizes the client i generator of the cross-modal loss, l represents the sequence length of the reconstructed feature representation, characterizes the discriminator of client j, characterizes the client j discriminator of the trainable parameters; The cross - modal loss of the discriminator of the client j is calculated from the reconstructed feature representation and the corresponding missing - modality feature representation of the client j locally. The formula of the loss function is as follows: Among them, characterize the cross-modal loss of the client j discriminator and characterize the missing modal feature representation corresponding to the client j locally at time t.

4. The distributed collaborative learning method according to claim 1, characterized in that, The generator of the client i generates a pseudo - sample feature representation. The formula is as follows: Among them, The generator characterizing client i The generated pseudo-sample feature representation, Characterizes the initialization noise representation randomly sampled from the standard normal distribution N(0, 1), Characterizes the generator The trainable parameters of The discriminator of the client i classifies and trains on the unlabeled sample feature representation, labeled sample feature representation, and pseudo - sample feature representation for optimization. The formula of the loss function is as follows: Among them, represents the discriminator of client i of the cross-task loss, is the feature representation of the labeled samples of client i, is the feature representation of the unlabeled samples, is the feature representation of the pseudo-samples, is the discriminator of client i of the trainable parameters; The predictor of the client i predicts and trains based on the labeled sample feature representation for optimization. The formula of the loss function is as follows: Among them, represents the cross-task loss of the predictor T of client i i , is the labeled category of the labeled samples, is the predicted category of the unlabeled samples, is the predicted category of the pseudo-samples, is the trainable parameter of the predictor T of client i i ; The generator of the client i is optimized by the pseudo - sample feature representation and the labeled sample representation. The formula of the loss function is as follows: Among them, represents the generator of client i of the cross-task loss.

5. A distributed collaborative learning device, characterized in that, Including: A cross - domain transformation module, which is used to control that when the client j and the client i have data of different domains but the same modality, the generator of the client j generates a cross - domain feature representation mapped to the client i from the local original feature representation of the client j and sends it to the discriminator of the client i. The discriminator of the client i calculates the cross - domain loss of the generator of the client j from the cross - domain feature representation mapped to the client i, and calculates the cross - domain loss of the discriminator of the client i from the cross - domain feature representation mapped to the client i and the local original feature representation of the client i. The cross - domain loss of the generator of the client j is sent to the generator of the client j for optimization, and the cross - domain loss of the discriminator of the client i is used to optimize the discriminator of the client i; The generator of the client i generates a cross - domain feature representation mapped to the client j from the local original feature representation of the client i and sends it to the discriminator of the client j. The discriminator of the client j calculates the cross - domain loss of the generator of the client i from the cross - domain feature representation mapped to the client j, and calculates the cross - domain loss of the discriminator of the client j from the cross - domain feature representation mapped to the client j and the local original feature representation of the client j. The cross - domain loss of the generator of the client i is sent to the generator of the client i for optimization, and the cross - domain loss of the discriminator of the client j is used to optimize the discriminator of the client j; where the domain represents the business domain, and the modality represents the data type; The cross-modal reconstruction module is used to control that when client i lacks modal data relative to client j, client i generates a reconstructed feature representation from the cross-domain feature representation mapped to client i sent by client j and sends it to the discriminator of client j. The discriminator of client j calculates the cross-modal loss of the generator of client i from the reconstructed feature representation, and calculates the cross-modal loss of the discriminator of client j from the reconstructed feature representation and the corresponding missing modal feature representation locally at client j. The cross-modal loss of the generator of client i is sent to the generator of client i for optimization, and the cross-modal loss of the discriminator of client j is used to optimize the discriminator of client j; The cross-task adaptation module is used to control that when client j and client i have data of different tasks, client i receives the external feature representation from client j, and the external feature representation includes the cross-domain feature representation and / or the reconstructed feature representation. The external feature representation and the local feature representation of client i are average-pooled to obtain the unlabeled sample feature representation and the labeled sample feature representation. The generator of client i generates the pseudo-sample feature representation. The predictor of client i performs prediction training and optimization based on the labeled sample feature representation. The discriminator of client i performs classification training and optimization on the unlabeled sample feature representation, the labeled sample feature representation, and the pseudo-sample feature representation. The generator of client i is optimized by the pseudo-sample feature representation and the labeled sample feature representation; the training objective of the task representation model; where i and j are both natural numbers.

6. The distributed collaborative learning device according to claim 5, wherein The cross-domain transformation module is used to control that when client j and client i have data of different domains but the same modality, the generator of client j generates the cross-domain feature representation mapped to client i from the local raw data of client j, and the formula is as follows: Among them, The generator characterizing client j The cross-domain feature representation output and mapped to client i The local original feature representation characterizing client j The label characterizing the domain corresponding to client i The trainable parameters of the generator of client j, m ∈ M i ∩M j Indicates the set of the same modalities shared by client i and client j; The discriminator of client i calculates the cross-domain loss of the generator of client j from the cross-domain feature representation mapped to client i, and the loss function formula is as follows: Among them, represents the cross-domain loss of the client j generator, is the discriminator of client i, are the trainable parameters of the client i discriminator; The cross-domain loss of the discriminator of client i is calculated from the cross-domain feature representation mapped to client i and the local raw feature representation of client i, and the loss function formula is as follows: Among them, is the cross-domain loss of the client i discriminator and is the local original feature representation of client i.

7. The distributed collaborative learning device according to claim 5, wherein The cross-modal reconstruction module is used to control that client i generates a reconstructed feature representation from the cross-domain feature representation mapped to client i sent by client j, and the formula is as follows: Among them, represents the reconstructed feature representation generated at time t, represents the sequence of historical generated reconstructed feature representations from time 0 to t - 1, represents the cross - domain feature representation mapped from client j to client i, represents the generator of client i with trainable parameters, m ∈ M i ∩M j represents the set of the same modalities shared by client i and client j, n ∈ M i \M j represents the set of modalities that client j has but client i lacks; The discriminator of client j calculates the cross-modal loss of the generator of client i from the reconstructed feature representation, and the loss function formula is as follows: Among them, characterizes the client i generator of the cross-modal loss, l represents the sequence length of the reconstructed feature representation, characterizes the discriminator of client j, characterizes the client j discriminator of the trainable parameters; The cross-modal loss of the discriminator of client j is calculated from the reconstructed feature representation and the corresponding missing modal feature representation locally at client j, and the loss function formula is as follows: Among them, characterizes the cross-modal loss of the client j discriminator and characterizes the missing modal feature representation corresponding to the client j locally at time t.

8. The distributed collaborative learning device according to claim 5, wherein, The cross-task adaptation module is used to control that the generator of client i generates the pseudo-sample feature representation, and the formula is as follows: Among them, The generator characterizing client i The generated pseudo-sample feature representation, Characterizes the initialization noise representation randomly sampled from the standard normal distribution N(0, 1), Characterizes the generator Of the trainable parameters; The discriminator of client i performs classification training and optimization on the unlabeled sample feature representation, the labeled sample feature representation, and the pseudo-sample feature representation, and the loss function formula is as follows: Among them, The discriminator characterizing client i The cross-task loss of Is the feature representation of the labeled samples of client i, Is the feature representation of the unlabeled samples, Is the feature representation of the pseudo samples, Is the discriminator of client i The trainable parameters of The predictor of client i performs prediction training and optimization based on the labeled sample feature representation, and the loss function formula is as follows: Among them, represents the cross-task loss of the predictor T for client i, i where is the labeled category of the labeled samples, is the predicted category of the unlabeled samples, is the predicted category of the pseudo-samples, and i are the trainable parameters of the predictor T for client i. The generator of the client i is optimized by the pseudo-sample feature representation and the labeled sample representation, and the loss function formula is as follows: Among them, The generator characterizing client i Cross-task loss.

9. An electronic device, characterized in that, It includes a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps of the distributed collaborative learning method described in any one of claims 1 to 4 are implemented.

10. A computer-readable storage medium, characterized in that, At least one instruction, at least one segment of program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one segment of program, the code set or the instruction set is loaded and executed by the processor to implement the steps of the distributed collaborative learning method described in any one of claims 1 to 4.