Federal learning method and device for scene with inconsistent domain distribution, and medium

By adopting the method of clustering and domain prototype fusion in federated learning, the knowledge loss problem caused by inconsistent domain distribution is solved, and the good performance of the global model in multiple domains and the bias elimination of classifiers is achieved.

CN120278299APending Publication Date: 2025-07-08SOUTH CHINA UNIV OF TECH +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510238913.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In federated learning scenarios where domain distribution is inconsistent, it is difficult for the existing technology to effectively retain knowledge from multiple fields, resulting in poor performance of models on different feature distributions and classifiers are prone to bias.

Method used

Using a knowledge-protected model aggregation and a classifier retraining method based on domain prototypes, the client is divided into the same domain clusters through cluster federated learning, and a global feature extractor is used to form a global feature extractor, and a domain prototype fusion is used to train the global classifier to avoid model average and classifier bias.

Benefits of technology

It effectively retains knowledge in different fields, improves the discrimination and generalization capabilities of the model in multiple fields, and reduces communication and computing overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278299A_ABST
    Figure CN120278299A_ABST
Patent Text Reader

Abstract

The invention discloses a federated learning method and device for a scene with inconsistent domain distribution, and a medium, and the method comprises the steps: firstly, employing a clustering federated learning strategy in a feature extractor training stage of a global model, and carrying out the cooperative training in a cluster to eliminate the influence of domain deflection; and a global feature extractor is obtained by collecting feature extractors of the cluster models, so that the global model can reserve knowledge of different fields and has intra-field discrimination ability and inter-field generalization ability. Secondly, a classifier retraining scheme based on a domain prototype is designed for matching higher-dimensional features output by the global feature extractor, a client randomly fuses local features and the domain prototype to obtain virtual features, and a server trains a global classifier by using a virtual feature set uploaded by the client; according to the method, classifier prejudice caused by cooperative training is avoided, the communication and calculation overhead of the client is saved, and the robustness of model training to scenes with inconsistent field distribution is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular, to a federated learning method, device, and medium for a scenario with inconsistent domain distributions. Background Art

[0002] Centralized large-scale training data has enabled deep learning models to achieve remarkable results. However, in reality, due to the extensive use of mobile devices, the generated data is often scattered among these edge devices. Limited by privacy protection policies and the increasing concern for privacy, the server cannot require clients to upload local data for centralized model training. Federated Learning (FL) has been proposed to address this problem. It enables clients to collaboratively train a shared model without revealing the original private data.

[0003] Although FL has achieved great success, it still faces some research problems. A key and inherent challenge is data heterogeneity. In actual FL tasks, since private data is usually collected from different sources or scenarios, the data distributions of different clients may be non-independent and identically distributed (non-IID). Existing work has shown that due to the inconsistency of the local objective functions and optimization directions of clients caused by non-IID data, the performance of the averaged global model is much worse than that on IID data.

[0004] An important type of data heterogeneity is domain skew, where private data comes from different domains, resulting in different feature distributions. In this scenario, the optimization goal requires the global model to perform satisfactorily on multiple different feature distributions and avoid poor discrimination ability in a certain domain. However, the typical federated learning paradigm makes it difficult to achieve this goal. We know that federated learning can generally be regarded as an iterative process consisting of two steps, namely local update and model averaging; in the local update phase, the model is only trained on private data, which easily leads to overfitting to its own domain knowledge and forgetting the knowledge of other domains learned before. This catastrophic forgetting makes it difficult for different domain knowledge to be transmitted between clients. In addition, data heterogeneity can lead to differences in model parameters, and domain skew may exacerbate this effect, thus weakening the effect of model averaging. Therefore, the above two defects hinder the generalization ability of the global model across multiple domains.

[0005] In summary, a key problem under domain skew is: how to effectively retain the knowledge of multiple domains. Summary of the Invention

[0006] To at least to some extent solve one of the technical problems existing in the prior art, the purpose of the present invention is to provide a federated learning method, device and medium for the scenario of inconsistent domain distributions.

[0007] The first technical solution adopted by the present invention is:

[0008] A federated learning method for the scenario of inconsistent domain distributions, including two stages: model aggregation with knowledge protection and retraining of the classifier based on domain prototypes.

[0009] The model aggregation with knowledge protection includes:

[0010] A1: The client performs truncated singular value decomposition on the local data and uploads the decomposition result to the server.

[0011] A2: The server clusters the clients according to the principal angle similarity of the local data subspaces, and the clients with the same domain distribution are grouped into the same cluster.

[0012] A3: The clients in each cluster have the same domain distribution, and each cluster performs federated learning to train the global model within the cluster, and these models fit the specific domain distribution.

[0013] A4: Each cluster uploads the feature extractor part of the globally trained model to the server, and the server collects these feature extractors as the global feature extractor instead of using traditional model averaging, thereby retaining different domain knowledge.

[0014] A5: The server broadcasts the global feature extractor to the clients, and the clients fix the parameters of the global feature extractor, which is part of the final global model.

[0015] The retraining of the classifier based on domain prototypes includes:

[0016] B1: To process the data features output by the global feature extractor with higher dimensions, the server initializes a larger-scale global classifier.

[0017] B2: The client takes the average of the local features of the same class to obtain the local prototype of each class and uploads it to the server.

[0018] B3: The server takes the average of the local prototypes of the clients with the same domain distribution according to the clustering information and broadcasts the domain prototype to the clients.

[0019] B4: The client randomly mixes the local features with the domain prototypes of the same class to obtain a virtual feature set and uploads it to the server.

[0020] B5: The server uses the virtual features uploaded by the client to train a new global classifier, thereby eliminating classifier bias and imposing no additional communication and computational burden on the client.

[0021] Furthermore, in step A1, are the p most important left singular vectors U obtained by client m, where p is a hyperparameter; they span a data subspace as the principal basis and concisely capture the main features of the local data D m of the basic distribution.

[0022] Furthermore, the said step A2 includes:

[0023] The server constructs a similarity matrix A according to where the matrix element A i,j is the minimum principal angle between two data subspaces, and the expression is as follows:

[0024]

[0025] In the formula, represents the set of the p most important left singular vectors of client i, represents the set of the p most important left singular vectors of client j, and u i represents the vector in the set and u j represents the vector in the set ;

[0026] The server uses agglomerative hierarchical clustering on matrix A to obtain disjoint clusters where v is the number of clusters and β is a parameter of the clustering algorithm HC, representing the Euclidean distance between two clusters.

[0027] Furthermore, in step A4, the global feature extractor is defined as an ensemble of these cluster model feature extractors, that is where θ1,...,θ V are the parameters of the cluster model feature extractors, represents the model structure; the output of ε(Θ) is the concatenation of the features output by each cluster model feature extractor, that is where x represents an input sample, represents the concatenation operation and

[0028] Furthermore, the said step B1 includes:

[0029] The server initializes a global classifier which maps to a logits output where Φ are the parameters of the global classifier, represents the model structure, and I is the number of classes in the classification task.

[0030] Further, in step B2, the calculation formula for the k-th local prototype of client m is:

[0031]

[0032] where k is the sample set of the k-th class in the local dataset and the local prototype set of client m is denoted as

[0033]

[0034] Further, in step B3, the calculation formula for the k-th domain prototype of domain v is: where is the cluster of the domain distribution v, and the prototype set of domain v is denoted as Furthermore, the total domain prototype set is denoted as

[0035] Further, step B4 includes:

[0036] The client inputs the local data into the downloaded global feature extractor to obtain local features. The k-th class feature of client m is denoted as:

[0037]

[0038] Then, each feature in the set is randomly fused with each domain prototype of class k to obtain the virtual feature set of class k where α is a random number sampled from the uniform distribution U(a, b), and a, b are hyperparameters; then the virtual feature set of client m is denoted as where

[0039] Further, step B5 includes:

[0040] On the virtual features uploaded by all clients, the server updates the global classifier by gradient descent according to the cross-entropy loss function: where Φ t represents the model parameters obtained in the t-th epoch, η is the learning rate, g i is a virtual feature, y i is its corresponding label, is the set of virtual features of all clients.

[0041] The second technical solution adopted by the present invention is:

[0042] An electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a federated learning method for a scenario with inconsistent domain distributions as described above.

[0043] The third technical solution adopted by the present invention is:

[0044] A computer-readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a federated learning method for a scenario with inconsistent domain distributions as described above.

[0045] The fourth technical solution adopted by the present invention is:

[0046] A computer program product or a computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions so that the computer device executes the above method.

[0047] The beneficial effects of the present invention include:

[0048] (1) For the problem of domain skew, the present invention adopts clustering federated learning, which can prevent interference of knowledge between domains. By adopting clustering federated learning, knowledge of different domains is effectively retained. By retaining the cluster model, the formed global feature extractor has both in-domain discrimination ability and inter-domain generalization ability.

[0049] (2) The present invention does not adopt the traditional model averaging as the model aggregation scheme, but proposes a model aggregation with domain knowledge protection, which can prevent the loss of domain knowledge caused by model aggregation.

[0050] (3) The present invention adopts the method of fusing local features and domain prototypes to generate virtual features for training the classifier, which can eliminate classifier bias and save communication and computing overhead of the client. Description of the Drawings

[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings of the relevant technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings in the following introduction are only for clearly expressing some embodiments of the technical solutions in the present invention for convenience. For those skilled in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0052] Figure 1 It is a flowchart of a federated learning method for a scenario with inconsistent domain distributions in an embodiment of the present invention. Detailed implementation manners

[0053] The embodiments of the present invention are described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0054] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the accompanying drawings, and it is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, so it should not be construed as a limitation to the present invention.

[0055] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is two or more, and understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or the sequence of the indicated technical features.

[0056] In the description of the present invention, unless otherwise clearly defined, words such as setting, installation, connection, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.

[0057] Term explanation:

[0058] FedAvg: Federated Averaging, is a classic federated learning algorithm.

[0059] In view of the existing technical problems, the present invention proposes a federated learning solution for the scenario of inconsistent domain distribution, including two stages: model aggregation with knowledge protection and retraining of the classifier based on domain prototypes; in the stage of model aggregation with knowledge protection, in this embodiment, several cluster models fitting specific domain distributions are first obtained by using clustering federated learning, and the server collects the feature extractors of these cluster models as the global feature extractor, so as to avoid the loss of domain knowledge caused by model averaging. In the stage of retraining the classifier based on domain prototypes, in this embodiment, the global classifier is trained on the server side using a virtual feature set fused by domain prototypes to match the higher-dimensional features output by the global feature extractor and avoid classifier bias caused by collaborative training.

[0060] Embodiment 1

[0061] As Figure 1 shown, this embodiment provides a federated learning method for the scenario of inconsistent domain distribution, which can ensure that the trained global model has both in-domain discrimination ability and inter-domain generalization ability. The method includes two stages: model aggregation with knowledge protection and retraining of the classifier based on domain prototypes:

[0062] The model aggregation with knowledge protection is specifically implemented as follows:

[0063] Step A1: The client performs truncated singular value decomposition on the local data and uploads the result to the server.

[0064] In one implementation, are the p most important left singular vectors U obtained by client m, where p is a hyperparameter; they span a data subspace as the main basis and concisely capture the main features of the basic distribution of the local data D m basic distribution.

[0065] Step A2: The server clusters the clients according to the principal angle similarity of the local data subspaces, and the clients with the same domain distribution are grouped into the same cluster.

[0066] In this embodiment, the server constructs a similarity matrix A according to where, is the minimum principal angle between two data subspaces. The more similar the data sets are, finally, the server uses agglomerative hierarchical clustering on this matrix to obtain non-overlapping clusters where V is the number of clusters and β is a parameter of the clustering algorithm HC, representing the Euclidean distance between two clusters.

[0067] Step A3: The clients in each cluster have the same domain distribution, and the domain skew problem within the cluster is solved. In this embodiment, the commonly used FedAvg can effectively fit the distribution of this domain, and these cluster models fit the specific domain distribution.

[0068] Step A4: Each cluster uploads the feature extractor part of the globally trained model to the server. The server collects these feature extractors as the feature extractor of the global model instead of using traditional model averaging, thus retaining different domain knowledge.

[0069] In this embodiment, the global feature extractor is defined as the integration of these cluster model feature extractors, that is where θ1,...,θ V are the parameters of the cluster model feature extractor, represents the model structure; the output of ε(Θ) is the concatenation of the features output by each cluster model feature extractor, that is where x represents an input sample, represents the concatenation operation and

[0070] Step A5: The server broadcasts the global feature extractor to the clients, and the clients fix the parameters of the global feature extractor, which is part of the final global model.

[0071] The retraining of the domain prototype-based classifier specifically includes the following steps:

[0072] Step B1: To process the data features with higher dimensions output by the global feature extractor, the server initializes a global classifier with a larger scale.

[0073] In some embodiments, the server initializes a global classifier which will map to a logits output where Φ are the parameters of the global classifier, represents the model structure, and I is the number of categories in the classification task.

[0074] Step B2: The clients take the average of the local features of the same class to obtain the local prototype of each class and upload it to the server.

[0075] Exemplarily, the calculation formula for the local prototype of the k-th class of client m is:

[0076]

[0077] where, k is the local dataset The sample set of the k-th class in

[0078]

[0079] Step B3: According to the clustering information, the server takes the average of the local prototypes of the clients with the same domain distribution to obtain the domain prototypes and broadcasts them to the clients.

[0080] Specifically, the calculation formula for the k-th domain prototype of domain v is where For the clusters of domain distribution v, the prototype set of domain v can be expressed as Furthermore, the total domain prototype set is expressed as

[0081] Step B4: The client randomly mixes the local features with the domain prototypes of the same category to obtain a virtual feature set and uploads it to the server.

[0082] In some embodiments, the client inputs the local data into the downloaded global feature extractor to obtain local features. The k-th class features of client m are denoted as:

[0083]

[0084] Then each feature in the set is randomly fused with each domain prototype of category k to obtain the virtual feature set of category k where α is a random number sampled from the uniform distribution U(a, b), and a, b are hyperparameters; then the virtual feature set of client m is expressed as where

[0085] Step B5: The server uses the virtual features uploaded by the clients to train the new global classifier, thereby eliminating classifier bias and not imposing additional communication and computational burdens on the clients.

[0086] Specifically, on the virtual features uploaded by all clients, the server performs gradient descent update on the global classifier according to the cross-entropy loss function: where Φ t represents the model parameters obtained in the t-th epoch, η is the learning rate, g i is a virtual feature, y i is its corresponding label, is the set of virtual features of all clients.

[0087] The method of this embodiment is further elaborated through specific experiments below.

[0088] In this embodiment, the Office-Caltech-10 and PACS datasets are selected, which have 10 and 7 categories respectively. 10 and 20 clients are set for the Office-Caltech-10 and PACS tasks respectively, and a specific number of clients are assigned to each domain. The local data is randomly sampled from the corresponding domain at different ratios to ensure that the number of local samples is relatively uniform. Specifically, Office-Caltech-10 contains four domains: Caltech, Amazon, DSLR, and Webcam; PACS contains four domains: Photo, Art Painting, Cartoon, and Sketch.

[0089] In this experiment, the number of communication rounds of federated learning is set to 100 rounds, and the number of local training times is 10. In addition, the SGD optimizer is used in this experiment, the learning rate η is set to 0.01, and the batch size is 64.

[0090] To prove the effectiveness of the present invention in heterogeneous model scenarios, this experiment compares the present invention with state-of-the-art methods, including FedAvg, FedProx, FedProc, FedProto, and FPL.

[0091] Table 1: Comparison table of test accuracies under the Office-Caltech-10 dataset

[0092]

[0093] Table 2: Comparison table of test accuracies under the PACS dataset

[0094]

[0095] As can be seen from Table 1 and Table 2, the present invention is superior to the existing strategies on two commonly used image classification datasets. It achieves the highest average accuracy (AVG) and the lowest standard deviation (STD) on the Office-Caltech-10 dataset, and the highest average accuracy and a relatively low standard deviation on the PACS dataset. Therefore, the robustness of the present invention to inconsistent domain distributions can be proved.

[0096] Example 2

[0097] The embodiment of the present invention also provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement Figure 1 a federated learning method for a scenario with inconsistent domain distributions as shown.

[0098] It can be understood that the memory may include a Random Access Memory (RAM), or may also include a Read-Only Memory. Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function, instructions for implementing the above various method embodiments, etc.; the data storage area can store data created according to the use of the server, etc.

[0099] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts within the entire server. By running or executing instructions, programs, code sets or instruction sets stored in the memory, and by calling data stored in the memory, it executes various functions of the server and processes data. Optionally, the processor can be implemented in at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor can integrate one or several combinations of a Central Processing Unit (CPU) and a modem, etc. Among them, the CPU mainly processes the operating system and application programs, etc.; the modem is used for processing wireless communication. It can be understood that the above modem may not be integrated into the processor and can be implemented separately by a single chip.

[0100] Since this electronic device is an electronic device corresponding to a federated learning method for a scenario with inconsistent domain distributions in an embodiment of the present invention, and the principle by which this electronic device solves problems is similar to that of this method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and repeated parts will not be elaborated.

[0101] Embodiment 3

[0102] An embodiment of the present invention further provides a computer-readable storage medium. At least one instruction, at least one segment of program, code set or instruction set is stored in the storage medium. The at least one instruction, the at least one segment of program, the code set or instruction set is loaded and executed by a processor to implement Figure 1 a federated learning method for a scenario with inconsistent domain distributions as shown.

[0103] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium, which includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium that can be used to carry or store data.

[0104] Since this storage medium is the storage medium corresponding to the federated learning method for the scenario of inconsistent domain distribution in the embodiments of the present invention, and the principle of solving problems by this storage medium is similar to that of this method, the implementation of this storage medium can refer to the implementation process of the above method embodiments, and the repeated parts will not be described again.

[0105] Embodiment 4

[0106] In some possible implementation manners, each aspect of the method in the embodiments of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a federated learning method for the scenario of inconsistent domain distribution according to various exemplary implementation manners described in this specification above. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0107] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0108] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0109] The above embodiments are only for illustrating the technical concept and characteristics of the present invention, and the purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered by the protection scope of the present invention.

Claims

1. A federated learning method for the scenario of domain distribution inconsistency, characterized in that, It includes two stages: model aggregation with knowledge protection and retraining of the classifier based on domain prototypes: The model aggregation with knowledge protection includes: A1: The client performs truncated singular value decomposition on the local data and uploads the decomposition result to the server; A2: The server clusters the clients according to the principal angle similarity of the local data subspaces, and the clients with the same domain distribution are grouped into the same cluster; A3: The clients in each cluster have the same domain distribution, and each cluster performs federated learning to train the global model within the cluster, and these models fit a specific domain distribution; A4: Each cluster uploads the feature extractor part of the global model it has trained to the server, and the server collects these feature extractors as the global feature extractor; A5: The server broadcasts the global feature extractor to the clients, and the clients fix the parameters of the global feature extractor; The retraining of the classifier based on domain prototypes includes: B1: To process the data features output by the global feature extractor with higher dimensions, the server initializes a larger-scale global classifier; B2: The client takes the average of the local features of the same class to obtain the local prototype of each class and uploads it to the server; B3: The server takes the average of the local prototypes of the clients with the same domain distribution according to the clustering information to obtain the domain prototype and broadcasts it to the clients; B4: The client randomly mixes the local features with the domain prototypes of the same class to obtain a virtual feature set and uploads it to the server; B5: The server uses the virtual features uploaded by the clients to train the new global classifier, thereby eliminating classifier bias and not imposing additional communication and computational burdens on the clients.

2. The federated learning method for a domain distribution inconsistent scenario according to claim 1, wherein In step A1, are the p most important left singular vectors U obtained by the client m, where p is a hyperparameter; they span a data subspace as the principal basis and capture the local data D m The main characteristics of the basic distribution.

3. A federated learning method for a domain distribution inconsistent scenario according to claim 1, characterized in that, The step A2 includes: The server, according to constructs a similarity matrix A, where the matrix element A i,j is the minimum principal angle between two data subspaces, and the expression is as follows: wherein, represents the set of p most important left singular vectors of client i, represents the set of p most important left singular vectors of client j, and u i represents the set in the vector, u j represents the set in the vector; The server uses agglomerative hierarchical clustering on matrix A to obtain disjoint clusters where v is the number of clusters and β is a parameter of the clustering algorithm HC, representing the Euclidean distance between two clusters.

4. A federated learning method for a domain distribution inconsistent scenario according to claim 1, characterized in that In step A4, the global feature extractor is defined as an ensemble of these cluster model feature extractors, i.e., where θ1,...,θ V are the parameters of the cluster model feature extractor. Represents the model structure; the output of ε(Θ) is the concatenation of the features output by each cluster model feature extractor, that is where x represents an input sample, Represents the concatenation operation.

5. A federated learning method for a domain distribution inconsistent scenario according to claim 1, characterized in that, The step B1 includes: The server initializes a global classifier It will be mapped to a logits output where Φ is the parameter of the global classifier, represents the model structure, and I is the number of categories in the classification task.

6. The federated learning method for a domain distribution inconsistent scenario according to claim 1, characterized in that, In step B2, the calculation formula for the local prototype of the k-th class of client m is: Among them, is the sample set of the k-th class in the local dataset and the local prototype set of client m is denoted as In step B3, the calculation formula for the k-th type of domain prototype in domain v is as follows: where For the cluster of domain distribution v, the prototype set of domain v is expressed as Furthermore, the total domain prototype set is expressed as 7. A federated learning method for a domain distribution inconsistent scenario according to claim 1, wherein The step B4 includes: The client inputs the local data into the downloaded global feature extractor to obtain local features, and the k-th class features of client m are denoted as: Then the set Each feature in is randomly fused with each domain prototype of class k to obtain the virtual feature set of class k where α is a random number sampled from the uniform distribution U(a, b), and a, b are hyperparameters; then the virtual feature set of client m is expressed as where 8. A federated learning method for a domain distribution inconsistent scenario according to claim 1, characterized in that The step B5 includes: On all the virtual features uploaded by the clients, the server updates the global classifier by performing gradient descent according to the cross-entropy loss function: where Φ t represents the model parameters obtained in the t-th epoch, η is the learning rate, g i is a virtual feature, and y i is its corresponding label, and is the set of virtual features of all clients.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, At least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Model heterogeneous federal learning method and system based on double dynamic prototypes

    CN121257782A

  • A model heterogeneous federated learning method and system based on double dynamic prototypes

    CN121257782B