Pseudo-label-based clustering federal learning method

By using pseudo-labels and fuzzy hierarchical clustering algorithms in cluster federated learning, the cumbersome problems of information leakage and clustering division in traditional methods are solved, and a more efficient and flexible clustering process and better model performance are achieved.

CN120218186APending Publication Date: 2025-06-27UNIV OF ELECTRONICS SCI & TECH OF CHINA +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510276693.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The traditional clustering federated learning method has the problems of information leakage risks, cumbersome clustering and division processes and inflexible results, especially when the number of clients is huge and the data distribution is very different.

Method used

A clustering federated learning method based on pseudo-labels is adopted, and the initial model is sent to the client through the central server. The client generates pseudo-labels and uncertainty scores, calculates the similarity between clients, and uses a fuzzy hierarchical clustering algorithm for clustering to ensure that the clients within the cluster have highly consistent pseudo-label features, and introduces consistency optimization in the model training stage.

Benefits of technology

It effectively prevents data leakage, improves clustering efficiency, reduces computational overhead, optimizes the clustering process, and improves the performance and generalization capabilities of the model, especially when the client data distribution varies greatly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218186A_ABST
    Figure CN120218186A_ABST
Patent Text Reader

Abstract

The invention provides a clustering federal learning method based on pseudo labels, which comprises the following steps: firstly, training each client by using own local data set to obtain respective convergent local models, then predicting a public data set to generate the pseudo labels, and finally, according to a pseudo label data set, carrying out clustering on each client to obtain a clustering federal learning model; a similarity matrix between the clients is calculated, the similarity matrix is based on pseudo labels and uncertainty scores, after the similarity matrix is obtained, the clients are clustered by adopting a fuzzy hierarchical clustering algorithm, the clients are distributed to a plurality of clusters after clustering is completed, the clients in each cluster have similar data distribution and pseudo labels, and the clients in each cluster are distributed to a plurality of clusters. And finally, on the basis of the clusters, carrying out personalized model training in the clusters. According to the scheme, data leakage is effectively prevented, the clustering efficiency is improved, the calculation overhead is reduced, and the performance and generalization ability of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of federated learning, and particularly relates to a clustering federated learning method based on pseudo-labels. Background Art

[0002] With the increasingly strict protection of data by citizens and the promulgation of relevant personal privacy laws, the problem of data islands has become increasingly serious. Fortunately, the birth of Federated Learning (FL) provides a solution to this problem. Federated learning trains a high-quality shared global model from data scattered across a large number of different clients through a central server, thus breaking the situation of data islands. Initially, federated learning was proposed by Google and was introduced as a technology widely used in the Gboard application at the Google I / O Developer Conference in 2019. Different from traditional recommendation systems, the training of the Gboard model largely depends on the mobile device itself, which means that the recommendation process can be completed without infringing on user privacy. In recent years, federated learning has been widely applied in many fields such as finance, healthcare, and social networks.

[0003] However, although federated learning can effectively protect data privacy, it still faces some challenges, especially in complex application environments. Widely distributed devices generate a large amount of non-independent and identically distributed (non-IID) data, which may lead to poor performance of the global model on some clients, and even its performance is lower than that of the local model. Therefore, the traditional unified global model is not suitable for all clients, especially when the data distribution differences among clients are large. This makes Clustered Federated Learning (CFL) a more suitable method. Clustered federated learning divides clients with similar data distributions into the same group, making the data distribution within the group closer to independent and identically distributed (IID), and training personalized models for each group, which can effectively alleviate the impact of non-IID data on the performance of the global model, thereby improving the training effect and personalized performance of the model.

[0004] Clustering federated learning has shown great potential in dealing with non-IID data and personalized requirements, but traditional clustering federated learning methods still have the following limitations. In traditional clustering federated learning methods, clients usually cluster by calculating the similarity of model parameters or directly using the similarity of local datasets. This method has a risk of information leakage because by comparing client model weights or local data, sensitive information of clients may be indirectly revealed. For example, the model parameters may contain important features about the local data distribution. Attackers may infer data characteristics or preferences of certain clients by analyzing the similarity between models. Although some methods try to reduce the risk of data leakage through encryption techniques, these techniques still increase additional computational overhead, and encrypted computations cannot completely avoid the leakage of sensitive information through model similarity.

[0005] At the same time, in traditional clustering federated learning methods, the number of clients is large, and the difference in data distribution makes the clustering process cumbersome. The partitioning results are inflexible. There may be clients that are not in the optimal clusters. How to select an appropriate number of clusters directly affects the clustering effect and the performance of the personalized model. Too many clusters will lead to overfitting and cannot effectively improve the generalization ability of the model; while too few clusters will fail to capture the subtle differences between clients, resulting in a decrease in the accuracy of the personalized model. Summary of the Invention

[0006] To solve the above problems, the present invention proposes a clustering federated learning method based on pseudo-labels. This method improves the clustering efficiency, reduces the computational overhead, and optimizes the clustering process while ensuring privacy. The method includes:

[0007] Step S1, the central server sends the initial model to each client. Each client uses its own local dataset for training to obtain its own converged local model. Each client uses the converged local model to predict a public common dataset to generate pseudo-labels. Finally, each client generates a pseudo-label dataset. The pseudo-label of each sample in the public dataset is determined by the maximum probability category output by the converged local model of each client. At the same time, the uncertainty score is calculated to measure the confidence of the model in predicting this sample.

[0008] Step S2, the server receives the pseudo-label datasets sent by all clients. By using the pseudo-labels and uncertainty scores of each client, it calculates the similarity between clients to obtain a similarity matrix, and designs a fuzzy hierarchical clustering algorithm to cluster clients with high similarity into the same cluster. For clients with high similarity to multiple clusters, they are divided into multiple clusters to ensure that the clients within the cluster have highly consistent pseudo-label features to support efficient federated learning training.

[0009] In step S3, according to the clustering results, each client participates in the corresponding cluster for federated learning training. The clients within each cluster cooperate in training to jointly optimize the cluster model to better adapt to the common data features within the cluster. Meanwhile, consistency optimization is introduced to further optimize the consistency of the pseudo-labels during the training phase of the model, ensuring that the models of each cluster can be more stably generalized to similar feature distributions, thereby improving the overall model performance and adaptability.

[0010] Furthermore, step S1 includes:

[0011] First, the central server initializes the global model parameters θ and sends them to all clients. Each client c i uses its local dataset to train the model. Here, x j represents sample j, y j represents the label of sample j, N i represents the number of samples in the local dataset of client c i . After the client completes the local model training, each client c i obtains the converged local model and uses the trained model to provide prediction labels, i.e., pseudo-labels, for the unlabeled data. The expression of the pseudo-label is as follows:

[0012]

[0013] where k represents the category, and f i (x j ; θ i ) k is the prediction probability of the converged model of client c i for category k. When generating the pseudo-labels, the uncertainty score of the label is introduced to mark the uncertainty of the model for this output result. The uncertainty score is measured by calculating the entropy of the probability distribution. The larger the entropy value, the more uncertain the model's prediction of the sample. The calculation method of the uncertainty score is as follows:

[0014]

[0015] where K is the total number of categories. Finally, each client generates a pseudo-label dataset, which includes samples, pseudo-labels, and uncertainty scores. The pseudo-label dataset generated by client c i is expressed as follows:

[0016]

[0017] where, represents client c iThe generated pseudo-label dataset, represents the publicly available common dataset.

[0018] Furthermore, the step S2 includes:

[0019] Define the similarity between two pseudo-label datasets to measure the similarity degree of clients c i and client c k :

[0020]

[0021] where α is the weight parameter, is the sample-level similarity, is the label distribution similarity;

[0022] The sample-level similarity is calculated as follows:

[0023]

[0024] where, represents the number of samples in the common dataset, and the sample-level similarity s j (i, k) represents the contribution of sample to the similarity between two clients c i and c k . The calculation method of the sample-level similarity contribution s j (i, k) is as follows:

[0025]

[0026] where cert j (i, k) represents the average certainty of client c i and client c k for sample , and the calculation method is:

[0027]

[0028] The average certainty factor cert j (i, k) ensures that when both clients are very certain about the prediction, the consistency or inconsistency of their predictions has a greater weight. When the pseudo-labels of the two clients are consistent, if their uncertainty scores are closer (i.e., The smaller the value, the greater the similarity contribution. When the pseudo-labels of two clients are inconsistent, if their certainty is higher, the similarity contribution is a larger negative value, indicating that their differences are more significant.

[0029] Define the label distribution similarity as:

[0030]

[0031] where D JS (P i , P k ) is the Jensen-Shannon divergence between clients c i and c k , measuring the difference in their label distributions; represents the pseudo-label distribution of client c i over all samples, represents the proportion of class k on client c i , represents the pseudo-label distribution of client c k over all samples, represents the proportion of class k on client c k , and K is the total number of classes.

[0032] The calculation of the Jensen-Shannon divergence is as follows:

[0033]

[0034] where is the average of the two distributions; D KL (P i ∥M) and D KL (P k ∥M) are the Kullback-Leibler divergences:

[0035]

[0036] The label distribution similarity is finally expressed as:

[0037]

[0038] This formula ensures that when the label distributions of two clients are exactly the same, D JS (P i , P k ) = 0, and the label distribution similarity When the label distributions of two clients are completely different, the label distribution similarity decreases.

[0039] Then, a similarity matrix is constructed based on the similarity. After obtaining the similarity matrix S, fuzzy hierarchical clustering is performed. The specific process of the fuzzy hierarchical clustering algorithm is as follows:

[0040] Initialize each client as an independent cluster to form an initial cluster set Let N be the total number of clients, that is, at the beginning, each client corresponds to a cluster; initialize the distance matrix D. Since in the initial cluster set, each client is an independent cluster, the initial distance matrix Initialize the virtual connection record table V i is an empty set, V i represents the virtual connection record of each client, denoted as

[0041] Through an iterative method, gradually merge similar clusters and update the cluster set and virtual connection records. In each merge iteration t, select the two clusters C a and C b with the smallest distance in the distance matrix D, and merge these two clusters C a and C b into a new cluster C ab =C a ∪C b , and update the cluster set Record the merge distance, denoted as d t =D ab , where D ab represents the distance between clusters C a and C b . Update the cluster set by removing clusters C from the cluster set a and C b , and adding the new cluster C ab , which is expressed as:

[0042]

[0043] While performing cluster merging, it is necessary to record the virtual connections of the clients corresponding to the two merged clusters, that is, cluster according to the similarity matrix S. Which clients should each client be connected to, that is, for each client c a and C b in, for each client c i ∈{C a ,C b}, select the client c i in descending order of similarity according to the similarity between c j and other clients. The number of selected clients is the same as that of client c iThe number of clients in the clusters to be merged with the current cluster, add client c j to the virtual connection record table of client c i and update V i = V i ∪{c j}, during the entire clustering process, client c i will generate multiple records, but the selected client needs to continue from the previous record instead of starting over.

[0044] After each merge, the distance matrix D needs to be updated. Remove the rows and columns corresponding to the merged clusters C a and C b from the distance matrix, and add the distances between the new cluster C ab and other clusters. The calculation method for the distance between the new cluster C ab and other clusters C c is as follows:

[0045]

[0046] where, |C ab | represents the number of clients in cluster C ab , |C c | represents the number of clients in cluster C c , and D ij represents the distance between client c i and client c j .

[0047] Continue the iteration until all clients are merged. Finally, a large cluster is obtained. During this process, a dendrogram is generated to record the process of each cluster merge and the corresponding merge distance, and then a cut-off height is determined to cut the final large cluster. The selection of the cut-off height is based on the maximum gap principle, that is, find the maximum gap between the merge distances in the dendrogram. The difference in merge distances is expressed as Δ t =(d t+1 -d t ), t ∈ {1, 2,..., N - 2}. Set the cut-off height h as:

[0048]

[0049] That is, find the iteration number t * with the largest difference in merge distances, and cut the dendrogram at this merge. After cutting the generated dendrogram at the cut-off height, L = t * + 1 clusters are obtained. At this time, all clients are divided into L clusters, and the virtual connection records generated after t * are also cancelled.

[0050] Finally, an improved two-stage clustering adjustment mechanism is introduced. By analyzing the virtual connection situation of each client, boundary point identification and clustering adjustment are carried out. For each client, according to its similarity with clients in other clusters, it is judged whether it should be classified into multiple clusters or reallocated to a cluster with higher similarity. Specifically: by judging the virtual connection situation of each client, boundary points are judged. If a client is in the edge part of a cluster, it is very likely to be similar to multiple clusters, and its virtual connection formed according to the similarity matrix will inevitably connect it with clients in other clusters. Therefore, through the objects of the virtual connection, it can be divided into multiple clusters with which it has a virtual connection to achieve the effect of fuzzy clustering and obtain the final clustering division.

[0051] Further, the step S3 includes:

[0052] In the initial stage, for each cluster C m , first generate a shared pseudo-label dataset within the cluster For each sample j ∈ D public , the pseudo-labels of each client are weighted-voted according to their uncertainty scores to obtain the final shared pseudo-label The weights are determined by the uncertainty scores. The lower the uncertainty of a label, the greater its weight in the aggregation result. The calculation method of the shared pseudo-label is as follows:

[0053]

[0054] where k represents the category, C m represents the set of clients in the m-th cluster, I(·) is the indicator function, which takes the value of 1 when the condition in the parentheses holds, otherwise 0, that is represents whether the prediction of client c i for sample j is the category k.

[0055] During the training process of the clustering model, in order to improve the consistency of the output results of each client within the same cluster for the shared pseudo-label dataset D shared , a consistency loss term is added to the total loss function. Specifically, for a cluster C m , the consistency loss L consistency (θ i ) is defined as the expected prediction error of client c i on the shared pseudo-label dataset D shared :

[0056]

[0057] where Denote the client as c i The predicted output of the model for the input x j Let l(·,·) denote the loss function [·] represents the calculation of the expected value on the shared pseudo-label dataset D shared That is, the average of the losses of all samples in D share The calculation formula of the loss function l(f θ (x),y) is as follows:

[0058]

[0059] By minimizing the above consistency loss, the client model will be forced to approach the prediction results of the overall clustering, thereby reducing the output differences of the models within the clustering. Further, when the consistency loss and the classification loss are combined into the total loss function, the optimization objective is expressed as:

[0060]

[0061] where λ is a hyperparameter used to balance the weights of the classification loss and the consistency loss. At the same time, the update formula of the client model parameters is expressed as:

[0062]

[0063] where η is the learning rate, z is the output of the neural network is the gradient of the total loss function is the model parameter after the t-th round of training for client c i Due to the introduction of the consistency loss, the update of the model parameter θ i not only depends on the classification performance of the local data but also is affected by the prediction results of the shared pseudo-label data determined by the clustering. This will force the client model parameters to approach the clustering aggregation parameters, thereby reducing the differences in the model parameters within the clustering. Compared with the traditional method without introducing the consistency loss, after introducing the consistency loss, the variance of the parameter distribution within the clustering will be significantly reduced.

[0064] During the training process of the clustering model, each clustering uses the Federated Averaging (FedAvg) algorithm for parameter aggregation, but the aggregation is only carried out within the clustering, not globally. For the set of clients in clustering C m The server collects the model parameters uploaded by all clients after each round of communication and updates the clustering model in a weighted average manner:

[0065]

[0066] where, |D i | represents client c iThe size of the local dataset of, |D j represents the local dataset size of client c j ; is the model parameter of client c i after the (t + 1)-th round of training, and is the aggregated model parameter of cluster C m in the (t + 1)-th round.

[0067] For the case where a client is identified as a boundary client, that is, a client belongs to multiple clusters simultaneously, these clients can participate in the training processes of multiple clusters. Specifically, if client c i simultaneously belongs to cluster C m and cluster C n , then this client will alternately use the models of the corresponding clusters for training in different training rounds and upload the training results to the corresponding cluster servers for aggregation respectively. This mechanism enables boundary clients to make full use of the knowledge of multiple relevant clusters, thereby improving the model performance and generalization ability.

[0068] The present invention reflects the similarity degree between clients by using pseudo-labels, avoids using local datasets or model parameters, effectively prevents data leakage, and uses a fuzzy hierarchical clustering algorithm to divide clients with high similarity into one cluster and train their respective cluster models for each cluster, effectively improving the model performance and generalization ability. The beneficial technical effects of the present invention are as follows:

[0069] Calculate the similarity between clients through pseudo-labels and uncertainty scores, and accurately cluster clients with similar data distributions. Different from traditional methods that rely on directly sharing models or local data to calculate similarities, the present invention avoids the risk of privacy leakage by generating pseudo-labels, thereby effectively reducing the negative impact brought by non-IID data. This method improves the clustering accuracy and enhances the training effect of the global model. Especially in the case where the client data distributions vary greatly, it can significantly improve the performance of the global model;

[0070] Replace the similarity calculation based on model parameters or local data in traditional methods with the similarity calculation based on pseudo-labels, greatly reducing the amount of data to be transmitted. Clients only need to transmit pseudo-labels and uncertainty scores, without uploading a large amount of model parameters or local data, greatly reducing the computational and communication overhead;

[0071] By introducing the fuzzy clustering method, the present invention solves the problem that traditional hard clustering methods cannot fully utilize the diversity of client data. Different from the rigid division of clients into a single cluster, fuzzy clustering allows each client to participate in training in multiple clusters, ensuring that the data characteristics of each client can be fully utilized in multiple clusters. This flexible clustering method improves the adaptability of the personalized model, enabling the model to better reflect the unique needs of each client and enhancing the accuracy and generalization ability of the personalized model. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0073] Figure 1 It is the overall framework diagram of a clustering federated learning method based on pseudo-labels in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0075] The present invention proposes a clustering federated learning method based on pseudo-labels, and the overall framework is as Figure 1 shown, and the method includes the following steps:

[0076] Step 1: The central server sends the initial model to each client, and each client uses its own local dataset for training to obtain its own converged local model. Each client uses the converged local model to predict a publicly available common dataset to generate pseudo-labels. Finally, each client generates a pseudo-label dataset. The pseudo-label of each sample in the public dataset is determined by the maximum probability category output by the converged local model of each client, and at the same time, the uncertainty score is calculated to measure the confidence of the model's prediction of this sample.

[0077] First, the central server initializes the global model parameters θ and sends them to all clients. Each client c i uses its local dataset to train the model. Among them, x j represents sample j, and yj Denotes the label of sample j, N i Denotes client c i The number of samples in the local dataset of client c. The training objective is to minimize the loss function between the predicted label and the true label. When the change of the loss function is less than the preset threshold ∈ in several consecutive iterations or after a fixed maximum number of training rounds T, it can be considered that the model has converged. This convergence criterion helps to control the quality of the training process, thus ensuring the accuracy of the generated pseudo-labels.

[0078] After the local model training is completed on the client side, each client c i Obtains the converged local model Uses the trained model to provide predicted labels, i.e., pseudo-labels, for the unlabeled data. The expression of the pseudo-labels is as follows:

[0079]

[0080] Where k represents the category, and f i (x j ; θ i ) k Is the predicted probability of category k by the converged model of client c i When generating the pseudo-labels, the uncertainty score of the label is introduced to mark the uncertainty of the model for the output result. The uncertainty score is measured by calculating the entropy of the probability distribution. The larger the entropy value, the more uncertain the model's prediction of the sample. The calculation method of the uncertainty score is as follows:

[0081]

[0082] Where K is the total number of categories. Finally, each client generates a pseudo-label dataset, which includes samples, pseudo-labels, and uncertainty scores. The pseudo-label dataset generated by client c i Is represented as follows:

[0083]

[0084] Where Denotes the pseudo-label dataset generated by client c i And Denotes the publicly available common dataset.

[0085] Step 2: The server receives the pseudo-label datasets sent by all clients. By using the pseudo-labels and uncertainty scores of each client, it calculates the similarity between clients to obtain a similarity matrix, and designs a fuzzy hierarchical clustering algorithm to cluster clients with high similarity into the same cluster. For clients with high similarity to multiple clusters, they are assigned to multiple clusters to ensure that the clients within the cluster have highly consistent pseudo-label features to support efficient federated learning training. The core idea of fuzzy clustering is that each client can belong to multiple clusters, allowing clients within the cluster to share similar data features while enhancing the flexibility and accuracy of the clustering results.

[0086] To calculate the similarity between clients, two pseudo-label datasets are defined and the similarity between to measure the similarity degree between client c i and client c k :

[0087]

[0088] where α is the weight parameter, is the sample-level similarity, is the label distribution similarity;

[0089] The sample-level similarity is calculated as follows:

[0090]

[0091] where, represents the number of samples in the common dataset. By accumulating and averaging the contribution values of each sample, the sample-level similarity s j (i,k) represents the similarity contribution of sample to two clients c i and c k . For the sample-level similarity contribution s j (i,k), the following factors are considered: 1) Pseudo-label consistency: Whether the pseudo-label predictions of client c i and client c k for sample are consistent; 2) Prediction certainty: The certainty level calculated based on the uncertainty score; 3) Uncertainty difference: The degree of difference in the uncertainty predicted by the two clients, which is calculated as follows:

[0092]

[0093] where cert j(i, k) represents client c i and client c k 's average certainty for the samples, calculated as:

[0094]

[0095] Average certainty factor cert j (i, k) ensures that when both clients are very certain about the prediction, the agreement or disagreement of their predictions has a greater weight. When the pseudo-labels of the two clients are the same, if their uncertainty scores are closer (i.e., smaller), the similarity contribution is greater; when the pseudo-labels of the two clients are different, if their certainty is higher, the similarity contribution is a larger negative value, indicating that their differences are more significant.

[0096] To measure the similarity of the label distributions of client c i and c k , we define:

[0097]

[0098] where D JS (P i , P k ) is the Jensen-Shannon divergence between client c i and c k , measuring the difference in their label distributions; represents the pseudo-label distribution of client c i over all samples, represents the proportion of class k on client c i , represents the pseudo-label distribution of client c k over all samples, represents the proportion of class k on client c k , and K is the total number of classes.

[0099] The calculation of the Jensen-Shannon divergence is as follows:

[0100]

[0101] where, is the average of the two distributions; D KL (P i ∥M) and D KL (P k ∥M) are the Kullback-Leibler divergences:

[0102]

[0103] The final expression of the label distribution similarity is as follows:

[0104]

[0105] This formula ensures that when the label distributions of two clients are exactly the same, D JS (P i , P k ) = 0, and the label distribution similarity When the label distributions of two clients are completely different, the label distribution similarity decreases.

[0106] After obtaining the similarity matrix S, fuzzy hierarchical clustering needs to be performed. The specific process of the fuzzy hierarchical clustering algorithm is as follows:

[0107] Initialize each client as an independent cluster to form an initial cluster set Let N be the total number of clients, that is, at the beginning, each client corresponds to a cluster; initialize the distance matrix D. Since in the initial cluster set, each client is an independent cluster, the initial distance matrix Initialize the virtual connection record table V i is an empty set. V i represents the virtual connection record of each client, denoted as

[0108] Through an iterative method, gradually merge similar clusters and update the cluster set and virtual connection records. In each merge iteration t, select the two clusters C a and C b with the smallest distance in the distance matrix D, and merge these two clusters C a and C b into a new cluster C ab = C a ∪ C b , and update the cluster set Record the merge distance, denoted as d t = D ab , where D ab represents the distance between clusters C a and C b . Update the cluster set by removing clusters C and C a from the cluster set b , and adding the new cluster C ab , which is expressed as:

[0109]

[0110] While performing cluster merging, it is necessary to record the virtual connections of the clients corresponding to the two merged clusters, that is, cluster according to the similarity matrix S, and determine which clients each client should be connected to, that is, for cluster C a and C b For each client c i ∈{C a ,C b}, according to the similarity between c i and other clients, select clients c j from high to low according to the similarity. The number of selected clients is the number of clients in the cluster that is merged with the cluster where the client c i is currently located. Add the client c j to the virtual connection record table of the client c i , and update V i = V i ∪ {c j}. During the entire clustering process, the client c i will generate multiple records, but the selected clients need to be selected continuously from the previous record instead of starting over.

[0111] After each merge, it is necessary to update the distance matrix D, remove the rows and columns corresponding to the merged clusters C a and C b in the distance matrix, and add the distances between the new cluster C ab and other clusters. The calculation method for the distance between the new cluster C ab and other clusters C c is as follows:

[0112]

[0113] Where, |C ab | represents the number of clients in cluster C ab , |C c | represents the number of clients in cluster C c , and D ij represents the distance between the client c i and the client c j .

[0114] Continue to iterate until all clients are merged, and finally obtain a large cluster. During this process, a dendrogram will be generated to record the process of each cluster merge and the corresponding merge distance, and then a cut-off height is determined to trim the final large cluster. The selection of the cut-off height is based on the maximum gap principle, that is, find the maximum gap between the merge distances in the dendrogram. The difference in merge distances is expressed as Δ t = (d t+1 - d t), where \(t\in\{1,2,\ldots,N - 2\}\), set the truncation height \(h\) as follows:

[0115]

[0116] That is, find the iteration number \(t\) with the largest difference in the merging distance * , and truncate the dendrogram at this merging point. After truncating the generated dendrogram at the truncation height, we get \(L=t\) * + 1 clusters. At this time, all clients are divided into \(L\) clusters, and the virtual connection records generated after \(t\) * are also cancelled together.

[0117] Finally, an improved two-stage clustering adjustment mechanism is introduced. By analyzing the virtual connection situation of each client, boundary point recognition and clustering adjustment are carried out. For each client, according to its similarity with clients in other clusters, it is judged whether it should be assigned to multiple clusters or reallocated to a cluster with higher similarity. Specifically: by judging the virtual connection situation of each client, boundary points are judged. If a client is in the edge part of a cluster, it is very likely to be similar to multiple clusters, and its virtual connection formed according to the similarity matrix will inevitably connect it to clients in other clusters. Therefore, through the objects of the virtual connection, it can be divided into multiple clusters with which it has a virtual connection to achieve the effect of fuzzy clustering and obtain the final clustering division.

[0118] Step 3: According to the clustering results, each client participates in the corresponding cluster for federated learning training. The clients within each cluster cooperate in training to jointly optimize the clustering model to better adapt to the common data characteristics within the cluster. At the same time, consistency optimization is introduced. During the training stage of the model, by further optimizing the pseudo-label consistency, it is ensured that the models of each cluster can be more stably generalized to similar feature distributions, thereby improving the overall model performance and adaptability.

[0119] In the initial stage, for each cluster \(C\) m , first generate a shared pseudo-label dataset within the cluster For each sample \(j\in D\) public , the pseudo-labels of each client are weighted-voted according to their uncertainty scores to obtain the final shared pseudo-label The weights are determined by the uncertainty scores. The lower the uncertainty of a label, the greater its weight in the aggregation result. The calculation method of the shared pseudo-label is as follows:

[0120]

[0121] where \(k\) represents the category, \(C\) mDenote the set of clients in the \(m\)-th cluster. \(I(\cdot)\) is the indicator function, which takes the value of 1 when the condition in the parentheses holds, and 0 otherwise, that is denote client \(c\) i whether the prediction of sample \(j\) is class \(k\).

[0122] During the training process of the clustering model, in order to improve the consistency of the output results of each client within the same cluster for the shared pseudo-label dataset \(D\) shared a consistency loss term is added to the total loss function. This consistency loss is used to constrain the prediction results of each client model to be as close as possible to the shared pseudo-label prediction formed within the cluster, thereby reducing the prediction differences between different clients. Specifically, for a cluster \(C\) m , define the consistency loss \(L\) consistency \((\theta\) i ) as the expected prediction error of client \(c\) i on the shared pseudo-label dataset \(D\) shared :

[0123]

[0124] where denote the prediction output of the model of client \(c\) i for the input \(x\) j , \(l(\cdot, v)\) represents the loss function, denote calculating the expected value on the shared pseudo-label dataset \(D\) shared , that is, taking the average of the losses of all samples in \(D\) share . The calculation formula of the loss function \(l(f\) θ (x), y)\) is as follows:

[0125]

[0126] By minimizing the above consistency loss, the client model will be forced to approach the prediction result of the whole cluster, thereby reducing the output differences of the models within the cluster. Further, when the consistency loss and the classification loss are combined into the total loss function, the optimization objective is expressed as:

[0127]

[0128] where \(\lambda\) is a hyperparameter used to balance the weights of the classification loss and the consistency loss. At the same time, the update formula of the client model parameters is expressed as:

[0129]

[0130] where \(\eta\) is the learning rate, \(z\) is the output of the neural network, is the gradient of the total loss function, and \(c\) is the clienti The model parameters after the t-th round of training. Due to the introduction of the consistency loss, the model parameters θ i are updated not only depending on the classification performance of the local data but also affected by the prediction results of the shared pseudo-label data determined by the clustering. This will force the client model parameters to approach the clustering aggregation parameters, thereby reducing the difference in the model parameters within the clustering. Compared with the traditional method without introducing the consistency loss, after introducing the consistency loss, the variance of the parameter distribution within the clustering will be significantly reduced.

[0131] During the training process of the clustering model, each clustering uses the Federated Averaging (FedAvg) algorithm for parameter aggregation, but the aggregation is only carried out within the clustering, not globally. For the set of clients in clustering C m the server collects the model parameters uploaded by all clients after each round of communication and updates the clustering model in a weighted average manner:

[0132]

[0133] where |D i | represents the size of the local dataset of client c i | represents the size of the local dataset of client c j | represents the size of the local dataset of client c j is the model parameter of client c after the (t + 1)-th round of training, i is the aggregated model parameter of clustering C at the (t + 1)-th round. m

[0134] For the case of clients identified as boundary clients, that is, a client belongs to multiple clusters at the same time. These clients can participate in the training processes of multiple clusters. Specifically, if client c i belongs to both clustering C m and clustering C n , then this client will alternately use the models of the corresponding clusters for training in different training rounds and upload the training results to the corresponding cluster servers for aggregation respectively. This mechanism enables boundary clients to make full use of the knowledge of multiple relevant clusters, thereby improving the model performance and generalization ability.

[0135] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.​

Claims

1. A clustering federated learning method based on pseudo labels, characterized in that: The method comprises the following steps: Step S1: The central server sends the initial model to each client. Each client uses its own local data set for training to obtain its own converged local model. Each client uses the converged local model to predict a public public data set and generate pseudo labels. Finally, each client generates a pseudo-labeled data set. The pseudo label of each sample in the public data set is determined by the maximum probability category output by the local model of each client. At the same time, the uncertainty score is calculated to measure the confidence of the model in the prediction of the sample. Step S2: The server receives the pseudo-label data sets sent by all clients, calculates the similarity between the clients by using the pseudo-labels and uncertainty scores of each client, obtains a similarity matrix, and designs a fuzzy hierarchical clustering algorithm to cluster clients with high similarity into the same cluster. For clients with high similarity to multiple clusters, they are divided into multiple clusters to ensure that the clients within the cluster have highly consistent pseudo-label features to support efficient federated learning training. In step S3, according to the clustering results, each client participates in the federated learning training in the corresponding cluster. The clients within each cluster collaborate in training and jointly optimize the clustering model to better adapt to the common data features within the cluster. At the same time, consistency optimization is introduced. During the training phase of the model, the consistency of pseudo labels is further optimized to ensure that the model of each cluster can be more stably generalized to similar feature distributions, thereby improving the overall model performance and adaptability.

2. The method according to claim 1, characterized in that The step S1 further comprises: First, the central server initializes the global model parameters θ and sends them to all clients. Each client c i Leverage its local dataset Train the model, where x j represents sample j, y j represents the label of sample j, N i Represents client c i After the client completes the local model training, each client c i Get the converged local model The trained model is used to provide predicted labels, namely pseudo labels, for unlabeled data. The pseudo labels are expressed as follows: Among them, k represents the category, f i (x j θ i ) k Is the client c i The predicted probability of the converged model for category k; When generating pseudo labels, the uncertainty score of the label is introduced to mark the uncertainty of the model for the output result. The uncertainty score is measured by calculating the entropy of the probability distribution. The larger the entropy value, the more uncertain the model's prediction of the sample. The calculation method of the uncertainty score is: Where K is the total number of categories; Finally, each client generates a pseudo-label dataset, which includes samples, pseudo-labels, and uncertainty scores. i The generated pseudo-labeled dataset is represented as: in, Represents client c i The pseudo-labeled dataset generated, Represents a publicly available public dataset.

3. The method according to claim 1, characterized in that The step S2 further comprises: Step S21, define two pseudo-label data sets and The similarity between To measure the client c i and client c k The similarity: Among them, α is the weight parameter, is the sample-level similarity, is the label distribution similarity; Step S22, construct a similarity matrix S according to the similarity, and then perform fuzzy hierarchical clustering according to the similarity matrix. The specific process of the fuzzy hierarchical clustering algorithm is as follows: Initialize each client as an independent cluster to form an initial cluster set N is the total number of clients, that is, each client corresponds to one cluster at the beginning; Initialize the distance matrix D, the initial distance matrix Initialize virtual connection record table V i is an empty set, V i Represents the virtual connection record of each client, denoted as Through iteration, similar clusters are gradually merged, and cluster sets and virtual connection records are updated until all clients are merged, and finally a large cluster is obtained. During the iteration, a dendrogram is generated to record the process of each cluster merging and the corresponding merging distance. Finally, a truncation height is determined to trim the final large cluster. Step S23, introduces an improved two-stage clustering adjustment mechanism, and determines the boundary points by judging the virtual connection status of each client. If a client is at the edge of the cluster, that is, it has similarities with multiple clusters, the virtual connection formed according to the similarity matrix will inevitably connect it with the clients in other clusters. Therefore, through the virtually connected objects, it is divided into multiple clusters with which it has virtual connections, so as to achieve the effect of fuzzy clustering and obtain the final cluster division.

4. The method according to claim 3, characterized in that The sample-level similarity in step S21 The calculation method is: in, Represents the number of samples in the public data set. The sample-level similarity is obtained by accumulating and averaging the contribution value of each sample. s j (i,k) represents the sample For two clients c i and c k Similarity contribution, sample-level similarity contribution s j The calculation of (i,k) is as follows: Among them, cert j (i,k) represents client c i and client c k For samples The average certainty is calculated as:

5. The method according to claim 3, characterized in that: The label distribution similarity in step S21 is calculated as follows: Among them, D JS (P i ,P k ) is the client c i and c k The Jensen-Shannon divergence of , which measures the difference in its label distribution, Represents client c i The pseudo-label distribution over all samples, Represents category k on client c i The proportion of Represents client c k The pseudo-label distribution over all samples, Represents category k on client c k The proportion of above, K is the total number of categories; The Jensen-Shannon divergence is calculated as: in, is the average of the two distributions; D KL (P i ∥M) and D KL (P k ∥M) is the Kullback-Leibler divergence, calculated as: The label distribution similarity is finally expressed as: When the label distributions of the two clients are exactly the same, D JS (P i ,P k )=0, label distribution similarity When the label distributions of two clients are completely different, label distribution similarity decline.

6. The method according to claim 3, characterized in that The process of each iteration is: In each merging iteration t, select the two clusters C with the smallest distance in the distance matrix D a and C b , these two clusters C a and C b Merge into a new cluster C ab =C a ∪C b , and update the clustering set Record the merge distance, denoted as d t =D ab , where D ab Represents cluster C a and C b The distance between them is used to update the clustering set From the cluster set Remove cluster C a and C b , and add a new cluster C ab , expressed as: While merging clusters, virtual connection records are made for the clients corresponding to the two merged clusters. a and C b Each client c i ∈{C a ,C b }, according to c i Similarity with other clients, select client c from high to low based on similarity j , the number of choices is the same as the number of clients c i The number of clients in the cluster to be merged with the current cluster, and the client c j Add to client c i In the virtual connection record table, update V i =V i ∪{c j }; After each merge, the distance matrix D needs to be updated to remove the merged clusters C in the distance matrix. a and C b Corresponding rows and columns, and add a new cluster C ab The distance to other clusters, the new cluster C ab and other clusters C c The distance between is calculated as: Among them, |C ab | represents cluster C ab The number of clients in |C c | represents cluster C c The number of clients in D ij Represents client c i and client c j The distance between.

7. The method according to claim 3, characterized in that The selection of the cutoff height is based on the maximum gap principle, that is, finding the maximum gap between the merged distances in the dendrogram, and the difference in merged distances is expressed as Δ t =(d t+1 -d t ),t∈{1,2,…,N-2}, set the cutoff height h to: That is, find the number of iterations t with the largest merge distance difference * , and cut off the dendrogram at the merge, and get L = t * +1 cluster, all clients are divided into L clusters, and at t * The virtual connection records generated afterwards are also cancelled.

8. The method according to claim 1, characterized in that The step S3 further comprises: In the initial stage, for each cluster C m , first generate a shared pseudo-label dataset within the cluster For each sample j∈D public , the pseudo-label of each client Score based on its uncertainty Perform weighted voting to obtain the final shared pseudo-label The weight is determined by the uncertainty score. The lower the uncertainty, the greater the weight of the label in the aggregation result. The calculation method of shared pseudo-label is as follows: Among them, k represents the category, C m represents the client set in the mth cluster, I(·) is an indicator function, which takes the value 1 when the condition in the brackets is met, otherwise it takes the value 0, that is, Represents client c i Whether the prediction for sample j is category k; In the clustering model training process, the consistency loss term is added to the total loss function. m , define the consistency loss L consistency (θ i ) is the client c i On the shared pseudo-label dataset D shared Expected prediction error on : in, Represents client c i The model is for input x j The predicted output of , l(·,·) represents the loss function, Indicates that in the shared pseudo-label dataset D shared Calculate the expected value, loss function l(f θ The calculation formula of (x), y) is as follows: When the consistency loss and classification loss are combined into the total loss function, the optimization objective is expressed as: Among them, λ is a hyperparameter used to balance the weights of classification loss and consistency loss. At the same time, the update formula of the client model parameters is expressed as: Among them, η is the learning rate, z is the output of the neural network, is the gradient of the total loss function, Is the client c i Model parameters after tth round of training.

9. The method according to claim 8, characterized in that During the clustering model training process, each cluster uses the federated average algorithm to aggregate parameters, but the aggregation is only performed within the cluster. m The server collects the model parameters uploaded by all clients after each round of communication and updates the clustering model by weighted average: Among them, |D i | indicates client c i The size of the local dataset, |D j | indicates client c j The size of the local dataset, Is the client c i The model parameters after the t+1th round of training, is cluster C m Aggregate model parameters at round t+1; For clients that are identified as border clients, that is, a client that belongs to multiple clusters at the same time, these clients will participate in the training process of multiple clusters. Specifically, if client c i Also belongs to cluster C m and cluster C n , the client will alternately use the corresponding cluster model for training in different training rounds, and upload the training results to the corresponding cluster server for aggregation.