Federal learning method for realizing client selection based on data feature clustering

Through singular value decomposition and hierarchical clustering, and combining contribution metrics and dynamic adjustment of learning rate, the problem of low global model accuracy caused by data heterogeneity in federated learning is solved, improving model accuracy and avoiding overfitting.

CN120387077APending Publication Date: 2025-07-29HANGZHOU DIANZI UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510446689.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the federated learning scenario, the global model accuracy is low due to data heterogeneity.

Method used

The client data features are extracted through singular value decomposition, and bottom-up hierarchical clustering is carried out. Combining the client sample capacity, data quality and number of selected times, the client contribution is quantified, and the client with the highest contribution is selected in each training to participate in training, and the learning rate is dynamically adjusted.

Benefits of technology

Effectively alleviate data heterogeneity, improve the prediction accuracy of the global model, avoid overfitting, and improve the overall performance of federated learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387077A_ABST
    Figure CN120387077A_ABST
Patent Text Reader

Abstract

The invention discloses a federated learning method for realizing client selection based on data feature clustering, which comprises the following steps of: firstly, performing singular value decomposition on local data by a client, and combining left singular vectors corresponding to first k singular values as local data features; secondly, performing hierarchical clustering on clients by taking included angles between local data features as similarity measurement standards; and then, constructing a client contribution quantification model, comprehensively considering data quality, sample capacity and selected times, refining contributions of the clients, and selecting h clients with the highest current contribution degree to participate in global training. And finally, during each global training, the server selects h clients from each group to carry out current global training, and the server carries out weighted aggregation to obtain a new global model. According to the method, the data isomerism is effectively relieved, the global model prediction precision is improved, and the problem of low global model precision caused by the data isomerism in a federated learning scene is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of federated learning, and specifically relates to a federated learning method for client selection based on data features. Background Art

[0002] With the acceleration of the digitalization process, the amount of data has increased explosively. However, data privacy and security issues have become increasingly prominent, resulting in data existing in the form of "islands", making it difficult to manage and share centrally. Traditional centralized machine learning methods are difficult to implement due to privacy regulations, while federated learning, as a distributed machine learning technology, has emerged. It allows for joint modeling by sharing model parameters without sharing the original data.

[0003] The architecture of federated learning mainly includes horizontal federated learning, vertical federated learning, and federated transfer learning. Horizontal federated learning is applicable to scenarios where there is a large overlap of samples and a small overlap of features across data centers; vertical federated learning is applicable to scenarios where there is a large overlap of features and a small overlap of samples across data centers; federated transfer learning is used for cases where both the samples and features have a small overlap across data centers. However, federated learning faces the problem of data heterogeneity, that is, the data distributions of different data centers or institutions are uneven, showing non-independent and identically distributed and non-balanced distributions, which will affect the performance of the global model.

[0004] Federated learning, as an emerging distributed machine learning technology, provides an effective way to solve the problems of data islands and privacy protection. Although data heterogeneity poses challenges, through methods such as asynchronous communication, device sampling, multi-task learning, and model heterogeneity, federated learning can make full use of distributed data resources across data centers to improve model performance while protecting privacy. In the future, federated learning is expected to be widely applied in more fields, promoting a new model of privacy protection and data sharing. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention designs and implements a federated learning method for client selection based on data feature clustering to solve the problem of low accuracy of the global model caused by data heterogeneity in the federated learning scenario.

[0006] The present invention uses singular value decomposition to extract the local client data features, and based on this, performs bottom-up hierarchical clustering to achieve client grouping; then, comprehensively considering the client sample capacity, data quality, and the number of times selected, the client contribution degree is quantified, and the client with the highest current contribution degree in each group is selected for the next global training of federated learning. The present invention selects the most representative clients to participate in the training in each training iteration and dynamically adjusts the learning rate, thereby improving the accuracy of the global model.

[0007] First, extract the client data features. Specifically, the local client uses singular value decomposition on the data of each category of the client to obtain the left singular vector, the singular value matrix, and the right singular vector, and combines the left singular vectors corresponding to the top K singular values before merging as the data features of the client. On this basis, the hierarchical clustering algorithm is used, with the angle between the left singular vectors as the similarity measure, and the similarity matrix between client pairs is calculated iteratively. With the constraint that the angle is not higher than the preset threshold, the client pairs with the smallest angle are merged layer by layer to form clustering groups, and finally a client topology with a hierarchical structure is generated.

[0008] Second, quantify the client contributions. Under the client grouping architecture generated by hierarchical clustering, the contribution of the client is quantified by relying on the data quality, sample size, and number of selections of the client. Specifically, the arithmetic sum of the top K singular values obtained by singular value decomposition is used to measure the data quality of the client; the sample size is used to represent the impact of the data scale on the global training; the number of selections is used as a negative feedback factor to avoid overfitting of some nodes. By constructing a client contribution quantification mechanism for these three factors, the contribution of the client is refined reasonably.

[0009] Finally, implement client selection with dynamically changing learning rates. Only the top n clients with the highest current contribution in each group are selected to participate in the global training. At the same time, when the cumulative participation times of the client exceed half of the total number of current global training rounds, the local learning rate of the client will be reduced.

[0010] A federated learning method for client selection based on data feature clustering includes the following steps:

[0011] S1, The client performs singular value decomposition on the local data to obtain the left singular vector, the singular value matrix, and the right singular vector, and combines the left singular vectors corresponding to the first k singular values as the local data features;

[0012] S2, Implement client hierarchical clustering, with the angle between the local data features as the similarity measure. Initially, each client is regarded as a separate client pair. By iteratively updating the similarity matrix between client pairs and using the angle not exceeding the preset threshold as the constraint condition, gradually merge the client pairs with the smallest angle, perform hierarchical clustering on the clients, construct clustering groups, and finally form a client clustering with a hierarchical structure;

[0013] S3. Construct a client contribution quantization model, comprehensively considering three factors: data quality, sample size, and the number of selections. Reflect the characteristics of the data by calculating the arithmetic sum of the first k singular values of the singular value decomposition; use the sample size to measure the impact of the data scale on global training; use the number of selections as a negative feedback factor to avoid overfitting. By constructing the client contribution quantization mechanism for these three factors, reasonably refine the contributions of the clients.

[0014] S4. In each group, only select the h clients with the highest current contribution degree to participate in global training. At the same time, when there is a client whose number of selections exceeds half of the global training times, reduce its local learning rate.

[0015] S5. During each global training, the server first selects the h clients with the highest contribution degree from each group, sends the global model, and conducts this global training; after the target client completes the training, the server performs weighted aggregation on the local model of the target client to obtain a new global model.

[0016] Furthermore, the specific process of using singular value decomposition for the local dataset in step S1 is as follows:

[0017] S11. Client i traverses the datasets of each category j For Perform singular value decomposition to obtain the left singular vector matrix and the singular value matrix

[0018] S12. Client i extracts the first k singular values of the singular value matrix of each category j, and calculates their arithmetic sum At the same time, extract the left singular vectors corresponding to the first k singular values of the left singular vector matrix of each category j Concatenate them vertically to form D i Take D i to represent the local data characteristics of client i and transmit them to the central server.

[0019] Furthermore, the process of hierarchical clustering of clients in step S2 is as follows:

[0020] S21. Calculate the angle between the data characteristics D <x m and D <x n of any two clients m and n as a similarity measure;

[0021] S22. Initialize each client as an independent clustering cluster;

[0022] S23. Calculate the similarity matrix between all clusters, and select the two clusters with the highest similarity (the smallest included angle) for merging;

[0023] S24. Update the data feature representation of the merged cluster, and recalculate the similarity between this cluster and other clusters;

[0024] S25. Repeat steps S23 and S24 until the preset stop condition is met, and generate the final clustering result.

[0025] Further, the specific formula for constructing the client contribution quantization model in step S3 is as follows:

[0026]

[0027] where dataSize i , selectedCount i are the sample size and the number of times selected for client i respectively, S i represents the arithmetic sum of the singular values of client i, and a, b, c are hyperparameters, and con i client i contribution value

[0028] Further, the specific process of step S4 is as follows:

[0029] S41. The number of clients participating in training for each group i is calculated by the following formula, where g i represents the number of clients in group i, and totalSize is the total number of clients available for selection in federated learning:

[0030]

[0031] S42. The local learning rate of each client i is determined by the following formula, where baseLR is the initial learning rate and globalSteps is the current global training times:

[0032]

[0033] Beneficial effects: The present invention represents clients with the singular value decomposition results, measures the similarity between clients with the included angle, and performs client hierarchical clustering, so that similar clients are clustered together; comprehensively considering the client data features, sample size, and the number of times selected, a contribution degree quantization model is designed, and by selecting some clients with the highest contribution degree in each group to participate in training, the data heterogeneity is effectively alleviated to improve the prediction accuracy of the global model; at the same time, the present invention also dynamically modifies the client local training learning rate to avoid the occurrence of overfitting during training. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is the flowchart of the method of the present invention. Specific embodiments

[0035] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation steps.

[0036] A federated learning method for client selection based on data feature clustering, as Figure 1 shown, includes the following steps:

[0037] Step 1: Data feature extraction. The client performs singular value decomposition on the local data set according to categories, and respectively obtains the left singular vectors and singular value matrices of each category. The left singular vectors corresponding to the top Top singular values are merged to obtain a singular value matrix, which is used as the data feature of the client; at the same time, the top Top singular values are added together as the data quality degree of the client. The client uploads the data feature and the data quality degree to the central server.

[0038] In this example, the MNIST data set is used for training, which has 10 categories of data; the MLP model is selected for the model, and the central server initializes the global model ω global . There are 100 clients that can participate in the training. The server randomly selects 23 clients each time to participate in the global training

[0039] At the same time, the client i that can participate in the training traverses the data sets of each category j For perform singular value decomposition to obtain the left singular vector matrix and the singular value matrix The formula for singular value decomposition is as follows:

[0040]

[0041] Then, the client i that can participate in the training calculates the singular value matrix of each category j The sum of the first k singular values corresponding to Vertically splice the left singular vector matrices of each category j The left singular vectors corresponding to the first k singular values of to form D i , represented by D i represent the data feature of client i and transmit it to the central server. The formula for splicing the left singular vectors is as follows:

[0042]

[0043] Step 2: Client hierarchical clustering. In this example, the included angle not exceeding a preset threshold is used as a constraint condition for client hierarchical clustering. The central server uses a bottom-up hierarchical clustering method with the data feature Di The included angle between them is used as a similarity measurement criterion. By iteratively updating the similarity matrix between clients, and using the condition that the included angle does not exceed a preset threshold as a constraint, gradually merge the pair of clients with the smallest included angle to construct clustering groups, and finally form a client topology with a hierarchical structure.

[0044] First, the server calculates the data features D of any two clients m and n m and D n The included angle between them is used as a similarity metric, and each client is initialized as an independent clustering cluster. The formula for calculating the included angle is as follows:

[0045]

[0046] Then, calculate the similarity matrix between all clustering clusters, and select the two clustering clusters with the highest similarity (the smallest included angle) for merging. Specifically, update the data feature representation of the merged clustering cluster, and recalculate the similarity between this clustering cluster and other clustering clusters.

[0047] The server repeats the above steps until the preset stop condition is met, and generates the final clustering result.

[0048] After that, each group G m Adopt the above steps to select clients repeatedly until there are 4 clients or all clients have completed grouping.

[0049] Step 3: Client contribution quantization modeling. Build a client contribution quantization model on the server side, comprehensively considering the data quality S i , sample size dataSize i and the number of selected times selectedCount i . Reflect the characteristics of the data by calculating the arithmetic sum of the top TopK singular values of the singular value decomposition; use the sample size to measure the impact of the data size on global training; use the number of selected times as a negative feedback factor to avoid overfitting. The formula for contribution quantization modeling is as follows:

[0050]

[0051] Step 4: Selection of clients with dynamically changing learning rates. Only select the several clients with the highest current contribution degree in each group to participate in global training, and at the same time weaken the learning rate of the clients that frequently participate in training.

[0052] First, for each group, its participation quota is determined by the following formula, where g i represents the number of clients in group i, and totalSize is the total number of clients available for selection, which is 100 in this example:

[0053]

[0054] Secondly, for each selected client, its local learning rate is dynamically adjusted according to the number of times it is selected. The calculation formula of the local learning rate is as follows, where baseLR is the initial learning rate, set to 0.3 in this example, and globalSteps is the current global training times:

[0055]

[0056] Step 5: Global model aggregation

[0057] The selected client i receives the global model issued by the server, trains it using local data according to the corresponding local learning rate, and obtains a new local model ω i , and sends it to the server. The server receives the local model ω i , and aggregates them with weights to form a new global model ω global , which is used for the next round of global training. The aggregation formula is as follows, where DataSize is the comprehensive sample size of the currently selected client:

[0058]

[0059] Step 6: Repeat steps 3 to 5 until the set number of global iterations is completed or the set classification accuracy is reached.

[0060] According to the steps described above, the experimental results of this embodiment are shown in Table 1. The MLP model is trained using the MNIST dataset, and the test accuracy in the figure is measured on the MNIST test set. Judging from the results, the federated optimization method of this embodiment can effectively improve the global classification accuracy of federated learning in a data heterogeneous environment. Compared with the baseline, in different data heterogeneous environments, the accuracy can be improved by up to 4.08%, showing very broad application prospects.

[0061] Table 1 Test accuracies of FedAvg, FedProx, POC, S-FedAvg, GreedyFed and this embodiment on the MNIST dataset

[0062]

Claims

1. A federated learning method for client selection based on data feature clustering, characterized in that, It includes the following steps: S1. The client performs singular value decomposition on local data to obtain left singular vectors, a singular value matrix, and right singular vectors, and combines the left singular vectors corresponding to the top k singular values as local data features; S2. Using the angle between local data features as a similarity measurement criterion, hierarchical clustering is performed on the clients to construct clustering groups, forming a hierarchical client clustering; S3. A client contribution quantification model is constructed, comprehensively considering three factors: data quality, sample size, and the number of selection times. By constructing a client contribution quantification mechanism for these three factors, the contribution of clients is refined; S4. In each group, select the top h clients with the highest current contribution degree to participate in global training. When there is a client whose number of selection times exceeds half of the global training times, reduce its local learning rate; S5. Each time global training is performed, the server first sends the global model from the top h clients with the highest contribution degree within each group for this global training; after the target client completes training, the server performs weighted aggregation on the local model of the target client to obtain a new global model.

2. The federated learning method for client selection based on data feature clustering according to claim 1, wherein The specific process of performing singular value decomposition on local data in step S1 is as follows: S11, the client i traverses the data sets of each category j For perform singular value decomposition to obtain the left singular vector matrix and the singular value matrix S12, the client i extracts the singular value matrix of each category j for the first k singular values and calculates their arithmetic sum Meanwhile, extract the left singular vector matrix of each category j for the left singular vectors corresponding to the first k singular values Concatenate them vertically to form D i , where D i represents the local data features of client i and is transmitted to the central server.

3. The federated learning method for client selection based on data feature clustering according to claim 2, wherein Step S2 is specifically implemented as: using the angle between local data features as a similarity measurement criterion. Initially, each client is regarded as a separate client pair. By iteratively updating the similarity matrix between client pairs and using the angle not exceeding a preset threshold as a constraint condition, merge the client pair with the smallest angle, perform hierarchical clustering on the clients, construct clustering groups, and finally form a hierarchical client clustering.

4. The federated learning method for client selection based on data feature clustering according to claim 3, wherein The specific implementation process of performing hierarchical clustering on clients is as follows: S21, calculate the data feature D of any two clients m and n m and D n the included angle between them as the similarity measure; S22. Initialize each client as an independent clustering cluster; S23. Calculate the similarity matrix between all clustering clusters, and select the two clustering clusters with the highest similarity for merging; S24. Update the data feature representation of the merged clustering cluster, and recalculate the similarity between this clustering cluster and other clustering clusters; S25. Repeat steps S23 and S24 until a preset stop condition is met, generate the final clustering result, and complete hierarchical clustering of the clients.

5. The federated learning method for implementing client selection based on data feature clustering according to claim 4, wherein The client contribution quantification mechanism for the three factors of data quality, sample size, and the number of selection times in step S3 is specifically as follows: calculate the arithmetic sum of the top k singular values of the singular value decomposition to reflect the characteristics of the data; use the sample size to measure the impact of the data size on global training; use the number of selection times as a negative feedback factor to avoid overfitting.

6. The federated learning method for client selection based on data feature clustering according to claim 5, characterized in that, The construction of the client contribution quantification model in step S3 is as follows: Among them, dataSize i , selectedCount i are the sample size and the number of times selected for client i respectively, S i represents the arithmetic sum of singular values of client i, and a, b, c are hyperparameters, con i Contribution value of client i.

7. The federated learning method for client selection based on data feature clustering according to claim 6, wherein The specific implementation process of step S4 is as follows: S41, the number of clients for each group \(i\) participating in training is calculated by the following formula, where \(g\) i represents the number of clients within group \(i\), and totalSize is the total number of clients available for federated learning: S42. The local learning rate of each client i is determined by the following formula, where baseLR is the initial learning rate and globalSteps is the current global training times:

Citation Information

Cited By

  • Government affair SaaS architecture and method based on cloud edge collaboration

    CN120598757A

  • Model training method, system, equipment and medium

    CN120781110A

  • A model training method, system, device and medium

    CN120781110B