A Federated Aggregation Method and Device for Data Imbalance
By constructing data quality vectors and clustering analysis, the model oscillation problem of the federal aggregation algorithm in the case of data imbalance is solved, and a more efficient federal aggregation and training process is achieved.
Patent Information
- Application Number
- CN202310182577.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-01
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-03-01
AI Technical Summary
In the face of data imbalance, existing federal aggregation algorithms cannot effectively evaluate data distribution and quality, resulting in increased model oscillation and training costs.
By constructing a data quality vector, combining gradient factors, distribution factors and quantity factors, cluster analysis of participants is carried out to realize global gradient calculation of grouped aggregation gradients.
This method measures data set differences through multiple angles, improves communication efficiency, optimizes federal aggregation methods, reduces model oscillation, and improves training efficiency.
Smart Images

Figure CN116340790B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of federated learning, and in particular provides a federated aggregation method and device for data imbalance. Background Art
[0002] With the development and application of artificial intelligence, the value of data has become increasingly prominent. In practice, how to protect user privacy while using data is a major challenge in data applications. Against this background, federated learning has emerged. Federated learning uses model gradient transmission to train information, ensuring that the training task is completed without the user data leaving the local area, achieving the purpose of protecting data privacy.
[0003] During the data accumulation process, due to factors such as usage habits and task differences, the data categories of data owners vary significantly, and the data quality of each party is uneven. The currently most widely used federated aggregation algorithm is the federated average algorithm, which uses the number of data sets owned by the participating parties as the weight of the gradients of the participating parties to achieve gradient aggregation. When facing data-imbalanced data sets, it often performs poorly.
[0004] The main problems it faces include:
[0005] 1. Lack of evaluation, quantification, and modeling of data distribution and data quality. In practice, only calculating weights from the quantity cannot solve the model oscillation caused by distribution differences in multi-classification.
[0006] 2. Unable to solve the problem of non-independent and identically distributed data, which makes it difficult for traditional federated aggregation algorithms to grasp the model update direction and increases the training cost. Summary of the Invention
[0007] The present invention aims at the above-mentioned deficiencies of the prior art and provides a practical federated aggregation method for data imbalance.
[0008] A further technical task of the present invention is to provide a federated aggregation device for data imbalance with reasonable design, safety, and applicability.
[0009] The technical solution adopted by the present invention to solve its technical problems is:
[0010] A federated aggregation method for data imbalance has the following steps:
[0011] S1. Construct a data quality vector, which consists of a gradient factor, a distribution factor, and a quantity factor;
[0012] S2. Use the data quality vector as the clustering feature to perform clustering analysis on the participating parties to achieve grouping of the participating parties;
[0013] S3. Based on the method of grouping and aggregating gradients, complete the calculation of the global gradient.
[0014] Further, in step S1, it further includes:
[0015] S101. The participating party obtains the global gradients of the model trained in the previous t - 1 rounds from the central server and updates the local model parameters
[0016] S102. The participating party conducts the t - th round of model training based on the local data to obtain the gradients of each neuron in the local model Meanwhile, take the global gradient value of the previous round Divide the gradients with each network layer in the neural network as the basic unit and calculate and gradient offset;
[0017] S103. The participating parties count the data volumes of their respective data sets and perform normalization;
[0018] S104. Each participating party calculates the KL divergence between its own data set and the uniform distribution as the distribution difference of the data sets when the participating parties participate in training evenly, denoted as
[0019] S105. Construct a data quality vector, denoted as
[0020] Further, in step S102, calculate and gradient offset as a factor of the data quality vector, denoted as to measure the influence of the current data set on the optimization direction of the model. The measurement criterion for gradient offset selects the inner product of vectors, and each value is between [0, 1];
[0021] Select a 3 - layer fully - connected neural network. After dividing the gradients with the network layer as the basic unit, it contains 3 vectors as follows. Each vector represents the gradient information of the corresponding network layer:
[0022] represents the gradients of each neuron in the first - layer fully - connected layer, which is a vector;
[0023] represents the gradients of each neuron in the first - layer fully - connected layer after local training, which is a vector;
[0024]
[0025] where the symbol represents the inner - product operation, and the offset result The result example is: [0.2, 0.5, 0.8], and each layer structure of the network corresponds to a value.
[0026] Further, in step S103, the participating parties count the data volume of their respective data sets and perform normalization, denoted as As the second factor of the data quality vector, where Its calculation formula is as follows:
[0027]
[0028] where n represents the number of participating parties, i represents the i-th participating party, and D represents the data ownership of the participating party.
[0029] Further, in step S2, it further includes:
[0030] S201. The participating parties upload the data quality vector, to the central server;
[0031] S202. Assign different weights α, β, γ to the three features in the quality vector, where α > β > γ and α + β + γ = 1;
[0032] S203. At the central server, based on the clustering algorithm, complete the clustering process to obtain the clusters cluster and the participating parties within the clusters. The distance metric in the clustering process uses the weighted Euclidean distance, and the weights are α, β, γ set in step 2;
[0033] Further, in step S203, after clustering, each cluster and the participating parties under the cluster are obtained. Taking the clustering of five participating parties A, B, C, D, and E as an example;
[0034] Through clustering, the participating parties are divided into different clusters according to the data quality. The participating parties in the same cluster are similar in data distribution and contribution to the gradient, and the data quality is consistent.
[0035] Further, in step S3, it further includes:
[0036] S301. Each participating party uploads the training gradient of this round to the central server;
[0037] S302. At the central server, traverse the clusters generated by the clustering in step S2, aggregate the gradients of the participating parties in the same cluster. Since the data of each participating party in the same cluster is similar, the mean value is used for calculation;
[0038] S303. Aggregate the gradients between clusters. The differences between clusters reflect the differences in data of different qualities. Use the federated averaging algorithm to aggregate to obtain the global gradient of this round, denoted as
[0039] S304. The central server distributes the global gradient to each participating party, and each participating party updates the weights to complete this round of training.
[0040] Further, in step S302, after the aggregation is completed, each cluster obtains its own aggregated gradient, and the gradient aggregation method within the cluster is as follows:
[0041]
[0042] In the formula represents the aggregated gradient value of the c-th cluster, n represents that there are n participating parties in this cluster, represents the gradient calculated by the participating party in this training.
[0043] Further, in step S303, the difference is the aggregation weight, which changes from the node data volume to the data volume of all nodes within the cluster.
[0044] A federated aggregation device for data imbalance includes: at least one memory and at least one processor;
[0045] The at least one memory is used to store machine-readable programs;
[0046] The at least one processor is used to call the machine-readable program to execute a federated aggregation method for data imbalance.
[0047] Compared with the prior art, a federated aggregation method and device for data imbalance of the present invention has the following outstanding beneficial effects:
[0048] The present invention constructs a data quality description vector, which weighs the quantity, quality, and model contribution of the data of each participating party in the case of data imbalance, measures the differences between data sets from multiple angles, and the clustering analysis based on this vector can greatly improve the communication efficiency.
[0049] This method optimizes the federated aggregation method. The aggregation method adopts a grouped aggregation method, so that approximate data is uniformly trained and different data is aggregated later, reducing the model oscillation caused by the data distribution difference. Description of the Drawings
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0051] Appendix Figure 1 It is a schematic flow diagram of a federated aggregation method for data imbalance. Specific implementation manner
[0052] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] The following gives an optimal embodiment:
[0054] As Figure 1 shown, in this embodiment, a federated aggregation method for data imbalance has the following steps:
[0055] S1. Construct a data quality vector, which consists of a gradient factor, a distribution factor, and a quantity factor;
[0056] It further includes:
[0057] S101. The participating party obtains the global gradient of the model trained in the (t - 1)th round from the central server and uses it to update the local model parameters
[0058] S102. The participating party performs the t-th round of model training based on local data to obtain the gradients of each neuron in the local model Meanwhile, take the global gradient value of the previous round Divide the gradients with each network layer in the neural network as the basic unit, and calculate and the gradient offset, as a factor of the data quality vector, denoted as It measures the influence of the current data set on the optimization direction of the model. The measurement criterion for the gradient offset selects the inner product of vectors, and its values are all between [0, 1].
[0059] Taking a 3-layer fully connected neural network as an example, the gradients divided by network layer as the basic unit are as follows, including a total of 3 vectors, and each vector represents the gradient information of the corresponding network layer:
[0060] represents the gradients of each neuron in the first-layer fully connected layer, and is a vector.
[0061] represents the gradients of each neuron in the first-layer fully connected layer after local training, and is a vector.
[0062]
[0063] Among them, the symbol represents the inner product operation, and the result example of the offset result is: [0.2, 0.5, 0.8]. Each layer structure of the network corresponds to a value.
[0064] S103. The participating parties count the data volume of their respective data sets and perform normalization, denoted as as the second factor of the data quality vector, where its calculation formula is as follows:
[0065]
[0066] where n represents the number of participating parties, i represents the i-th participating party, and D i represents the data ownership of the participating party.
[0067] S104. Each participating party calculates the KL divergence between its own data set and the uniform distribution (or a standard data set specified by the user) as the distribution difference of the data sets when the participating parties participate in training, denoted as
[0068] S105. Construct the data quality vector, denoted as
[0069] So far, the construction of the data quality vector is completed. The vector measures the data set from three levels: data content, data distribution quality, and model change.
[0070] S2. Use the data quality vector as the clustering feature to perform clustering analysis on the participating parties to achieve grouping of the participating parties;
[0071] Further including:
[0072] S201. The participating parties upload the data quality vector: to the central server. The advantage of using the quality vector is that the number of parameters is small and the communication cost is low.
[0073] S202. Assign different weights α, β, γ to the three features in the quality vector where α > β > γ and α + β + γ = 1.
[0074] S203. At the central server, based on the clustering algorithm, complete the clustering process to obtain the cluster cluster and each participating party within the cluster. The distance metric in the clustering process uses the weighted Euclidean distance, and the weights are α, β, γ set in step 2.
[0075] After clustering, each cluster and the participating parties under the cluster are obtained. Taking the clustering of five participating parties ABCDE as an example, the result is as follows: {1: [A B C], 2: [D E]}, indicating that the clustering generates 2 clusters. The first cluster contains three participating parties A, B, and C, and the second cluster contains two participating parties D and E.
[0076] Through clustering, the participating parties are divided into different clusters according to data quality. The participating parties in the same cluster are approximate in data distribution and contribution to the gradient, and the data quality is consistent.
[0077] S3. Complete the global gradient calculation based on the method of grouped aggregation gradient;
[0078] Further includes:
[0079] S301. Each participating party uploads the gradient of this training to the central server;
[0080] S302. At the central server, traverse the clusters generated by the clustering in step S2, and aggregate the gradients of the participating parties in the same cluster. Since the data of each participating party in the same cluster is similar, the mean value is used for calculation. This step can ensure that similar data shares the same weight and is not determined by the data volume. After aggregation, each cluster obtains its own aggregated gradient. The gradient aggregation method within the cluster is as follows:
[0081]
[0082] In the formula represents the aggregated gradient value of the c-th cluster, and n represents that there are n participating parties in this cluster;
[0083] represents the gradient calculated by the participating party in this training.
[0084] S303. Aggregate the gradients between clusters. The differences between clusters reflect the differences in data of different qualities. Use the federated averaging algorithm for aggregation to obtain the global gradient of this round, denoted as The difference is the aggregation weight, which changes from the node data volume to the data volume of all nodes within the cluster, ensuring that nodes with consistent distribution but small data volume can also contribute their gradients.
[0085] S304. The central server distributes the global gradient to each participating party, and each participating party updates the weight to complete this round of training.
[0086] Based on the above method, a federated aggregation device for data imbalance in this embodiment includes: at least one memory and at least one processor;
[0087] The at least one memory is used to store machine-readable programs;
[0088] The at least one processor is used to call the machine-readable programs and execute a federated aggregation method for data imbalance.
[0089] The above specific embodiments are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific embodiments. Any appropriate changes or substitutions made by any person of ordinary skill in the art that meet the claims of a federated aggregation method and device for data imbalance of the present invention shall fall within the patent protection scope of the present invention.
[0090] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A federated aggregation method for data imbalance, characterized in that It has the following steps: S1. Construct a data quality vector, which consists of a gradient factor, a distribution factor, and a quantity factor; It further includes: S101. The participant obtains the global gradients of the model for the (t - 1)-th round of training from the central server and updates the local model parameters S102. The participating parties perform the t-th round of model training based on local data to obtain the gradients of each neuron in the local model Meanwhile, take the global gradient value of the previous round Divide the gradients with each network layer in the neural network as the basic unit and calculate and gradient offset; Calculation and Gradient offset, as a factor of the data quality vector, is denoted as Measure the influence of the current data set on the optimization direction of the model. The measurement criterion for gradient offset selects the inner product of vectors, and each value is between [0, 1]; Select a 3-layer fully connected neural network. After dividing by the network layer as the basic unit, the gradients are as follows, and there are a total of 3 vectors, each vector representing the gradient information of the corresponding network layer: Indicates the gradient of each neuron in the first fully connected layer, which is a vector; Indicates the gradients of each neuron after the first - layer fully - connected local training, which is a vector; where the symbol ° represents the inner product operation, and the result of the offset The result example is: [0.2, 0.5, 0.8], and each layer structure of the network corresponds to a value; S103. The participating parties count the data volume of their respective data sets and perform normalization; The participants count the data volumes of their respective data sets and perform normalization, denoted as as the second factor of the data quality vector, where its calculation formula is as follows: where n represents the number of participants, i represents the i-th participant, and D i represents the data ownership of the participant; S104. Each participant calculates the KL divergence between its own dataset and the uniform distribution as the distribution difference of the dataset when the balanced participants participate in the training, denoted as S105. Construct a data quality vector, denoted as S2. Use the data quality vector as the clustering feature to perform clustering analysis on the participating parties to achieve grouping of the participating parties; It further includes: S201. The participating parties upload the data quality vectors to the central server. To the central server; S202. Assign different weights α, β, and γ to the three features in the quality vector, where α > β > γ and α + β + γ = 1; S203. At the central server, based on the clustering algorithm, complete the clustering process to obtain clusters cluster and each participating party within the cluster. The distance metric in the clustering process uses the weighted Euclidean distance, and the weights are αβγ set in step 2; After clustering, each cluster and the participating parties under the cluster are obtained. Taking the clustering of five participating parties A, B, C, D, and E as an example; Through clustering, the participating parties are divided into different clusters according to data quality. The participating parties in the same cluster are approximate in data distribution and contribution to the gradient, and the data quality is consistent; S3. Based on the method of aggregating gradients by grouping, complete the global gradient calculation; It further includes: S301. Each participating party uploads the training gradients for this time to the central server; S302. At the central server, traverse the clusters generated by the clustering in step S2, aggregate the gradients of the participating parties in the same cluster. Since the data of each participating party in the same cluster is similar, the mean value is used for calculation; S303. Inter-cluster gradient aggregation. The differences in data of different qualities are reflected between clusters. Aggregate using the federated averaging algorithm to obtain the global gradient of this round, denoted as S304. The central server distributes the global gradient to each participant, and each participant updates the weights to complete this round of training.
2. The federated aggregation method for data imbalance according to claim 1, wherein, In step S302, after aggregation, each cluster obtains its own aggregated gradient. The gradient aggregation method within the cluster is as follows: In the formula represents the aggregation gradient value of the c-th cluster, and n represents that there are n participants in this cluster. represents the gradient calculated by the participant in this training.
3. The federated aggregation method for data imbalance according to claim 2, wherein, In step S303, the difference is the aggregation weight, which changes from the node data volume to the data volume of all nodes within the cluster.
4. A federated aggregation device for data imbalance, characterized in that, It includes: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program and execute the method described in any one of claims 1 to 3.
Citation Information
Patent Citations
Improved gas impermeability for injection molded containers
CN103003048A
Complex environment indoor fingerprint positioning method based on integrated federated learning
CN114205905A