Group Training Method, Server and Client Based on Distributed Machine Learning
By grouping and updating grouping according to the client's local optimization target gradient in a distributed machine learning system, the problems of sub-optimization target conflict and data distribution imbalance in traditional systems are solved, and model performance and training stability are improved.
Patent Information
- Application Number
- CN202111232230.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-10-22
AI Technical Summary
Traditional distributed machine learning systems are unable to effectively deal with the hidden suboptimal target conflicts in tasks, resulting in a degradation of model performance and increase the bias of machine learning models due to the imbalance in the distribution of training data.
By grouping clients according to the local optimization target gradient of each client, creating corresponding group processes and group consensus models, each client is assigned to the group process with the most similar optimization target gradient for training, and updating the grouping when the training data distribution is offset to reduce model bias caused by data distribution imbalance.
It effectively alleviates the problems of decreased training convergence speed, decreased model performance and increased training jitter due to sub-optimization target conflicts, and reduces machine learning model bias due to imbalance in the distribution of training data.
Smart Images

Figure CN114118210B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and particularly to a grouped training method, a server and a client based on distributed machine learning. Background Art
[0002] With the advent of the third wave of artificial intelligence, more and more deep learning applications have emerged in various fields. The training of these deep learning models depends on a large amount of training data, and the service provider collects the training data from user devices to perform model training.
[0003] Traditional distributed machine learning systems only train a consensus model and cannot handle the situation of conflicts in sub-optimization goals implicit in tasks, further leading to a decline in the model performance of the training system. In addition, in traditional distributed machine learning systems, the client storing the training data undertakes the model training task, and the server collects and integrates the models. However, the distribution of the training set in a distributed machine learning system is usually unbalanced, and the data imbalance will increase the bias of the machine learning model, resulting in a decline in model performance. Summary of the Invention
[0004] Embodiments of the present invention provide a server and a client in multiple aspects to solve the problems that traditional distributed machine learning systems cannot handle the conflicts in sub-optimization goals in training and the training data distribution is unbalanced.
[0005] The grouped training method based on distributed machine learning provided by embodiments of the present invention is applied to a server and includes:
[0006] Grouping the clients according to the local optimization target gradients of each client to obtain N client groups and the group optimization target gradient corresponding to each client group, and creating a corresponding group process and a corresponding group consensus model based on each client group, so that each client is assigned to the group process with the most similar optimization target gradient for training the group consensus model; N is greater than or equal to 1, where the local optimization target gradient is calculated by the client based on local training data and an initial model sent by the server;
[0007] After the previous training round ends and before the current training round starts, receiving the grouped update information returned by the client with an offset in the training data distribution; where the grouped update selection information of the client is the group process with the highest similarity of the optimization target gradient selected by the client with an offset in the training data distribution based on the comparison result of the similarity between the current local optimization target gradient and each group optimization target gradient;
[0008] Update the grouping of the clients according to the grouping update information, so that each client is assigned to the corresponding group process in the current training round for training the group consensus model.
[0009] As an improvement to the above solution, grouping the clients according to the local optimization target gradients of each client, obtaining N client groups and the group optimization target gradient corresponding to each client group, and creating a corresponding group process and a corresponding group consensus model for each client group, so that each client is assigned to the group process with the most similar optimization target gradient for training the group consensus model, specifically including:
[0010] Send the initial model to all clients, so that all clients calculate the corresponding local optimization target gradients based on the local training data and the initial model;
[0011] Receive the local optimization target gradients uploaded by all clients, and cluster the optimization target gradients of all clients to obtain N client groups and the group optimization target gradient corresponding to each client group;
[0012] Based on each client group, create a group process and a group consensus model corresponding to each client group, so that each client is assigned to the group process with the most similar optimization target gradient for training the group consensus model.
[0013] As an improvement to the above solution, grouping the clients according to the local optimization target gradients of each client, obtaining N client groups and the group optimization target gradient corresponding to each client group, and creating a corresponding group process and a corresponding group consensus model for each client group, so that each client is assigned to the group process with the most similar optimization target gradient for training the group consensus model, specifically including:
[0014] Select M first clients among all clients, and send the initial model to all the first clients, so that each first client obtains the corresponding local optimization target gradient based on the local training data and the initial model, where M is less than N;
[0015] Receive the local optimization target gradients of each first client, and cluster the local optimization target gradients of all the first clients to obtain multiple client groups and the group optimization target gradient corresponding to each client group;
[0016] Send the initial model and each group optimization target gradient to the second clients, so that the second clients obtain the local optimization target gradients based on the local training data and the initial model; where the second clients refer to the clients other than the first clients among all clients;
[0017] Receive the group selection results returned by the second clients, update the client groups, and assign each of the second clients to the corresponding group process for training the group consensus model; wherein, the group selection result of the second client is the group process with the highest similarity of the optimization gradient selected by the second client based on the comparison result of the similarity between the local optimization objective gradient and each group optimization objective gradient.
[0018] As an improvement to the above solution, the client with the offset in the training data distribution is determined by the following method:
[0019] When the sum of the deviation distances of all label categories of the client in the current training round is greater than the preset deviation distance, it is determined that the training data distribution of the client has an offset; wherein, the label category refers to the category of the training sample, and the deviation distance of each label category is the deviation distance between the label category probability distribution of the client in the current training round and the label category probability distribution at the time of the last offset.
[0020] As an improvement to the above solution, the deviation distance between the label category probability distribution of the client in the current training round and the label category probability distribution at the time of the last offset is calculated by the following formula:
[0021]
[0022] where d c represents the deviation distance, EMD(·) represents the EMD distance between two distributions, and P c is the label category probability distribution of the client in the current training round, is the label category probability distribution of the client at the time of the last offset.
[0023] As an improvement to the above solution, the steps of receiving the local optimization objective gradients uploaded by all clients and clustering the optimization objective gradients of all clients to obtain N client groups and the corresponding group optimization objective gradients for each client group specifically include:
[0024] Receive the optimization objective gradients uploaded by each client, and calculate the Euclidean distance based on the compressed cosine similarity between each client and other clients according to the following formula:
[0025]
[0026] V = SVD(ΔW T , N)
[0027] Among them, EDC(i,j) is the Euclidean distance between client i and client j based on compressed cosine similarity, N is the number of client groups, V is the gradient matrix compressed by the machine learning algorithm SVD, v is the column vector of the gradient matrix V, S(i,v T ) is the model update vector of client i, S(j,v T ) is the model update vector of client j, ΔW T is the transpose matrix of the matrix composed of the client model update vectors;
[0028] Execute a clustering algorithm based on all the calculated Euclidean distances to obtain N client groups and the group-optimized target gradients corresponding to each client group.
[0029] As an improvement of the above solution, the similarity between the local optimized target gradient and each group-optimized target gradient is calculated by the following formula:
[0030]
[0031] Among them, represents the group-optimized target gradient of group process g j , represents the local optimized target gradient of client i, sim(i,j) represents the similarity between the local optimized target gradient of client i and the group-optimized target gradient of group process g j . represents the angle between and
[0032] The second aspect of the embodiments of the present invention provides a grouping training method based on distributed machine learning, which is applied to a client and includes:
[0033] Based on the local training data and the initial model sent by the server, obtain the local optimized target gradient, and send the local optimized target gradient to the server, so that the server performs operations on client grouping, calculates the group-optimized target gradients of each client group, creates group processes corresponding to each client group, and initializes the group consensus models corresponding to each client group according to the local optimized target gradients of each client; wherein, the client group is obtained based on the grouping result of the clients;
[0034] Receive the grouping result and the corresponding group consensus model returned by the server, and perform training on the group consensus model in the corresponding group process;
[0035] After the previous training round ends and before the current training round starts, calculate the local training data distribution offset result;
[0036] When the training data distribution offset result is an offset, calculate the similarity between the current local optimization objective gradient and each group of optimization objective gradients, and send group selection information to the group process with the highest similarity of the optimization objective gradients, so that the server updates the grouping of the clients and assigns the clients to the corresponding group processes in the current training round for training the group consensus model.
[0037] A third aspect of the embodiments of the present invention provides a server, which includes:
[0038] A group process creation module, configured to group clients according to the local optimization objective gradients of each client, obtain N client groups and the group optimization objective gradients corresponding to each client group, and create corresponding group processes and corresponding group consensus models based on each client group, so that each client is assigned to the group process with the most similar optimization objective gradient for training the group consensus model; N is greater than or equal to 1, where the local optimization objective gradient is calculated by the client based on local training data and the initial model sent by the server;
[0039] A grouping update information receiving module, configured to receive the grouping update information returned by the client with an offset in the training data distribution after the training of the previous training round ends and before the training of the current training round starts; where the client's grouping update selection information is the group process with the highest similarity of the optimization objective gradients selected by the client with an offset in the training data distribution based on the comparison result of the current local optimization objective gradient and each group of optimization objective gradients;
[0040] A client grouping update module, configured to update the grouping of the clients according to the grouping update information, so that each client is assigned to the corresponding group process in the current training round for training the group consensus model.
[0041] A fourth aspect of the embodiments of the present invention provides a client, including:
[0042] A local optimization objective gradient obtaining module, configured to obtain the local optimization objective gradient based on local training data and the initial model sent by the server, and send the local optimization objective gradient to the server, so that the server performs operations on client grouping, calculates the group optimization objective gradients of each client group, creates group processes corresponding to each client group, and initializes the group consensus models corresponding to each client group according to the local optimization objective gradients of each client; where the client group is obtained based on the grouping result of the clients;
[0043] A group consensus model training module, configured to receive the grouping result and the corresponding group consensus model returned by the server, and train the group consensus model in the corresponding group process;
[0044] A data distribution offset result calculation module, configured to calculate a local training data distribution offset result after the end of the previous training round and before the start of the current training round;
[0045] A grouping selection module, configured to calculate the similarity between the current local optimization objective gradient and the optimization objective gradient of each group when the training data distribution offset result is an offset, and send grouping selection information to the group process with the highest similarity of the optimization objective gradient, so that the server updates the grouping of the client and assigns the client to the corresponding group process in the current training round for training the group consensus model.
[0046] Compared with the prior art, the grouping training method, server and client provided by the embodiments of the present invention have the following beneficial effects:
[0047] The grouping training method based on distributed machine learning provided by the embodiments of the present invention includes grouping clients according to the local optimization objective gradient of each client to obtain N client groups and the corresponding group optimization objective gradient of each client group, and creating corresponding group processes and corresponding group consensus models based on each client group, so that each client is assigned to the group process with the most similar optimization objective gradient for training the group consensus model; after the end of the previous training round and before the start of the current training round, receiving the grouping update information returned by the client whose training data distribution has shifted; updating the grouping of the client according to the grouping update information, so that each client is assigned to the corresponding group process in the current training round for training the group consensus model. It splits the training task into the corresponding group consensus models according to the optimization objective based on the local optimization objective gradient of the client, alleviating problems such as the decline in training convergence speed, the decline in model performance, and the increase in training jitter caused by the conflict of implicit sub-optimization objectives. At the same time, the present invention also monitors the local training data offset situation of the client during the training process. When the local training data of the client shifts, the grouping information of the client is updated, so that each client can train in the group model with the most similar optimization objective, effectively reducing the deviation problem of the machine learning model caused by the imbalance of the training data distribution. Correspondingly, the present invention also provides a server and a client. Description of the Drawings
[0048] Figure 1 is a flowchart of the grouping training method based on distributed machine learning provided by Embodiment 1 of the present invention;
[0049] Figure 2 is a flowchart of one implementation manner of the grouping training method based on distributed machine learning provided by Embodiment 1 of the present invention;
[0050] Figure 3 It is a structural block diagram of the grouped training method for distributed machine learning provided in Embodiment 5 of the present invention;
[0051] Figure 4 It is a system framework diagram of the grouped training method for distributed machine learning provided by the present invention. Detailed implementation manners
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] Example 1
[0054] See Figure 1 , Figure 1 It is a schematic flowchart of the grouped training method for distributed machine learning provided in Embodiment 1 of the present invention.
[0055] The grouped training method for distributed machine learning provided in Embodiment 1 of the present invention is applied to a server and includes:
[0056] Step S100: Group the clients according to the local optimization objective gradients of each client to obtain N client groups and the group optimization objective gradient corresponding to each client group, and create a corresponding group process and a corresponding group consensus model based on each client group, so that each client is assigned to the group process with the most similar optimization objective gradient for training the group consensus model; N is greater than or equal to 1, where the local optimization objective gradient is calculated by the client based on local training data and the initial model sent by the server;
[0057] Step S110: After the previous training round ends and before the current training round starts, receive the grouped update information returned by the client with the shifted training data distribution; where the grouped update selection information of the client is the group process with the highest similarity of the optimization objective gradient selected by the client with the shifted training data distribution based on the comparison result of the similarity between the current local optimization objective gradient and each group optimization objective gradient;
[0058] Step S120: Update the grouping of the clients according to the grouped update information, so that each client is assigned to the corresponding group process for training the group consensus model in the current training round.
[0059] Specifically, the server refers to a cloud computing device with network communication capabilities and a processor cluster. The client refers to a mobile device with network communication capabilities and at least one processor. The group process, also known as the intermediate process, is an instance of a program running on a computer system and can be deployed on any device with communication and computing capabilities, such as a cloud server, a mobile device, or an edge server. In the model training stage, the group process broadcasts the group consensus model to the clients. Each client calculates the local model update and uploads the update back to the corresponding group process. The corresponding group process uses the federated averaging algorithm for model integration, which is a synchronous model update method.
[0060] Exemplarily, in step S100, the local optimized objective gradient can execute the neural network backpropagation algorithm based on the initial model and the local dataset, such as the stochastic gradient descent method with a small gradient, to obtain the model update vector, that is, the local objective optimized gradient direction. Finally, the client uploads the local objective optimized gradient to the server.
[0061] Specifically, before the server enters the training of the initial model, based on the client grouping result, the group optimized objective gradient, the group process, and the group consensus model corresponding to the group process obtained in step S100, the training of the clients is assigned to the group process with the most similar optimized objective, so that all the clients in this group process cooperate to train the group consensus model. In this process, the group process uses the federated averaging algorithm to refresh the group consensus model to obtain the group consensus model that has been trained for one round. At this time, each client will calculate the local training data distribution offset result before the next round of group consensus model training and determine whether to update the grouping according to the offset result. When the client determines that the local training data distribution has shifted, it calculates the similarity between the local optimized objective gradient and each group optimized objective gradient and selects the group process with the highest similarity to send the grouping selection information, so that the server re-groups the clients and repeats step S110 and step S120 until the training accuracy of all group consensus models meets the requirements or the training error can no longer continue to decrease. Then, all the trained group consensus models are integrated to obtain the trained initial model.
[0062] The grouped training method based on distributed machine learning provided by the embodiments of the present invention includes grouping clients according to the local optimization objective gradients of each client to obtain N client groups and the corresponding group optimization objective gradients for each client group, and creating corresponding group processes and corresponding group consensus models based on each client group, so that each client is assigned to the group process with the most similar optimization objective gradient for training the group consensus model; after the end of the previous training round and before the start of the current training round, receiving the grouped update information returned by the clients with offset training data distributions; updating the grouping of the clients according to the grouped update information, so that each client is assigned to the corresponding group process for training the group consensus model in the current training round. It splits the training tasks into corresponding group consensus models according to the optimization objectives based on the local optimization objective gradients of the clients, reducing problems such as the decline in training convergence speed, the decline in model performance, and the increase in training jitter caused by conflicts between implicit sub-optimization objectives. At the same time, the present invention also monitors the offset situation of the local training data of the clients during the training process. When the local training data of the client is offset, the grouping information of the client is updated, so that each client can train in the group model with the most similar optimization objective, effectively reducing the bias problem of the machine learning model caused by the imbalance of the training data distribution.
[0063] In one implementation manner, step S100 "grouping clients according to the local optimization objective gradients of each client to obtain N client groups and the corresponding group optimization objective gradients for each client group, and creating corresponding group processes and corresponding group consensus models based on each client group, so that each client is assigned to the group process with the most similar optimization objective gradient for training the group consensus model" specifically includes:
[0064] Sending an initial model to all clients so that all the clients calculate the corresponding local optimization objective gradients based on the local training data and the initial model;
[0065] Receiving the local optimization objective gradients uploaded by all clients, and clustering the optimization objective gradients of all clients to obtain N client groups and the corresponding group optimization objective gradients for each client group;
[0066] Creating a group process and a group consensus model corresponding to each client group based on each client group, so that each client is assigned to the group process with the most similar optimization objective gradient for training the group consensus model.
[0067] Further, the receiving the local optimization objective gradients uploaded by all clients, and clustering the optimization objective gradients of all clients to obtain N client groups and the corresponding group optimization objective gradients for each client group specifically includes:
[0068] Receive the optimized target gradients uploaded by each of the clients, and calculate the Euclidean distance between each client and other clients according to each of the optimized target gradients;
[0069] Execute a clustering algorithm based on all the calculated Euclidean distances to obtain N client groups and the group optimized target gradients corresponding to each client group.
[0070] Considering the complexity of the algorithm and the effect of clustering, in a preferred implementation manner of the embodiment of the present invention, the Euclidean distance between a client and other clients refers to the Euclidean distance based on the compressed cosine similarity between the client and other clients, and the Euclidean distance based on the compressed cosine similarity between the client and other clients is calculated by the following formula:
[0071]
[0072] V = SVD(ΔW T , N)
[0073] where EDC(i, j) is the Euclidean distance based on the compressed cosine similarity between client i and client j, N is the number of client groups, V is the gradient matrix compressed by the machine learning algorithm SVD, v is the column vector of the gradient matrix V, S(i, v T ) is the model update vector of client i, S(j, v T ) is the model update vector of client j, and ΔW T is the transpose matrix of the matrix composed of the client model update vectors.
[0074] In the embodiment of the present invention, the machine learning clustering algorithm K-Means++ can be executed according to the distance defined based on EDC to obtain N cluster centers as the group optimized target gradients Then, based on the initial model w 0 refresh to obtain N corresponding group consensus models
[0075] Alternatively, in an alternative implementation manner, the receiving the local optimized target gradients uploaded by all clients and clustering the optimized target gradients of all clients to obtain N client groups and the group optimized target gradients corresponding to each client group specifically includes:
[0076] Receive the optimized target gradients uploaded by each of the clients, and calculate the cosine similarity matrix between each client and other clients according to each of the optimized target gradients;
[0077] Execute a clustering algorithm based on all the calculated cosine similarity matrices to obtain N client groups and the group-optimized target gradients corresponding to each client group.
[0078] In the embodiments of the present invention, the cosine similarity matrix between clients can also be calculated, and the clients are clustered according to the cosine similarity matrix.
[0079] Specifically, the cosine similarity matrix between each client and other clients is calculated by the following formula:
[0080]
[0081] M ij = S(i, j)
[0082]
[0083] where M is the cosine similarity matrix, M ij is the element in the i-th row and j-th column of the matrix M, Δw i and Δw j are the model update vectors of client i and client j respectively, <Δw i , Δw j > represents the vector inner product of Δw i and Δw j , ||Δw i || represents the L2 norm of Δw i , and ||Δw j || represents the L2 norm of Δw j .
[0084] Alternatively, when calculating the cosine similarity matrix between each client and other clients, it can also be calculated according to the following formula:
[0085] M = K(Δw i , Δw j )
[0086] where M is the cosine similarity matrix and K is the cosine kernel function.
[0087] The client grouping operation provided by the above embodiment requires all clients to participate in the grouped training, with complex calculations and poor ability to handle the instability of the distributed machine learning system caused by clients exiting the training due to emergencies. Therefore, the embodiments of the present invention also provide another client grouping method.
[0088] Alternatively, in another embodiment, step S100 "group the clients according to the local optimization target gradients of each client to obtain N client groups and the group optimization target gradient corresponding to each client group, and create a corresponding group process and a corresponding group consensus model for each client group, so that each client is assigned to the group process with the most similar optimization target gradient for training the group consensus model" specifically includes:
[0089] Select M first clients among all clients and send the initial model to all the first clients, so that each first client obtains the corresponding local optimization target gradient based on the local training data and the initial model, where M is less than N;
[0090] Select M first clients among all clients and send the initial model to all the first clients, so that each first client obtains the corresponding local optimization target gradient based on the local training data and the initial model, where M is less than N;
[0091] Receive the local optimization target gradients of each first client, and cluster the local optimization target gradients of all first clients to obtain multiple client groups and the group optimization target gradient corresponding to each client group;
[0092] Send the initial model and each group optimization target gradient to the second clients, so that the second clients obtain the local optimization target gradients based on the local training data and the initial model; where the second clients refer to the clients other than the first clients among all clients;
[0093] Receive the grouping selection results returned by the second clients, update the client groups, and assign each second client to the corresponding group process for training the group consensus model; where the grouping selection result of the second client is the group process with the highest similarity of the optimization gradient selected by the second client based on the comparison of the similarity between the local optimization target gradient and each group optimization target gradient.
[0094] See Figure 2 , Figure 2It is a schematic flowchart of one implementation manner of the grouped training method based on distributed machine learning provided in the first embodiment of the present invention. In this implementation manner, the server can randomly select a preset proportion of clients to send the initial model, so that the clients calculate the corresponding local optimization target gradients, and execute a clustering algorithm based on the gradients of the local optimization targets of the clients to obtain a grouping result, group optimization target gradients, and obtain a group consensus model based on the clustering result and create corresponding group processes. The specific process can be seen in Figure (a). For the remaining clients or newly added clients, the corresponding group process can be selected based on the similarity between the group optimization target gradients and the optimization target gradients of the clients. The server updates the client grouping information. At the same time, each group process sends the new group consensus model to all clients within the group, so that all clients within the group refresh the group consensus model, and all group processes integrate and refresh the group consensus model for all new group consensus models. The specific process can be seen in Figure (b).
[0095] Specifically, the server creates a group process for each group, and the grouping rules of the clients participating in the clustering are consistent with the clustering result. Each client corresponds to only one group process. There are group processes without any clients, and the unassigned clients have no corresponding group processes.
[0096] In one implementation manner, in step S120, the client whose training data distribution has shifted is judged by the following method:
[0097] When the sum of the deviation distances of all label categories of the client in the current training round is greater than the preset deviation distance, it is determined that the training data distribution of the client has shifted; wherein, the label category refers to the category of the training sample, and the deviation distance of each label category is the deviation distance between the label category probability distribution of the client in the current training round and the label category probability distribution at the time of the last shift.
[0098] When the sum of the deviation distances of all label categories of the client in the current training round is less than or equal to the preset deviation threshold, the training data distribution of the client has not shifted.
[0099] In an embodiment of the present invention, considering that in the training of a machine learning model, the sample categories of the training set held by a client may be different, therefore, the offset distances of all label categories are statistically calculated to determine whether the client training data is offset. Specifically, the probability distribution of each label category should be understood as follows: Suppose there are samples 1, 2,..., n in the training set of client A, and sample 1 corresponds to label category 1, sample 2 corresponds to label category 2, then the probability distribution of label category 1 can be understood as the probability distribution of sample 1. The probability distribution of the label category at the time of the last offset should be understood as follows: For example, in the 5th round of group consensus model training, the training sample 1 of client A is offset, there is no offset in the 6th round, and an offset occurs in the 7th round. Then, when calculating the offset distance of label category 1 in the 7th round, it is based on the offset distance between the probability distribution of the label category of the samples belonging to label category 1 in the 7th round and the probability distribution of the label category of the samples belonging to label category 1 in the 5th round.
[0100] In one implementation, the offset distance between the probability distribution of the label category of the client in the current training round and the probability distribution of the label category at the time of the last offset is calculated by the following formula:
[0101]
[0102] where d c represents the offset distance, EMD(·) represents the EMD distance between two distributions, P c is the probability distribution of the label category of the client in the current training round, is the probability distribution of the label category at the time of the last offset of the client.
[0103] Furthermore, the preset offset threshold is calculated by the following formula:
[0104]
[0105] where T represents the number of local training categories of client i, n i represents the number of training samples of client i, C is a constant between 0 and 1. The smaller C is, the more subtle the distribution change will be determined to be an offset (that is, the greater the tendency to determine an offset). Preferably, C is set to 0.2.
[0106] In one implementation, in step S110, "after the end of the training in the previous training round and before the start of the training in the current training round, receive the grouped update information returned by the client with a shifted training data distribution; wherein, the grouped update selection information of the client is the optimization target gradient with the highest similarity selected by the client with a shifted training data distribution based on the comparison result of the similarity between the current local optimization target gradient and each group optimization target gradient", the similarity between the local optimization target gradient and each group optimization target gradient is calculated by the following formula:
[0107]
[0108] Wherein, represents the group optimization target gradient of group process g j ; represents the local optimization target gradient of client i, and sim(i, j) represents the similarity between the local optimization target gradient of client i and the group optimization target gradient of group process g j . represents the included angle between and
[0109] Specifically, the smaller sim(i, j) is, the higher the similarity between the local optimization target gradient of client i and the group optimization target gradient of group process g j is.
[0110] Example 2
[0111] Embodiment 2 of the present invention provides a grouped training method based on distributed machine learning, which is applied to a client and includes:
[0112] Obtain a local optimization target gradient based on local training data and an initial model sent by a server, and send the local optimization target gradient to the server, so that the server performs operations on client grouping, calculates the group optimization target gradient of each client group, creates a group process corresponding to each client group, and initializes a group consensus model corresponding to each client group; wherein, the client group is obtained based on the grouping result of the clients;
[0113] Receive the grouping result and the corresponding group consensus model returned by the server, and perform training on the group consensus model in the corresponding group process;
[0114] Calculate the local training data distribution offset result after the end of the training in the previous training round and before the start of the training in the current training round;
[0115] When the training data distribution offset result is an offset, calculate the similarity between the current local optimization objective gradient and each group of optimization objective gradients, and send group selection information to the group process with the highest similarity of the optimization objective gradient, so that the server updates the grouping of the clients and assigns the clients to the corresponding group processes in the current training round for training the group consensus model.
[0116] It should be noted that the working principles and effects of the grouping training method based on distributed machine learning provided in the second embodiment of the present invention correspond one by one to those of the grouping training method based on distributed machine learning provided in the first embodiment. For the implementation method of client grouping, the calculation of the local training data distribution offset result of the client, etc., reference can be made to the first embodiment, and no further elaboration will be made here.
[0117] Example 3
[0118] The third embodiment of the present invention provides a server, including:
[0119] A group process creation module, configured to group clients according to the local optimization objective gradients of each client, obtain N client groups and the group optimization objective gradients corresponding to each client group, and create corresponding group processes and corresponding group consensus models based on each client group, so that each client is assigned to the group process with the most similar optimization objective gradient for training the group consensus model; N is greater than or equal to 1, where the local optimization objective gradient is calculated by the client based on local training data and the initial model sent by the server;
[0120] A grouping update information receiving module, configured to receive the grouping update information returned by the client with an offset in the training data distribution after the end of the previous training round and before the start of the current training round; where the client's grouping update selection information is the group process with the highest similarity of the optimization objective gradient selected by the client with an offset in the training data distribution based on the comparison result of the similarity between the current local optimization objective gradient and each group optimization objective gradient;
[0121] A client grouping update module, configured to update the grouping of the clients according to the grouping update information, so that each client is assigned to the corresponding group process in the current training round for training the group consensus model.
[0122] Example 4
[0123] The fourth embodiment of the present invention provides a client, including:
[0124] The local optimization target gradient acquisition module is used to obtain the local optimization target gradient based on the local training data and the initial model sent by the server, and send the local optimization target gradient to the server, so that the server performs operations such as client grouping, calculation of the group optimization target gradient for each client group, creation of the group process corresponding to each client group, and initialization of the group consensus model corresponding to each client group according to the local optimization target gradient of each client; wherein, the client group is obtained based on the client grouping result.
[0125] The group consensus model training module is used to receive the grouping result and the corresponding group consensus model returned by the server, and train the group consensus model in the corresponding group process.
[0126] The data distribution offset result calculation module is used to calculate the local training data distribution offset result after the end of the previous training round and before the start of the current training round.
[0127] The grouping selection module is used to calculate the similarity between the current local optimization target gradient and each group optimization target gradient when the training data distribution offset result is an offset, and send grouping selection information to the group process with the highest optimization target gradient similarity, so that the server updates the client grouping and assigns the client to the corresponding group process for group consensus model training in the current training round.
[0128] Example 5
[0129] See Figure 3 , Figure 3 FIG. is a structural block diagram of a grouping training method based on distributed machine learning provided in Embodiment 5 of the present invention. The process of joint participation in training by the client 31, the cloud server 33, and the group process 392 is reflected in this embodiment. Among them, the client 31 is composed of a variety of computing devices 504 with communication and storage capabilities, including portable personal computers, smart phones, sensors, smart cameras, etc. Each client 31 independently maintains a local training data set 311, and this data set 311 is stored in the memory of the device. The memory includes flash memory, solid-state drives, hard disk drives, etc.
[0130] The client 31 has wireless communication capabilities, and it sends and receives data to and from the group process 392 created by the cloud server 33 through the network 32.
[0131] The central server 33 runs a computer operating system. Using the application programming interfaces provided by the operating system, the system can create and train a task initialization module 34, a new client grouping module 35, an intermediate process creation module 38, and a grouped training module 39. Among them, the intermediate process creation module 38 includes an optimization target gradient collection process 381 and an optimization target clustering process 382. It obtains client grouping by collecting the local optimization target gradients of clients and executing a clustering algorithm, and writes the grouping rules into the grouping rule database 25. The grouped training module 39 contains multiple group processes 391. Each group process 391 includes a model collection process 3911 and a model integration process 3912. Its tasks include sending the model to the client devices ready for training through the network 12, obtaining the grouped consensus models after training for each client group, collecting the trained grouped consensus models by the model collection process 3911, taking the model as input by the model integration process 3912, and then executing a model integration algorithm such as Federated Averaging to output a group model, and saving this model in the grouped consensus group model group optimization gradient library 37. The new client grouping module 35 includes an optimization gradient collection process 351 and a gradient similarity grouping process 352. It reads the group optimization target gradients from the grouping rule database 36, calculates the gradient similarity to obtain new client grouping, and writes it into the grouping rule database 36.
[0132] See Figure 4 , Figure 4 shows the system framework diagram of the grouped training method for distributed machine learning of the present invention. For ease of understanding, the embodiments of the present invention also give the detailed process method of the grouped training method for distributed machine learning, which is applied to a distributed machine learning system and includes:
[0133] Step S1: The server distributes the initial model to a preset proportion of clients, and each client calculates its local optimization target gradient based on the local training data and the initial model.
[0134] Step S2: The server executes a clustering algorithm based on the gradients of the local optimization targets of the clients obtained in Step S1 to obtain the group optimization target gradients of the group model.
[0135] Step S3: The server obtains the grouped consensus model based on the clustering result of Step S2 and creates corresponding group processes.
[0136] Step S4: The ungrouped clients or the clients with data distribution deviation select the corresponding group processes based on the similarity between the group optimization target gradients of Step S2 and their local optimization target gradients, and the server updates the client grouping.
[0137] S5: The clients and the group processes jointly train the grouped consensus model based on the grouping rules obtained in Step S4.
[0138] S6: During the training process, monitor the data distribution deviation of each client. If it exceeds the predetermined threshold, re-execute step S4 to change the grouping rule.
[0139] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A grouped training method based on distributed machine learning, which is applied to a server, characterized in that, it includes: Sending an initial model to all clients, so that all the clients calculate corresponding local optimization objective gradients based on local training data and the initial model; Receiving the local optimization objective gradients uploaded by all clients, and clustering the optimization objective gradients of all clients to obtain N client groups and the group optimization objective gradients corresponding to each client group; Based on each client group, creating a group process and a group consensus model corresponding to each client group, so that each client is assigned to the group process with the most similar optimization objective gradients for training the group consensus model; N is greater than or equal to 1, where the local optimization objective gradient is calculated by the client based on local training data and the initial model sent by the server; After the end of the previous training round and before the start of the current training round, receiving the grouped update information returned by the clients with shifted training data distributions; wherein, the grouped update selection information of the client is the group process with the highest similarity of optimization objective gradients selected by the client with the shifted training data distribution based on the comparison result of the similarity between the current local optimization objective gradient and each group optimization objective gradient; Updating the grouping of the clients according to the grouped update information, so that each client is assigned to the corresponding group process in the current training round for training the group consensus model; The client with the shifted training data distribution is judged by the following method: When the sum of the deviation distances of all label categories of the client in the current training round is greater than a preset deviation distance, it is determined that the training data distribution of the client has shifted; wherein, the label category refers to the category of the training sample, and the deviation distance of each label category is the deviation distance between the label category probability distribution of the client in the current training round and the label category probability distribution at the time of the last shift; The deviation distance between the label category probability distribution of the client in the current training round and the label category probability distribution at the time of the last shift is calculated by the following formula: where d c represents the offset distance, EMD(·) represents the EMD distance between two distributions, and P c is the probability distribution of the label classes of the client at the current training round, and is the probability distribution of the label classes at the time of the last offset on the client; The step of receiving the local optimization objective gradients uploaded by all clients, and clustering the optimization objective gradients of all clients to obtain N client groups and the group optimization objective gradients corresponding to each client group specifically includes: Receiving the optimization objective gradients uploaded by each client, and calculating the Euclidean distance based on the compressed cosine similarity between each client and other clients according to the following formula: V = SVD(ΔW T , N) Among them, EDC(i,j) is the Euclidean distance between client i and client j based on compressed cosine similarity, N is the number of client groups, V is the gradient matrix compressed by the machine learning algorithm SVD, v is the column vector of the gradient matrix V, S(i,v T ) is the model update vector of client i, S(j,v T ) is the model update vector of client j, ΔW T is the transpose matrix of the matrix composed of client model update vectors; Performing a clustering algorithm based on all the calculated EDC distances to obtain N client groups and the group optimization objective gradients corresponding to each client group.
2. The grouped training method based on distributed machine learning according to claim 1, characterized in that, Group the clients according to the local optimized objective gradients of each client to obtain N client groups and the group optimized objective gradients corresponding to each client group, and create a corresponding group process and a corresponding group consensus model for each client group, so that each client is assigned to the group process with the most similar optimized objective gradients for training the group consensus model. Specifically, it includes: Select M first clients from all clients and send the initial model to all the first clients, so that each first client obtains the corresponding local optimized objective gradient based on the local training data and the initial model, where M is less than N; Receive the local optimized objective gradients of each first client, and cluster the local optimized objective gradients of all first clients to obtain multiple client groups and the group optimized objective gradients corresponding to each client group; Send the initial model and each group optimized objective gradient to the second clients, so that the second clients obtain the local optimized objective gradients based on the local training data and the initial model; where the second clients refer to the clients other than the first clients among all clients; Receive the grouping selection results returned by the second clients, update the client groups, and assign each second client to the corresponding group process for training the group consensus model; where the grouping selection result of the second client is the group process with the highest similarity of the optimized gradients selected by the second client based on the comparison of the similarity between the local optimized objective gradient and each group optimized objective gradient.
3. The grouped training method based on distributed machine learning according to claim 1, characterized in that, The similarity between the local optimized objective gradient and each group optimized objective gradient is calculated by the following formula: Among them, represents the group optimization objective gradient of group process g j ; represents the local optimization objective gradient of client i, and sim(i,j) represents the similarity between the local optimization objective gradient of client i and the group optimization objective gradient of group process g j ; represents the angle between and 4. A grouped training method based on distributed machine learning, which is applied to clients, characterized in that, It includes: Obtain the local optimized objective gradient based on the local training data and the initial model sent by the server, and send the local optimized objective gradient to the server; So that the server clusters the local optimized objective gradients corresponding to all clients to obtain N client groups and the group optimized objective gradients corresponding to each client group, and based on each client group, create the initialization of the group process corresponding to each client group and the group consensus model corresponding to each client group; where the client group is obtained based on the grouping result of the clients; Receive the grouping result and the corresponding group consensus model returned by the server, and train the group consensus model in the corresponding group process; After the end of the previous training round and before the start of the current training round, calculate the local training data distribution offset result; When the training data distribution offset result is an offset, calculate the similarity between the current local optimized objective gradient and each group optimized objective gradient, and send the grouping selection information to the group process with the highest similarity of the optimized objective gradients, so that the server updates the grouping of the clients and assigns the clients to the corresponding group process in the current training round for training the group consensus model; The determination of the training data distribution shift result is carried out in the following manner: When the sum of the deviation distances of all label categories in the current training round is greater than the preset shift distance, it is determined that the training data distribution shift result is a shift; wherein, the label category refers to the category of the training sample, and the deviation distance of each label category is the deviation distance between the label category probability distribution in the current training round and the label category probability distribution at the time of the last shift; The deviation distance between the label category probability distribution in the current training round and the label category probability distribution at the time of the last shift is calculated by the following formula: where d c represents the offset distance, EMD(·) represents the EMD distance between two distributions, and P c is the probability distribution of the label categories of the client in the current training round, and is the probability distribution of the label categories at the time of the last offset on the client; The server clusters the local optimized objective gradients corresponding to all clients to obtain N client groups and the group optimized objective gradients corresponding to each client group, specifically including: The server calculates the Euclidean distance based on the compressed cosine similarity between each client and other clients according to the following formula: V = SVD(ΔW T , N) Among them, EDC(i,j) is the Euclidean distance between client i and client j based on compressed cosine similarity, N is the number of client groups, V is the gradient matrix compressed by the machine learning algorithm SVD, v is the column vector of the gradient matrix V, S(i,v T ) is the model update vector of client i, S(j,v T ) is the model update vector of client j, ΔW T is the transpose matrix of the matrix composed of client model update vectors; The server performs a clustering algorithm based on all the calculated EDC distances to obtain N client groups and the group optimized objective gradients corresponding to each client group.
5. A server characterized in that it includes: A group process creation module, configured to send an initial model to all clients, so that all the clients calculate the corresponding local optimized objective gradients based on the local training data and the initial model; receive the local optimized objective gradients uploaded by all clients, and cluster the optimized objective gradients of all clients to obtain N client groups and the group optimized objective gradients corresponding to each client group; based on each client group, create a group process and a group consensus model corresponding to each client group, so that each client is assigned to the group process with the most similar optimized objective gradient for training the group consensus model; N is greater than or equal to 1, wherein the local optimized objective gradient is calculated by the client based on the local training data and the initial model sent by the server; A grouped update information receiving module, configured to receive the grouped update information returned by the client with a shifted training data distribution after the end of the previous training round and before the start of the current training round; wherein, the grouped update selection information of the client is the group process with the highest similarity of the optimized objective gradient selected by the client with a shifted training data distribution based on the comparison result of the similarity between the current local optimized objective gradient and each group optimized objective gradient; A client grouping update module, configured to update the grouping of the clients according to the grouped update information, so that each client is assigned to the corresponding group process in the current training round for training the group consensus model; The determination of the client with a shifted training data distribution is carried out in the following manner: When the sum of the deviation distances of all label categories of the client in the current training round is greater than the preset shift distance, it is determined that the training data distribution of the client has shifted; wherein, the label category refers to the category of the training sample, and the deviation distance of each label category is the deviation distance between the label category probability distribution of the client in the current training round and the label category probability distribution at the time of the last shift; The offset distance between the label category probability distribution of the client in the current training round and the label category probability distribution at the time of the last offset is calculated by the following formula: where d c represents the offset distance, EMD(·) represents the EMD distance between two distributions, and P c is the probability distribution of label categories of the client at the current training round, and is the probability distribution of label categories when the last offset occurred on the client; Receiving the local optimized objective gradients uploaded by all clients, and clustering the optimized objective gradients of all clients to obtain N client groups and the group optimized objective gradient corresponding to each client group, specifically including: Receiving the optimized objective gradient uploaded by each client, and calculating the Euclidean distance based on the compressed cosine similarity between each client and other clients according to the following formula: V = SVD(ΔW T , N) Among them, EDC(i,j) is the Euclidean distance between client i and client j based on compressed cosine similarity, N is the number of client groups, V is the gradient matrix compressed by the machine learning algorithm SVD, v is the column vector of the gradient matrix V, S(i,v T ) is the model update vector of client i, S(j,v T ) is the model update vector of client j, ΔW T is the transpose matrix of the matrix composed of client model update vectors; Performing a clustering algorithm based on all the calculated EDC distances to obtain N client groups and the group optimized objective gradient corresponding to each client group.
6. A client Characterized in that It includes: A local optimized objective gradient acquisition module, configured to obtain a local optimized objective gradient based on local training data and an initial model sent by the server, and send the local optimized objective gradient to the server; so that the server clusters the local optimized objective gradients corresponding to all clients to obtain N client groups and the group optimized objective gradient corresponding to each client group, and based on each client group, create an initialization of a group process corresponding to each client group and a group consensus model corresponding to each client group; wherein, the client group is obtained based on the grouping result of the clients; A group consensus model training module, configured to receive the grouping result and the corresponding group consensus model returned by the server, and perform training of the group consensus model in the corresponding group process; A data distribution offset result calculation module, configured to calculate the local training data distribution offset result after the end of the previous training round and before the start of the current training round; A grouping selection module, configured to, when the training data distribution offset result is an offset, calculate the similarity between the current local optimized objective gradient and each group optimized objective gradient, and send grouping selection information to the group process with the highest similarity of the optimized objective gradient, so that the server updates the grouping of the clients and assigns the clients to the corresponding group process in the current training round for training of the group consensus model; The training data distribution offset result is judged by the following method: When the sum of the deviation distances of all label categories in the current training round is greater than a preset offset distance, it is determined that the training data distribution offset result is an offset; wherein, the label category refers to the category of the training sample, and the deviation distance of each label category is the offset distance between the label category probability distribution in the current training round and the label category probability distribution at the time of the last offset; The offset distance between the label category probability distribution in the current training round and the label category probability distribution at the time of the last offset is calculated by the following formula: where d c represents the offset distance, EMD(·) represents the EMD distance between two distributions, and P c is the probability distribution of label classes of the client at the current training round, and is the probability distribution of label classes when the last offset occurred on the client; The server clusters the local optimized objective gradients corresponding to all clients to obtain N client groups and the group optimized objective gradient corresponding to each client group, specifically including: The server calculates the Euclidean distance based on the compressed cosine similarity between each client and other clients according to the following formula: V = SVD(ΔW T , N) Among them, EDC(i,j) is the Euclidean distance between client i and client j based on compressed cosine similarity, N is the number of client groups, V is the gradient matrix compressed by the machine learning algorithm SVD, v is the column vector of the gradient matrix V, S(i,v T ) is the model update vector of client i, S(j,v T ) is the model update vector of client j, ΔW T is the transpose matrix of the matrix composed of the client model update vectors; The server performs a clustering algorithm based on all the calculated EDC distances to obtain N client groups and the group optimization objective gradients corresponding to each client group.
Citation Information
Patent Citations
Lightweight clustering federal learning method and device and storage medium
CN118627597A
Failure Prediction In Distributed Environments
US20230023646A1