Federal learning model training method and device, equipment and storage medium
By performing two-stage federated learning and clustering processing on the server side, and using the attention mechanism to personalize the model parameters, the model performance degradation caused by inconsistent data distribution in different data centers is solved, and fast and accurate federated learning model training is achieved.
Patent Information
- Application Number
- CN202510205623.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-17
AI Technical Summary
When the data in different data centers are non-independent and homogeneously distributed or the data is large, conventional federated learning training methods may lead to degradation of model performance, and there is a lack of a technical solution that can quickly and accurately train federated learning models.
By performing the first stage of federated learning with each client on the server side, obtaining the global model parameter vector, clustering according to the data distribution characteristics of the client to form client clusters of various categories, and in the second stage, federated learning is carried out based on the attention mechanism, and model parameters are personalized to obtain a federated learning model suitable for each client cluster.
During the model aggregation process, global client data can be used to improve model accuracy and robustness, and the model can be adjusted based on personalized data of each client cluster, retain data specificity, improve model performance, and achieve fast and accurate model training.
Smart Images

Figure CN120163263A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of model training, and particularly to a method, apparatus, device, and storage medium for training a federated learning model. Background Art
[0002] For a single data center (also referred to as a client) with a large amount of data, a centralized training method can be used to obtain the model it needs. However, in the case where the data volume of a single data center is insufficient and it is desired to utilize the data of other data centers, considering data security issues, it is an excellent method to obtain the required model by using the federated learning training method.
[0003] However, when the data in different data centers is non-independent and identically distributed or the data differences are large, etc., the conventional federated learning training method may cause the performance of the model to decline. Therefore, it is urgent to explore a technical solution that can quickly and accurately train the federated learning model. Summary of the Invention
[0004] This application provides a method, apparatus, device, and storage medium for training a federated learning model to quickly and accurately train the federated learning model.
[0005] In a first aspect, this application provides a method for training a federated learning model. The method is applied to a server and includes:
[0006] Conduct federated learning with each client to obtain a globally trained model parameter vector;
[0007] For each category of client clusters obtained by clustering according to the data distribution characteristics of the clients, based on the globally trained model parameter vector, conduct federated learning with the client clusters of this category to obtain a federated learning model suitable for the client clusters of this category.
[0008] In a possible implementation manner, the conducting federated learning with the client clusters of this category includes:
[0009] In each round of iteration of conducting federated learning with the client clusters of this category, at least perform the following steps:
[0010] Receive the local model parameter vectors output in the previous round sent by each client in the client clusters of this category; and based on the attention mechanism, determine the weight of each client;
[0011] Based on the weight of each client and the local model parameter vector of each client, determine the global model parameter vector of this round to be sent to each client in this round, and send the global model parameter vector of this round to each client, so that each client uses the global model parameter vector of this round to update the local model parameter vector currently saved by the client itself, and enable each client to perform iterative training of the federated learning model to be trained based on the updated local model parameter vector.
[0012] In a possible implementation manner, determining the weight of each client based on the attention mechanism includes:
[0013] For each client, determine the distance between the local model parameter vector of this client and the global model parameter vector sent to each client in the previous round, and determine the weight of this client according to the distance.
[0014] In a possible implementation manner, the weight is proportional to the distance.
[0015] In a possible implementation manner, the process of obtaining the client cluster of each category includes:
[0016] Based on federated learning and clustering algorithms, obtain the client cluster of each category.
[0017] In a second aspect, the present application provides a federated learning model training device, which is applied to a server, and the device includes:
[0018] A first training module, configured to perform federated learning with each client to obtain a trained global model parameter vector;
[0019] A second training module, configured to perform federated learning with the client cluster of each category obtained by clustering according to the data distribution characteristics of the clients based on the global model parameter vector to obtain a federated learning model suitable for the client cluster of this category.
[0020] In a possible implementation manner, the second training module is specifically configured to:
[0021] In each iteration process of performing federated learning with the client cluster of this category, at least perform the following steps:
[0022] Receive the local model parameter vector output in the previous round sent by each client in the client cluster of this category; and determine the weight of each client based on the attention mechanism;
[0023] Based on the weights of each client and the local model parameter vectors of each client, determine the global model parameter vector for this round to be sent to each client, and send the global model parameter vector for this round to each client, so that each client updates the local model parameter vector currently saved by the client itself using the global model parameter vector for this round, and enable each client to perform iterative training for this round on the federated learning model to be trained based on the updated local model parameter vector.
[0024] In a possible implementation manner, the second training module is specifically configured to:
[0025] For each client, determine the distance between the local model parameter vector of the client and the global model parameter vector sent to each client in the previous round, and determine the weight of the client according to the distance.
[0026] In a possible implementation manner, the weight is proportional to the distance.
[0027] In a possible implementation manner, the second training module is specifically configured to:
[0028] Based on federated learning and a clustering algorithm, obtain the client clusters of each category.
[0029] In a third aspect, the present application provides an electronic device, which at least includes a processor and a memory. When the processor executes the computer program stored in the memory, it implements the steps of the method described in any one of the above.
[0030] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the method described in any one of the above.
[0031] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes: computer program code. When the computer program code runs on a computer, it causes the computer to execute the steps of the method described in any one of the above.
[0032] In the embodiments of the present application, the server can first perform the first stage of federated learning with each client to obtain the globally trained model parameter vector. Subsequently, for each client cluster obtained by clustering according to the data distribution characteristics of the clients, based on this globally trained model parameter vector, the server can perform the second stage of federated learning with the client cluster of this category to perform personalized adjustment of the globally trained model parameter vector for the client cluster of this category, thereby obtaining a federated learning model suitable for the client cluster of this category. Based on this, during the model aggregation (training) process, not only can all the data of all the clients globally be utilized, and the sufficient amount of data can improve the accuracy, generalization ability, and robustness of the model, but also the model can be personalized adjusted based on the personalized data of each client cluster, retaining the specificity of the data of each client cluster. Each client cluster can obtain a federated learning model exclusive to itself, improving the model performance and achieving the purpose of training the federated learning model quickly and accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] To more clearly illustrate the embodiments of the present application or the implementation manners in related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0034] Figure 1 FIG. 1 shows a schematic diagram of the training process of the first federated learning model provided by some embodiments;
[0035] Figure 2 FIG. 2 shows a schematic diagram of the training process of the second federated learning model provided by some embodiments;
[0036] Figure 3 FIG. 3 shows a schematic diagram of the training process of the third federated learning model provided by some embodiments;
[0037] Figure 4A FIG. 4 shows a schematic diagram of the performance comparison of the models obtained by training with a different model training method provided by some embodiments;
[0038] Figure 4B FIG. 5 shows a schematic diagram of the performance comparison of the models obtained by training with another different model training method provided by some embodiments;
[0039] Figure 5 FIG. 6 shows a schematic diagram of a federated learning model training device provided by some embodiments;
[0040] Figure 6 FIG. 7 shows a schematic diagram of the structure of an electronic device provided by some embodiments. Detailed implementation manners
[0041] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. Obviously, the embodiments described in the present application are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application.
[0042] It should be noted that the brief description of the terms in the present application is only for facilitating the understanding of the following described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and general meanings.
[0043] The terms "first", "second", "third", etc. in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.
[0044] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components does not necessarily have to be limited to all the components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0045] The term "module" refers to any known or later-developed hardware, software, firmware, artificial intelligence, fuzzy logic or a combination of hardware and / or software code that can perform functions related to the element.
[0046] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0047] In order to train the federated learning model quickly and accurately, the present application provides a method, device, equipment and storage medium for training the federated learning model. In this method, the server can first perform the first stage of federated learning with each client to obtain the globally trained model parameter vector; then, for each cluster of clients obtained by clustering according to the data distribution characteristics of the clients, based on this globally trained model parameter vector, perform the second stage of federated learning with the cluster of clients of this category to perform personalized adjustment of this globally trained model parameter vector for the cluster of clients of this category, so as to obtain a federated learning model suitable for the cluster of clients of this category. Based on this, in the process of model aggregation (training), not only can all the data of all the clients globally be utilized, and the sufficient amount of data can improve the model accuracy, generalization ability and robustness, but also the model can be personalized adjusted based on the personalized data of each cluster of clients, retaining the specificity of the data of each cluster of clients. Each cluster of clients of each category can obtain a federated learning model exclusive to itself, improving the model performance and achieving the purpose of quickly and accurately training the federated learning model.
[0048] All implementation manners of the embodiments of the present application comply with the relevant regulations of national laws and regulations regarding the acquisition, storage, use, processing, etc. of data.
[0049] Embodiment 1:
[0050] Figure 1 The first schematic diagram of the federated learning model training process provided by some embodiments is shown. This method is applied to the server. As Figure 1 shown, this process includes the following steps:
[0051] S101: Perform federated learning with each client to obtain the globally trained model parameter vector.
[0052] In a possible implementation manner, the server can perform the first stage of federated learning with all the clients globally to obtain a pre-trained global model, and the model parameter vector of this global model (for the convenience of description, called the global model parameter vector) can be used as the starting point for the second stage of federated learning model training to perform the second stage of federated learning in step S102.
[0053] Among them, the process of the server performing the first stage of federated learning with each client can be:
[0054] The server sends the initialized model parameter vector to each client. After that, in each round of iteration of the federated learning between the server and each client, it can at least include the following steps:
[0055] When each client receives the initialized model parameter vector sent by the server, it can perform iterative training for this round based on the initialized model parameter vector and local data. After obtaining the local model parameter vector output for this round, it can send the local model parameter vector to the server.
[0056] After the server receives the local model parameter vectors output by each client for this round (for the convenience of description, called the previous round), it can perform weighted averaging on the local model parameter vectors output by each client in the previous round to obtain the global model parameter vector for this round to be sent to each client, and send the global model parameter vector for this round to each client.
[0057] After each client receives the global model parameter vector for this round, it can use the global model parameter vector for this round to update the local model parameter vector currently saved by the client itself, and perform iterative training based on the updated local model parameter vector and local data to obtain the local model parameter vector output for this round, and can send the local model parameter vector to the server again.
[0058] After several rounds of such iterative training, the model parameter vector of the global model obtained from the first-stage training can be obtained.
[0059] Among them, this application does not specifically limit the termination training conditions for the first-stage federated learning, and can be flexibly set according to requirements. For example, it can be that the total number of training rounds reaches the set round threshold, or the model has converged, etc.
[0060] S102: For each cluster of clients obtained by clustering according to the data distribution characteristics of the clients, perform federated learning between the global model parameter vector and the cluster of clients of this category to obtain a federated learning model suitable for the cluster of clients of this category.
[0061] In a possible implementation manner, in order to improve the model accuracy, the clients can be clustered according to the data distribution characteristics of the clients to obtain a cluster of clients for each category. Then, taking the global model obtained from the first-stage federated learning training between the server and each client as a starting point, for each cluster of clients, perform the second-stage federated learning to perform personalized adjustment on the global model parameter vector of the global model to obtain a federated learning model suitable for each cluster of clients.
[0062] First, the process of clustering each client to obtain a cluster of clients for each category (this process can also be called the knowledge-driven federated clustering process, Knowledge-based Federated Clustering) will be explained below.
[0063] In a possible implementation, the server can cluster each client based on federated learning and a clustering algorithm (such as k-means) according to the data distribution characteristics of each client, so as to obtain a client cluster for each category after clustering. Taking the natural gas usage data of consumers as the data of each client as an example, each client can be clustered according to data distribution characteristics such as the gas consumption, gas usage type, and geographical location area of the consumers. Clients with the same or similar data distribution characteristics are clustered into the same category of client clusters. On the contrary, different clients with significantly different data distribution characteristics are classified into different categories of client clusters, and a client cluster for each category after clustering is obtained.
[0064] Specifically, before the server clusters each client based on federated learning and a clustering algorithm (such as k-means), for each client, the client can preprocess its local data (which can also be called time series data or temporal data) by using a time series decomposition algorithm (such as Seasonal-Trend Decomposition Procedure based on Loess, STL), and decompose the time series data into three parts: trend, seasonality, and residual. Among them, the process of preprocessing time series data based on STL can adopt existing technologies and will not be elaborated here.
[0065] After each client (which can also be referred to as the device owned by the retailer) preprocesses the local data respectively, a process of clustering each client based on federated learning and clustering algorithms can be carried out between the server and each client. Specifically, the server can initialize a clustering center (for the convenience of description, called the global clustering center), and this clustering center can contain parameters (which can also be called parameter feature vectors) that can reflect the data distribution characteristics of the clients, such as gas consumption, geographical location area, etc. The global clustering center is sent to each client respectively. Each client can update the global clustering center based on the distance between the data distribution characteristics (such as data feature vectors) of the local data and the parameter feature vectors in the global clustering center and the clustering algorithm, and obtain the local clustering center after this round of update. Among them, the client can use existing clustering algorithms to update the global clustering center, which will not be elaborated here. After that, each client can send the updated local clustering center to the server. The server can perform weighted averaging (aggregation) on the parameter feature vectors in the local clustering centers sent by each client in this round to obtain the global clustering center of the next round, and send this global clustering center to each client. Each client can update this global clustering center again to obtain the local clustering center after the next round of update. After that, each client can send the updated local clustering center of the next round to the server. The server can aggregate the parameter feature vectors in the local clustering centers of the next round sent by each client in this round to obtain the global clustering center of the round after next. Through several rounds of such iterative processes, until the update amplitude of the global clustering center sent by the client to the server is less than the set threshold, the federated learning iteration process is stopped. At this point, each client can be clustered into several categories of client clusters based on the trained global clustering center.
[0066] In a possible implementation manner, after clustering each client into multiple categories of client clusters, the second stage of federated learning can be carried out based on the global model parameter vector obtained through the first stage of federated learning between the server and each client in step S101. When carrying out the second stage of federated learning, for each category of client cluster, the server can carry out federated learning with this category of client cluster based on this global model parameter vector to obtain a federated learning model suitable for this category of client cluster.
[0067] Among them, for each category of client cluster, the server can carry out the second stage of federated learning with this category of client cluster in two ways.
[0068] Method 1:
[0069] For each category of client clusters, the server can first send the global model parameter vector of the global model obtained from the first-stage federated learning to the client cluster of this category. After that, in each round of the iterative process of federated learning between the server and the client cluster of this category, it can at least include the following steps:
[0070] When each client in the client cluster of this category receives the global model parameter vector sent by the server, it can use this global model parameter vector to update the local model parameter vector, and perform iterative training on the federated learning model to be trained based on the updated local model parameter vector and local data. After obtaining the local model parameter vector output in this round, it can send this local model parameter vector to the server. The server can receive the local model parameter vector output in this round (for the convenience of description, called the previous round) sent by each client in the client cluster of this category.
[0071] The server can perform weighted averaging on the local model parameter vectors output in the previous round by each client in the client cluster of this category to obtain the global model parameter vector of this round to be sent to each client in the client cluster of this category, and send this global model parameter vector of this round to each client in the client cluster of this category.
[0072] Each client in the client cluster of this category can use the global model parameter vector of this round to update the local model parameter vector currently saved by the client itself, and perform iterative training on the federated learning model to be trained based on the updated local model parameter vector and local data, obtain the local model parameter vector output in this round, and can send this local model parameter vector to the server again.
[0073] After several rounds of such iterative training, a federated learning model suitable for each category of client clusters can be obtained.
[0074] For the convenience of description, the federated learning model obtained based on Method 1 will be called the federated learning model obtained based on personalized federated learning and federated averaging (CPFL-FedAvg) later.
[0075] Method 2:
[0076] For each category of client clusters, the server can first send the global model parameter vector of the global model obtained from the first-stage federated learning to the client cluster of this category. After that, in each round of the iterative process of federated learning between the server and the client cluster of this category, it can at least include the following steps:
[0077] When each client in the client cluster of this category receives the global model parameter vector sent by the server, it can use this global model parameter vector to update the local model parameter vector. Based on the updated local model parameter vector and local data, it can perform iterative training on the federated learning model to be trained. After obtaining the local model parameter vector output in this round, it can send this local model parameter vector to the server. The server can receive the local model parameter vectors output in this round (for the convenience of description, called the previous round) sent by each client in the client cluster of this category.
[0078] The server can determine the weight of each client based on the attention mechanism. Exemplarily, for each client, the server can determine the distance (such as the Euclidean distance) between the local model parameter vector of this client and the global model parameter vector sent to the client cluster of this category in the previous round, and can determine the weight of this client according to this distance. In a possible implementation manner, it can be that the greater the distance, the greater the weight of this client, and the smaller the distance, the smaller the weight of this client. That is to say, the weight is proportional to the distance. Exemplarily, the Euclidean distance between the local model parameter vector of this client and the global model parameter vector can be directly determined as the weight of this client, or the product or sum value of this Euclidean distance and a set value can be determined as the weight of this client. This application does not make specific limitations on this.
[0079] After determining the weights of each client in the client cluster of a certain category, the weighted sum value of the weights of each client in the client cluster of this category and the local model parameter vectors of each client in the client cluster of this category can be determined as the global model parameter vector of this round sent to each client in the client cluster of this category.
[0080] Each client in the client cluster of this category can use the global model parameter vector of this round to update the local model parameter vector currently saved by the client itself, and perform iterative training on the federated learning model to be trained based on the updated local model parameter vector and local data, obtain the local model parameter vector output in this round, and can send this local model parameter vector to the server again.
[0081] After several rounds of such iterative training, a federated learning model suitable for the client cluster of each category can be obtained. Among them, this application does not make specific limitations on the termination training conditions of the second-stage federated learning, and can be flexibly set according to needs. For example, it can be that the total number of training rounds reaches a set round threshold, or the model has converged, etc.
[0082] For ease of description, the federated learning model obtained based on Method 2 will be hereinafter referred to as the federated learning model obtained based on personalized federated learning and attention mechanism federated learning (CPFL-FedAtt).
[0083] For ease of understanding, the model training process provided by the present application will be explained below through a specific embodiment. Refer to Figure 2 , Figure 2 which shows a schematic diagram of the training process of the second federated learning model provided by some embodiments. The process includes the following steps:
[0084] S201: Conduct the first stage of federated learning between the server and each client to obtain the globally trained model parameter vector. In addition, based on federated learning and clustering algorithms, the server clusters each client to obtain a client cluster for each category.
[0085] S202: For each client cluster of each category obtained by clustering according to the data distribution characteristics of the clients, the server conducts the second stage of federated learning with the client cluster of this category based on the globally trained model parameter vector obtained in the first stage. Among them, in each round of iteration of conducting federated learning with the client cluster of this category, at least the following steps are performed:
[0086] Receive the local model parameter vector output in the previous round sent by each client in the client cluster of this category; and for each client, determine the distance between the local model parameter vector of this client and the globally trained model parameter vector sent to the client cluster of this category in the previous round. According to the distance, determine the weight of this client. According to the weights of each client in the client cluster of this category and the local model parameter vector of each client, determine the globally trained model parameter vector of each client sent to the client cluster of this category in this round, and send the globally trained model parameter vector in this round to each client in the client cluster of this category. Each client in the client cluster of this category updates its currently saved local model parameter vector with the globally trained model parameter vector in this round, and conducts this round of iterative training on the federated learning model to be trained based on the updated local model parameter vector.
[0087] S203: For each client cluster of each category, obtain a federated learning model suitable for the client cluster of this category.
[0088] For ease of understanding, the model training process provided by the present application will be further explained below through a specific embodiment. Refer to Figure 3 , Figure 3 which shows a schematic diagram of the training process of the third federated learning model provided by some embodiments. The process includes the following steps:
[0089] Step A: Based on federated learning and clustering algorithms, the server clusters the clients of each retailer among all the clients of all retailers.
[0090] Among them, each retailer (such as retailer A, retailer B, and retailer C shown in the figure) may contain several clients, and this application does not make specific limitations in this regard. For each retailer, after clustering the clients of the retailer according to the data distribution characteristics of the clients of the retailer, client clusters of each category of the retailer are obtained. For example, it is assumed that each retailer contains client clusters of two categories, namely cluster1 and cluster2.
[0091] Step B: The server and all the clients of all retailers perform the first stage of federated learning to obtain the global model parameter vector obtained from the first stage of training.
[0092] Step C: For the client clusters of the cluster1 category in each retailer, respectively based on the global model parameter vector, perform the second stage of federated learning with the client clusters of cluster1 to obtain a federated learning model suitable for the client clusters of cluster1. For the client clusters of the cluster2 category in each retailer, respectively based on the global model parameter vector, perform the second stage of federated learning with the client clusters of cluster2 to obtain a federated learning model suitable for the client clusters of cluster2.
[0093] Among them, this application does not make specific limitations on the sequence between Step A and Step B. It is possible to perform Step A first, or Step B first, or perform Step A and Step B in parallel at the same time, and it can be flexibly set according to requirements.
[0094] In a possible implementation manner, different model training methods may have different impacts on the performance of the trained model. Exemplarily, please refer to Figure 4A , Figure 4AThe figure shows a schematic diagram of the performance comparison of models obtained by training with a different model training method provided by some embodiments. Using the monthly consumption data of 2,000 natural gas consumers, the accuracy of the predicted values of natural gas demand in the next 1-3 months of the models obtained by separately training (Solo), centralized training (Centralized), federated average (FedAvg), federated learning with attention mechanism (FedAtt), group personalized federated learning (CFL), personalized federated learning with separate fine-tuning (PFL-Solo), as well as the personalized federated learning and federated average (CPFL-FedAvg) in Method 1 and the personalized federated learning and attention mechanism federated learning (CPFL-FedAtt) in Method 2 in the embodiments of the present application is compared. Among them, during the training process, the data set is divided into a training set and a test set, and the mean square error (MSE), mean absolute error (MAE), and mean absolute percentage error (MAPE) are used to evaluate the model performance (accuracy).
[0095] It can be seen that the errors of the predicted results of the models obtained by training with Method 1 and Method 2 of the present application are both smaller than those of the models obtained by other training methods in the related art. In particular, the error of the predicted result of the model obtained by training with Method 2 is the smallest, and the model performance is the highest. For example, in the predicted result of the first month, the mean square error (MSE) of the predicted result of the model obtained by training with Method 2 (CPFL-FedAtt) is only 0.0126, the mean absolute error (MAE) is only 0.0834, and the mean absolute percentage error (MAPE) is only 0.1930, all of which are lower than those of the models obtained by other training methods, and will not be elaborated here.
[0096] In addition, referring to Figure 4B , Figure 4B The figure shows a comparison chart of the MAPE of the predicted results of the models obtained by separately training (Solo) and training with Method 2 (CPFL-FedAtt) provided by some embodiments. From Figure 4BIt can be seen that among the prediction results of the model trained by the CPFL-FedAtt method, the MAPE of the prediction data of 227 consumers does not exceed 0.1, and the MAPE of the prediction data of 963 consumers is between 0.1 and 0.2. Among the prediction results of the model trained by the solo training method, the MAPE of the prediction data of 190 consumers does not exceed 0.1, and the MAPE of the prediction data of 840 consumers is between 0.1 and 0.2. That is, in the range of relatively small MAPE values (not exceeding 0.2), the number of consumers of CPFL-FedAtt is more than that of Solo. On the contrary, in the range of relatively large MAPE values (0.2-0.4), the number of consumers of CPFL-FedAtt is less than that of Solo. It can be seen that the accuracy of the model trained based on CPFL-FedAtt is higher than that of the model trained based on Solo.
[0097] In the model training method provided by the embodiments of the present application, since the model is trained through federated learning between the server and the client, the data of the client can remain on the local device of the client, effectively protecting the privacy of the data of each client and avoiding data leakage. The model trained based on the model training method provided by the embodiments of the present application can effectively balance the privacy protection and data heterogeneity problems in natural gas load forecasting, improving the accuracy and practicality of model prediction.
[0098] Embodiment 2:
[0099] Based on the same technical concept, the present application provides a federated learning model training device, which is applied to a server. Refer to Figure 5 , Figure 5 which shows a schematic diagram of a federated learning model training device provided by some embodiments. The device includes:
[0100] A first training module 501, configured to perform federated learning with each client to obtain a globally trained model parameter vector;
[0101] A second training module 502, configured to perform federated learning with each client cluster of each category obtained by clustering according to the data distribution characteristics of the client, based on the globally trained model parameter vector, to obtain a federated learning model suitable for the client cluster of this category.
[0102] In a possible implementation manner, the second training module 502 is specifically configured to:
[0103] In each round of iteration of performing federated learning with the client cluster of this category, at least perform the following steps:
[0104] Receive the local model parameter vectors of the previous round of output sent by each client in the client cluster of this category; and determine the weight of each client based on the attention mechanism;
[0105] According to the weights of each client and the local model parameter vectors of each client, determine the global model parameter vectors of this round to be sent to each client in this round, send the global model parameter vectors of this round to each client, so that each client uses the global model parameter vectors of this round to update the local model parameter vectors currently saved by the client itself, and enable each client to perform iterative training of this round on the federated learning model to be trained based on the updated local model parameter vectors.
[0106] In a possible implementation manner, the second training module 502 is specifically configured to:
[0107] For each client, determine the distance between the local model parameter vector of this client and the global model parameter vector sent to each client in the previous round, and determine the weight of this client according to the distance.
[0108] In a possible implementation manner, the weight is proportional to the distance.
[0109] In a possible implementation manner, the second training module 502 is specifically configured to:
[0110] Based on federated learning and clustering algorithms, obtain the client clusters of each category.
[0111] Embodiment 3:
[0112] Based on the same technical concept, the present application also provides an electronic device, Figure 6 showing a schematic structural diagram of an electronic device provided by some embodiments, as Figure 6 shown, the electronic device includes: a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 complete communication with each other through the communication bus 604;
[0113] A computer program is stored in the memory 603, and when the program is executed by the processor 601, the processor 601 is caused to execute the following steps:
[0114] Perform federated learning with each client to obtain the trained global model parameter vectors;
[0115] For each client cluster of each category obtained after clustering according to the data distribution characteristics of the clients, based on the global model parameter vector, perform federated learning with the client cluster of this category to obtain a federated learning model suitable for the client cluster of this category.
[0116] In a possible implementation manner, the processor 601 is specifically configured to:
[0117] In each round of iteration of performing federated learning with the client cluster of this category, at least perform the following steps:
[0118] Receive the local model parameter vectors output in the previous round sent by each client in the client cluster of this category; and based on the attention mechanism, determine the weight of each client;
[0119] According to the weight of each client and the local model parameter vector of each client, determine the global model parameter vector of this round sent to each client, send the global model parameter vector of this round to each client, so that each client uses the global model parameter vector of this round to update the local model parameter vector currently saved by the client itself, and enable each client to perform iterative training of the federated learning model to be trained based on the updated local model parameter vector.
[0120] In a possible implementation manner, the processor 601 is specifically configured to:
[0121] For each client, determine the distance between the local model parameter vector of this client and the global model parameter vector sent to each client in the previous round, and determine the weight of this client according to the distance.
[0122] In a possible implementation manner, the weight is proportional to the distance.
[0123] In a possible implementation manner, the processor 601 is specifically configured to:
[0124] Based on federated learning and clustering algorithms, obtain the client clusters of each category.
[0125] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0126] The communication interface 602 is used for communication between the above-mentioned electronic device and other devices.
[0127] The memory may include a random access memory (RAM), or may also include a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0128] The above-mentioned processor may be a general-purpose processor, including a central processing unit, a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0129] Embodiment 4:
[0130] Based on the same technical concept, an embodiment of the present application provides a computer-readable storage medium, in which a computer program executable by an electronic device is stored. When the program runs on the electronic device, the electronic device is caused to execute the following steps when executed:
[0131] Perform federated learning with each client to obtain a globally trained model parameter vector;
[0132] For each client cluster obtained by clustering according to the data distribution characteristics of the clients, based on the globally trained model parameter vector, perform federated learning with the client cluster of this category to obtain a federated learning model suitable for the client cluster of this category.
[0133] In a possible implementation manner, the performing federated learning with the client cluster of this category includes:
[0134] In each round of iteration of performing federated learning with the client cluster of this category, at least the following steps are performed:
[0135] Receive the local model parameter vectors output in the previous round sent by each client in the client cluster of this category; and based on the attention mechanism, determine the weight of each client;
[0136] Determine the global model parameter vector for this round to be sent to each client according to the weight of each client and the local model parameter vector of each client, and send the global model parameter vector for this round to each client, so that each client uses the global model parameter vector for this round to update the local model parameter vector currently saved by the client itself, and enable each client to perform iterative training for this round on the federated learning model to be trained based on the updated local model parameter vector.
[0137] In a possible implementation manner, determining the weight of each client based on the attention mechanism includes:
[0138] For each client, determine the distance between the local model parameter vector of this client and the global model parameter vector sent to each client in the previous round, and determine the weight of this client according to the distance.
[0139] In a possible implementation manner, the weight is proportional to the distance.
[0140] In a possible implementation manner, the process of obtaining the client clusters of each category includes:
[0141] Based on federated learning and clustering algorithms, obtain the client clusters of each category.
[0142] The above computer-readable storage medium can be any available medium or data storage device accessible by the processor in the electronic device, including but not limited to magnetic memories such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc., optical memories such as CDs, DVDs, BDs, HVDs, etc., and semiconductor memories such as ROMs, EPROMs, EEPROMs, non-volatile memories (NANDFLASH), solid-state drives (SSD), etc.
[0143] Based on the same technical concept, the present application provides a computer program product, which includes: computer program code, when the computer program code runs on a computer, it enables the computer to execute the method described in any method embodiment applied to the electronic device above.
[0144] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof, and can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part.
[0145] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0146] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0147] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0148] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0149] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. A method for training a federated learning model, characterized in that The method is applied to a server, and the method includes: Performing federated learning with each client to obtain a globally trained model parameter vector; For each cluster of clients obtained by clustering according to the data distribution characteristics of the clients, based on the globally trained model parameter vector, performing federated learning with the cluster of clients of this category to obtain a federated learning model suitable for the cluster of clients of this category.
2. The method according to claim 1, characterized in that The performing federated learning with the cluster of clients of this category includes: In each round of iteration of performing federated learning with the cluster of clients of this category, at least the following steps are executed: Receiving the local model parameter vectors output in the previous round sent by each client in the cluster of clients of this category; and determining the weight of each client based on the attention mechanism; Determining the globally trained model parameter vector of this round sent to each client according to the weight of each client and the local model parameter vector of each client, sending the globally trained model parameter vector of this round to each client, enabling each client to update the local model parameter vector currently saved by the client itself with the globally trained model parameter vector of this round, and enabling each client to perform iterative training of the federated learning model to be trained based on the updated local model parameter vector.
3. The method according to claim 2, characterized in that The determining the weight of each client based on the attention mechanism includes: For each client, determining the distance between the local model parameter vector of this client and the globally trained model parameter vector sent to each client in the previous round, and determining the weight of this client according to the distance.
4. The method according to claim 3, characterized in that The weight is proportional to the distance.
5. The method according to any one of claims 1-4, characterized in that The process of obtaining each cluster of clients of this category includes: Based on federated learning and a clustering algorithm, obtaining each cluster of clients of this category.
6. A device for training a federated learning model, characterized in that The apparatus is applied to a server, and the apparatus includes: A first training module, configured to perform federated learning with each client to obtain a globally trained model parameter vector; A second training module, configured to, for each cluster of clients obtained by clustering according to the data distribution characteristics of the clients, based on the globally trained model parameter vector, perform federated learning with the cluster of clients of this category to obtain a federated learning model suitable for the cluster of clients of this category.
7. The device according to claim 6, characterized in that The second training module is specifically configured to: In each round of iteration of performing federated learning with the cluster of clients of this category, at least the following steps are executed: Receiving the local model parameter vectors output in the previous round sent by each client in the cluster of clients of this category; and determining the weight of each client based on the attention mechanism; Determining the globally trained model parameter vector of this round sent to each client according to the weight of each client and the local model parameter vector of each client, sending the globally trained model parameter vector of this round to each client, enabling each client to update the local model parameter vector currently saved by the client itself with the globally trained model parameter vector of this round, and enabling each client to perform iterative training of the federated learning model to be trained based on the updated local model parameter vector.
8. The device according to claim 6 or 7, characterized in that The second training module is specifically configured to: Based on federated learning and clustering algorithms, obtain the client clusters of each category.
9. An electronic device, characterized in that The electronic device includes at least a processor and a memory. When the processor executes the computer program stored in the memory, the steps of the method according to any one of claims 1-5 are implemented.
10. A computer-readable storage medium, characterized in that It stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1-5 are implemented.