Federated learning communication method, system and device
By using the client to train the quantizer and dequantizer in federated learning, quantizing and uploading the model parameters, and the server to perform dequantization and aggregation, the communication bottleneck problem in federated learning is solved, the efficiency is improved, and the accuracy of the model parameters is guaranteed.
Patent Information
- Application Number
- CN202311094452.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-28
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2043-08-28
AI Technical Summary
In federated learning, the heterogeneity and resource limitations between devices lead to communication bottlenecks, which affect communication efficiency and cost.
By training the quantizer and dequantizer on the client, quantizing and uploading the model parameters, and performing dequantization and aggregation on the server, it is ensured that the model parameters of each client can be effectively aggregated.
It reduces the client's communication volume and communication costs, improves the communication efficiency of federated learning, and ensures the accuracy of model parameters and the normal aggregation on non-independent and identically distributed data sets.
Smart Images

Figure CN119544695B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of federated learning technology, and in particular to a federated learning communication method, system, and device. Background Art
[0002] In the era of artificial intelligence, AI algorithms such as machine learning and deep learning are gradually becoming part of our lives. The foundation of these algorithms is more and better data. Data silos restrict the optimization of AI algorithms. To address these limitations, federated learning has emerged. It improves AI algorithm performance by leveraging more data (in terms of quantity and dimensionality) without leaving the data domain.
[0003] In the practical application of federated learning, the heterogeneity of devices and resource constraints exacerbate the communication bottlenecks faced by federated learning. Reducing communication costs within limited resources and improving the efficiency of federated modeling for massive client-side models are currently pressing challenges. Summary of the Invention
[0004] The main purpose of the present invention is to provide a federated learning communication method, system and device, aiming to propose a communication solution between the server and the client in the horizontal federated learning process to improve the efficiency of federated learning communication.
[0005] To achieve the above-mentioned object, the present invention provides a federated learning communication method, which is applied to a server participating in horizontal federated learning. The method comprises the following steps:
[0006] In any round of federated aggregation in horizontal federated learning, receiving quantization model parameters and dequantizers sent by each first client participating in the current round of federated aggregation, wherein, for any target client among the first clients, the quantization model parameters sent by the target client are obtained by quantizing the original model parameters of the target client using the quantizer of the target client, and the quantizer and dequantizer of the target client are obtained by training the target client using the original model parameters of the target client;
[0007] Dequantizing the quantization model parameters corresponding to the target client using the dequantizer corresponding to the target client to obtain the restored model parameters;
[0008] After obtaining the restoration model parameters corresponding to each of the first clients, the restoration model parameters of each of the first clients are aggregated to obtain an aggregate model parameter, and the aggregate model parameter is fed back to each of the first clients.
[0009] Optionally, the quantizer of the target client is obtained by training the original model parameters of the target client using the LQ-Nets algorithm, the quantizer includes a quantization reference vector, the dequantizer is used to divide the object of the dequantization operation by the quantization reference vector to dequantize the object of the dequantization operation, and the step of using the dequantizer corresponding to the target client to dequantize the quantization model parameters corresponding to the target client to obtain the restored model parameters includes:
[0010] The quantization model parameters corresponding to the target client are divided by the quantization reference vector corresponding to the target client using the inverse quantizer corresponding to the target client to obtain the restored model parameters.
[0011] Optionally, the federated learning communication method further includes:
[0012] receiving the model loss sent by each of the first clients;
[0013] Calculate the reduction in the model loss sent by the target client during the current round of federated aggregation compared to the model loss sent by the target client during the previous round of federated aggregation;
[0014] According to the reduction amount corresponding to each first client, each client participating in the next round of federated aggregation is selected from each second client participating in horizontal federated learning; wherein, when the reduction amount corresponding to the target client is greater than the average value of the reduction amounts corresponding to each first client, the probability of the target client being selected to participate in the next round of federated aggregation is greater than the probability of the target client being selected to participate in the current round of federated aggregation; when the reduction amount corresponding to the target client is less than or equal to the average value, the probability of the target client being selected to participate in the next round of federated aggregation is less than the probability of the target client being selected to participate in the current round of federated aggregation.
[0015] Optionally, the step of selecting clients to participate in the next round of federated aggregation from the second clients participating in horizontal federated learning according to the reduction amount corresponding to each first client includes:
[0016] When the reduction amount corresponding to the target client is greater than the average of the reduction amounts corresponding to the first clients, the α parameter value of the beta distribution corresponding to the target client is increased to update the beta distribution corresponding to the target client, wherein before the first round of federated aggregation, the same beta distribution is initialized for each of the second clients, and the initialization values of the α parameter and the β parameter in the beta distribution are equal;
[0017] When the reduction amount corresponding to the target client is less than or equal to the average value, increasing the β parameter value of the beta distribution corresponding to the target client to update the beta distribution corresponding to the target client;
[0018] After updating the beta distribution of each first client, a random number is generated using the beta distribution of each second client, and each second client is sorted in descending order according to the random number, and the second clients ranked in the front by a preset number are selected as the clients participating in the next round of federal aggregation.
[0019] Optionally, the quantization model parameters sent by the target client are obtained by quantizing the original model parameters of the target client using a quantizer of the target client and then encoding the quantized model parameters. The step of dequantizing the quantization model parameters corresponding to the target client using a dequantizer corresponding to the target client to obtain the restored model parameters includes:
[0020] The encoded quantization model parameters corresponding to the target client are decoded to obtain a decoding result, and then the decoding result is dequantized using a dequantizer corresponding to the target client to obtain a restored model parameter.
[0021] Optionally, the step of decoding the encoded quantization model parameters corresponding to the target client to obtain a decoding result includes:
[0022] Taking each quantized value of the quantized model parameter before encoding corresponding to the target client as a signal source symbol, and obtaining the cumulative probabilities corresponding to the various signal source symbols when arranged in a preset order;
[0023] The encoded quantization model parameters corresponding to the target client are used as a target decoding object, and a target source symbol is determined from various source symbols, so that the target decoding object is within an interval formed by a cumulative probability of the target source symbol and a cumulative probability of a next source symbol, wherein the next source symbol is a next source symbol among the target source symbols when the various source symbols are arranged in the preset order;
[0024] Using the parameter value corresponding to the target signal source symbol as a parameter value obtained by decoding, subtracting the cumulative probability of the target signal source symbol from the target decoding object and dividing the result by the preset occurrence probability of the target signal source symbol, updating the calculation result as the target decoding object, and returning to the step of determining the target signal source symbol from the various signal source symbols;
[0025] After the last parameter value is obtained through decoding, the quantization model parameters before encoding corresponding to the target client are obtained according to the decoded parameter values, and the quantization model parameters before encoding corresponding to the target client are used as the decoding result.
[0026] To achieve the above-mentioned object, the present invention provides a federated learning communication method, which is applied to any target client among various first clients participating in horizontal federated learning. The federated learning communication method includes the following steps:
[0027] In any round of federated aggregation in horizontal federated learning, the original model parameters are used to train the quantizer and dequantizer;
[0028] quantizing the original model parameters using the quantizer to obtain quantized model parameters;
[0029] The quantization model parameters and the dequantizer are sent to a server participating in horizontal federated learning, so that the server uses the dequantizer corresponding to the target client to dequantize the quantization model parameters corresponding to the target client to obtain the restored model parameters, and after obtaining the restored model parameters corresponding to each of the first clients, the restored model parameters of each of the first clients are aggregated to obtain the aggregated model parameters.
[0030] Optionally, the step of sending the quantization model parameters and the dequantizer to a server participating in horizontal federated learning includes:
[0031] Taking each quantized value in the quantized model parameter as a signal source symbol, and calculating the cumulative probabilities corresponding to the various signal source symbols when arranged in a preset order;
[0032] Initializing the coding interval, taking the first source symbol in the quantization model parameters as the target source symbol;
[0033] Adding the product of the length of the coding interval and the cumulative probability of the next source symbol to the lower limit of the coding interval, and using the calculation result to update the upper limit of the coding interval, wherein the next source symbol is the next source symbol among the target source symbols when various source symbols are arranged in the preset order;
[0034] Adding the product of the length of the coding interval and the cumulative probability of the target source symbol to the lower limit of the coding interval, and using the calculation result to update the lower limit of the coding interval;
[0035] Updating the next signal source symbol of the target signal source symbol in the quantization model parameters to the target signal source symbol, and returning to the step of adding the lower limit of the coding interval to the product of the length of the coding interval and the cumulative probability of the next signal source symbol based on the updated lower limit and upper limit of the coding interval, and updating the upper limit of the coding interval using the calculation result;
[0036] After obtaining an updated coding interval based on the last source symbol in the quantization model parameter, selecting a number from the coding interval, and obtaining a coding result of the quantization model parameter based on the selected number;
[0037] The encoding result and the dequantizer are sent to a server participating in horizontal federated learning, so that the server can decode the encoding result to obtain the quantization model parameters.
[0038] To achieve the above object, the present invention further provides a federated learning communication system, the federated learning communication system comprising a horizontal federated learning server and multiple clients;
[0039] The client is configured to use the original model parameters to train a quantizer and a dequantizer during any round of federated aggregation in horizontal federated learning;
[0040] The client is further configured to quantize the original model parameters using the quantizer to obtain quantized model parameters;
[0041] The client is further configured to send the quantization model parameters and the dequantizer to a server participating in horizontal federated learning;
[0042] The server is configured to receive, during any round of federated aggregation in horizontal federated learning, quantized model parameters and dequantizers sent by each client participating in the current round of federated aggregation, wherein, for any target client among the clients, the quantized model parameters sent by the target client are obtained by the target client using the quantizer of the target client to quantize the original model parameters of the target client, and the quantizer and dequantizer of the target client are obtained by training the target client using the original model parameters of the target client;
[0043] The server is further configured to use a dequantizer corresponding to the target client to dequantize the quantization model parameters corresponding to the target client to obtain restored model parameters;
[0044] The server is further configured to, after obtaining the restoration model parameters corresponding to each client participating in the current round of federated aggregation, aggregate the restoration model parameters of each client to obtain an aggregate model parameter, and feed the aggregate model parameter back to each client;
[0045] The client is further configured to receive aggregation model parameters fed back by the server, and use the aggregation model parameters to update the local federation model and then enter the next round of federation aggregation.
[0046] To achieve the above-mentioned objectives, the present invention also provides a federated learning communication device, which includes: a memory, a processor, and a federated learning communication program stored in the memory and executable on the processor. When the federated learning communication program is executed by the processor, the steps of the federated learning communication method described above are implemented.
[0047] In addition, to achieve the above-mentioned purpose, the present invention also proposes a computer-readable storage medium, on which a federated learning communication program is stored. When the federated learning communication program is executed by a processor, the steps of the federated learning communication method described above are implemented.
[0048] In an embodiment of the present invention, each first client uses the original model parameters to train a quantizer and a dequantizer, uses the quantizer to quantize the original model parameters to obtain quantized model parameters, and uploads them to the server. Compared with uploading the original model parameters, the communication volume of each client in the horizontal federated learning process is greatly reduced, the communication cost of each client is reduced, and the communication efficiency of the horizontal federated learning is improved; and each first client uses its own original model parameters to train its own exclusive quantizer, so that when the training data set of each client is not independent and identically distributed, the accuracy of the quantization of the model parameters of each client can be guaranteed; and each first client uploads the trained dequantizer corresponding to the quantizer to the server, and the server uses the dequantizer to dequantize the quantized model parameters and then aggregates them, avoiding the situation where the quantized model parameters obtained by each client using its own exclusive quantizer cannot be aggregated due to different quantization levels, thereby ensuring the normal progress of horizontal federated learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 This is a flowchart of the first embodiment of the federated learning communication method of the present invention;
[0050] Figure 2 This is a flow chart of the second embodiment of the federated learning communication method of the present invention;
[0051] Figure 3 This is a flowchart of the third embodiment of the federated learning communication method of the present invention;
[0052] Figure 4 Schematic diagram of a federated learning communication process between a server and a client involved in an embodiment of the present invention;
[0053] Figure 5 This is a flowchart of the fourth embodiment of the federated learning communication method of the present invention;
[0054] Figure 6This is a schematic diagram of the structure of the hardware operating environment involved in the embodiment of the present invention.
[0055] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0056] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0057] Reference Figure 1 , Figure 1 Schematic diagram of the first embodiment of the federated learning communication method of the present invention.
[0058] The embodiment of the present invention provides an embodiment of a federated learning communication method. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than here. In this embodiment, the execution subject of the federated learning communication method may be a server participating in horizontal federated learning. The server may be implemented using a personal computer, a smart phone, a server, or other devices, which is not limited in this embodiment. In this embodiment, the federated learning communication method includes the following steps:
[0059] Step S10: In any round of federated aggregation in horizontal federated learning, the quantization model parameters and dequantizers sent by each first client participating in the current round of federated aggregation are received, wherein, for any target client among each of the first clients, the quantization model parameters sent by the target client are obtained by the target client using the quantizer of the target client to quantize the original model parameters of the target client, and the quantizer and dequantizer of the target client are obtained by training the target client using the original model parameters of the target client.
[0060] Each client participating in horizontal federated learning locally maintains a federated model with the same model structure. Each client has its own training dataset. The feature dimensions of the training datasets of each client are the same, but the sample dimensions are different. Therefore, by combining the clients for horizontal federated learning, the number of samples can be expanded and the training effect of the federated model can be improved.
[0061] During horizontal federated learning, each client performs multiple rounds of federated aggregation with the server. In each round, each client trains the federated model it maintains using its own training dataset, obtains the local model parameters of the client in that round of federated aggregation, and then uploads the local model parameters to the server. The server aggregates the local model parameters of each client to obtain the aggregated model parameters, which are then distributed to each client. Each client updates its maintained federated model using the aggregated model parameters. Based on the updated federated model, the next round of federated aggregation is performed. This iterative cycle continues until a pre-set stop condition is met, and the final updated federated model is used as the trained federated model. In different application scenarios, the federated models trained by each client through horizontal federated learning can be applied to perform different tasks. That is, the clients can jointly construct a federated model based on different tasks and use the federated model to perform the task. The model structure of the federated model can be set or selected according to the needs of the specific application scenario, and is not limited in this embodiment. For example, a neural network model, a linear regression model, a logistic regression model, etc. can be used. Different tasks require different training datasets. For example, in one feasible implementation, the client can be a device deployed at a financial institution (such as a bank), and the server can be a device deployed at a third-party institution trusted by each financial institution. Each financial institution maintains user data for different users, and the user data includes data items that are associated with the user's credit risk, such as the number of loans taken by the user, the number of overdue loans, etc. Each financial institution can build a training data set based on the maintained user data, and each client uses its own training data set to conduct horizontal federated learning under the coordination of the server, and builds a federated model for predicting the user's credit risk without leaking the user's privacy. In other feasible implementations, each client can also be based on other types of tasks, use other types of training data sets, and train a federated model for performing other tasks; for example, each client can be a web camera, a mobile phone, or a computer, and the training data set can be a data set consisting of multiple images, multiple videos, or multiple audio clips. The trained federated model can be used for target recognition in images, video analysis in videos, speech recognition in audio, etc.
[0062] During federated learning communication, because data uplink efficiency is far lower than download efficiency, communication bottlenecks primarily occur when clients upload model parameters. Each client must upload its local model parameters to the server during each round of federated aggregation. Complex federated models often require a large amount of local model parameter data. Over tens of thousands of iterations on large datasets, the communication volume per client can exceed 1PB. This significant amount of communication leads to high communication costs and unpredictable impacts on the communication efficiency of horizontal federated learning.
[0063] In this embodiment, a method of quantizing model parameters is proposed to reduce the client's communication volume, reduce communication costs, and improve the communication efficiency of horizontal federated learning. The client's local model parameters are trained to obtain a quantizer and a dequantizer. The client's local model parameters are quantized by a client-specific quantizer, which adapts to the situation in which the training data sets of each client in a specific application scenario often present Non-IID (non-independent and identically distributed) characteristics, thereby improving the accuracy of quantization. The quantized model parameters are then dequantized by the dequantizer, so that the server can aggregate the model parameters of each client to ensure the normal operation of horizontal federated learning.
[0064] The following describes the federated communication method in this embodiment by taking any round of federated aggregation process as an example.
[0065] The clients participating in horizontal federated learning are referred to as second clients, and the clients participating in this round of federated aggregation are referred to as first clients for the purpose of distinction. In this embodiment, there is no restriction on whether the clients participating in each round of federated aggregation are the same. That is, in each round of federated aggregation, the clients that upload model parameters can be all second clients or some of the second clients. The operations performed by each first client in this round of federated aggregation are the same, so the following description will take a first client as an example, and the first client will be referred to as the target client for the purpose of distinction. The target client uses the training data set owned by the target client to train the federated model it maintains, and obtains the local model parameters of the target client in this round of federated aggregation. The model parameters are referred to as original model parameters for the purpose of distinction. In a feasible implementation, when the task of federated learning is to train a federated model for predicting user credit risk, the target client uses the training data set constructed by the user data it owns to train the federated model it maintains, and obtains the original model parameters of the target client in this round of federated aggregation.
[0066] The target client quantizes its original model parameters in this round of federated aggregation, and the quantization result is called quantized model parameters for distinction. Quantization of model parameters refers to approximating the model parameters represented by high bit width (such as Float32) with lower bit width (such as INT8, INT4, etc.), compressing the original network by reducing the number of bits required for the model parameters. In this embodiment, the original model parameters of the target client are used to train to obtain a quantizer and dequantizer specific to the target client. The quantizer is used to quantize the original model parameters, and the dequantizer is used to dequantize the original model parameters quantized by the quantizer. The purpose of dequantization is to restore the model parameters with the same number of bits as the original model parameters, that is, after the quantized object is quantized by the quantizer, the result obtained by dequantization by the dequantizer is the same as the number of bits of the quantized object. In this embodiment, there is no restriction on whether the inverse quantization is lossless, that is, the result of the inverse quantization can be completely consistent with the quantized object, that is, lossless inverse quantization, or there can be a certain amount of information loss compared to the quantized object, that is, lossy inverse quantization. Different effects are achieved according to the different training methods of the selected quantizer and inverse quantizer. In this embodiment, there is no restriction on the method of using the original model parameters to train the quantizer and inverse quantizer.
[0067] The target client sends the quantization model parameters and dequantizer obtained in this round of federated aggregation to the server. The server receives the quantization model parameters and dequantizer sent by each first client.
[0068] Step S20 : Dequantizing the quantization model parameters corresponding to the target client using the dequantizer corresponding to the target client to obtain the restored model parameters.
[0069] The server uses the dequantizer corresponding to the target client to dequantize the quantization model parameters corresponding to the target client. The dequantized results are referred to as restored model parameters for distinction. The number of bits in the restored model parameters is the same as the number of bits in the original model parameters. After the server dequantizes the quantization model parameters of each first client, it obtains the restored model parameters corresponding to each first client.
[0070] Step S30 : After obtaining the restoration model parameters corresponding to each of the first clients, aggregating the restoration model parameters of each of the first clients to obtain an aggregated model parameter, and feeding the aggregated model parameter back to each of the first clients.
[0071] In specific application scenarios, the training data sets owned by each first client are often non-independent and identically distributed, so the original model parameters of each first client may also be non-uniformly distributed. Therefore, the quantizers obtained by each first client using their own original model parameters for training are also different, which is specifically manifested in that the quantization levels may be different. For example, although the number of bits of the original model parameters of each first client is the same, the quantizer of a first client quantizes the original model parameters to 8 bits, and the quantizer of another first client may quantize the original model parameters to 4 bits. Since the number of bits of the quantized model parameters sent by each first client may be different, the server cannot directly aggregate the various quantized model parameters. In this embodiment, the server first uses an inverse quantizer to inverse quantize the quantized model parameters so that the restored model parameters corresponding to each first client have the same number of bits, thereby enabling the server to aggregate the restored model parameters of each first client to obtain aggregated model parameters.
[0072] In this embodiment, each first client uses the original model parameters to train a quantizer and a dequantizer, uses the quantizer to quantize the original model parameters to obtain quantized model parameters, and uploads them to the server. Compared with uploading the original model parameters, the communication volume of each client in the horizontal federated learning process is greatly reduced, the communication cost of each client is reduced, and the communication efficiency of the horizontal federated learning is improved; and each first client uses its own original model parameters to train its own exclusive quantizer, so that when the training data set of each client is not independent and identically distributed, the accuracy of the quantization of the model parameters of each client can be guaranteed; and each first client uploads the trained dequantizer corresponding to the quantizer to the server, and the server uses the dequantizer to dequantize the quantized model parameters and then aggregates them, avoiding the situation where the quantized model parameters obtained by each client using its own exclusive quantizer cannot be aggregated due to different quantization levels, thereby ensuring the normal progress of horizontal federated learning.
[0073] In one feasible embodiment, after receiving the aggregate model parameters fed back by the server, each first client uses the aggregate model parameters to update its local federated model. Specifically, the aggregate model parameters may be used as the model parameters of its local federated model. Based on the updated federated model, each first client performs the next round of federated aggregation, and repeats the cycle until a pre-set stop condition is met, and the final updated federated model is used as the trained federated model. In one feasible embodiment, when the task of federated learning is to train a federated model for predicting user credit risk, the trained federated model obtained by the client deployed at the financial institution can be used to predict the user's risk credit value. The client can obtain the user data of the target user, and the data items contained in the user data are the same as the data items in the user data used by the client when training the local federated model when participating in horizontal federated learning. The client inputs the user data of the target user into the trained federated model for prediction to obtain the credit risk value of the target user.
[0074] In one feasible implementation, the quantizer of the target client is obtained by training the original model parameters of the target client using the LQ-Nets algorithm, the quantizer includes a quantization reference vector, and the dequantizer is used to divide the object of the dequantization operation by the quantization reference vector to dequantize the object of the dequantization operation. Step S20 includes:
[0075] Step S201 : using a dequantizer corresponding to the target client, dividing the quantization model parameters corresponding to the target client by the quantization reference vector corresponding to the target client, to obtain restoration model parameters.
[0076] LQ-Nets is a neural network quantization method for non-uniform quantization. It uses a data-based dynamic quantization method. By statistically analyzing the data and dynamically updating the quantization parameters, it can adaptively adjust the quantization interval size according to the actual data distribution, thereby ensuring accuracy while reducing the computational complexity and storage space after quantization. In this embodiment, it is proposed that the LQ-Nets algorithm can be used to train the quantizer of the target client. Furthermore, based on the characteristics of the quantizer trained using the LQ-Nets algorithm, it is proposed that the quantization reference vector in the quantizer can be used for dequantization. Specifically, the dequantized object is dequantized by dividing the dequantized object by the quantization reference vector. Then, the dequantizer can refer to a functional module that divides the dequantized object by the quantization reference vector, or the dequantizer can represent a function model for implementing the computational process of dividing the dequantized object by the quantization reference vector. In a specific embodiment, the target client sending the dequantizer to the server can specifically refer to sending the trained quantization reference vector to the server. The server divides the quantization model parameters sent by the target client by the quantization reference vector sent by the target client to obtain the restoration model parameters corresponding to the target client.
[0077] For example, the quantization reference vector in the quantizer trained by the client is v = [1, 2, 4, 8], and the model parameter to be quantized is w = [0.3, 0.5, 0.5, 0.6]. The model parameter is quantized using the quantization reference vector to obtain the quantized model parameter Q(w) = v*w = [0.3, 1, 2, 4.8]. The server can restore the model parameter w by dividing the quantized model parameter by the quantization reference vector, that is, w = Q(w)*1 = [0.3, 0.5, 0.5, 0.6].
[0078] Assume that there are n clients. First, each client c i During the quantizer training phase, a dedicated quantizer Q is obtained. i , train the quantizer Q i The process is based on the client c i The model parameter w i , using the QEM (Quantization Error Minimization) algorithm to continuously update the optimal quantization bit B (ie, quantization level), since the model parameter w i It is calculated based on the client's unique training sample data. Therefore, when the data of each client participating in horizontal federated learning is heterogeneous, each client c i Different quantizers Q can be learned i .
[0079] After the quantizer training is completed, the client ci According to the quantizer Q i For the model parameter w i For example, when the quantizer Q i When the corresponding quantization level is 8, the data compression interval is [0, 255], and the model parameter w i Mapping to this range, we can get the quantized model weight Q(w i ).
[0080] Since each client is trained to obtain a different quantizer Q i Therefore, the quantization accuracy of the model parameters of each client is not exactly the same. Therefore, the server needs to dequantize the model parameters of each client before performing aggregation operations.
[0081] The client uses an adaptive model quantization method based on LQ-Nets to convert model parameters from high precision to low precision, which not only saves model storage space but also improves computing speed, ultimately achieving the goal of optimizing federated communication efficiency.
[0082] In one feasible implementation, the client and server can perform horizontal federated learning in the following manner.
[0083] w={w1,w2,…,w n} represents the model parameters, n is the number of model parameters in the client, v is the quantized reference vector, T is the total number of federated aggregation rounds, and M is the number of quantizer training rounds. Q(w) represents the quantized model parameters obtained by quantizing w using the quantizer.
[0084] On the server side, in each round of federation aggregation from round 1 to round T, the following steps are executed:
[0085] based on Dequantize Q(w) of each client to get w;
[0086] based on The model weights sent by the client are averaged and aggregated to obtain w, where k represents the total number of clients participating in this round of federated aggregation;
[0087] Send w to the client for the next round of model iteration.
[0088] On the client side, in each round of federated aggregation from round 1 to round T, the following operations are executed:
[0089] During the quantizer training phase:
[0090] Use v 0 =v=w / 2 n-1 Initialize the reference vector;
[0091] Then perform M rounds of iterative training on the quantizer. In each round of iterative training, execute:
[0092] Based on B=v T e l ,e l ∈{-1,1} K ,K=B t-1 , according to v t-1 Calculate B t ;
[0093] Based on v=(BB T ) -1 Bw, according to B t Calculate v t ;
[0094] Based on error = argming∫p(w)(Q(w)-w) 2 The dx calculation error determines whether the training has converged, where p(w) is the probability density function.
[0095] After M rounds of iterative training, based on v t+1 =αv t-1 +(1-α)v t ,α=0.9, update v by MV (moving average method).
[0096] During the quantizer training phase:
[0097] Based on Q(w)=v*w, Q(w) is calculated according to v;
[0098] Send the quantized model parameters Q(w) and v to the server.
[0099] Based on the above first embodiment, a second embodiment of the federated learning communication method of the present invention is proposed. In this embodiment, referring to Figure 2 , the federated learning communication method further includes:
[0100] Step S40: receiving the model loss sent by each of the first clients.
[0101] In this embodiment, in order to further improve the communication efficiency of horizontal federated learning, it is proposed to select some clients from each second client to participate in this round of federated aggregation during each round of federated aggregation, so as to improve the communication efficiency of horizontal federated learning by reducing the communication frequency.
[0102] All clients participating in horizontal federated learning are called second clients, and the clients participating in this round of federated aggregation among the second clients are called first clients to distinguish them.
[0103] During this round of federation, the server can select clients from among the second clients to participate in the next round of federation based on the model losses uploaded by each first client. Participating in a round of federation means uploading the local model parameters trained in that round of federation to the server for aggregation.
[0104] During this round of federated aggregation, each first client, after obtaining its local model parameters (i.e., the original model parameters), can calculate the model loss of its maintained federated model in this round of federated aggregation based on its local model parameters. Each first client sends its calculated model loss to the server, which then selects a client from among the second clients to participate in the next round of federated aggregation based on the model loss of each first client.
[0105] Step S50 , calculating a reduction in the model loss sent by the target client during the current round of federated aggregation compared to the model loss sent by the target client during the previous round of federated aggregation.
[0106] Take any one of the first clients as an example for explanation, and refer to the client as the target client for distinction. After receiving the model loss sent by the target client during the current round of federated aggregation (i.e., this round of federated aggregation), the server can calculate the reduction in the model loss compared to the model loss sent by the target client during the previous round of federated aggregation. In machine learning, the smaller the model loss, the better the model fits the data, and vice versa. The purpose of horizontal federated learning is to reduce the model loss of the federated model through multiple rounds of iterations. In each round of federated aggregation in horizontal federated learning, the difference in model loss represents the difference in the impact of the client on the global model during this round of federated aggregation. The reduction in the target client's model loss in this round compared to the model loss in the previous round reflects the contribution made by the target client to reducing the model loss of the federated model during this round of federated aggregation. The greater the reduction, the greater the contribution.
[0107] Step S60: Select the clients participating in the next round of federated aggregation from the second clients participating in the horizontal federated learning based on the reduction amount corresponding to each of the first clients; wherein, when the reduction amount corresponding to the target client is greater than the average value of the reduction amounts corresponding to each of the first clients, the probability of the target client being selected to participate in the next round of federated aggregation is greater than the probability of the target client being selected to participate in the current round of federated aggregation; when the reduction amount corresponding to the target client is less than or equal to the average value, the probability of the target client being selected to participate in the next round of federated aggregation is less than the probability of the target client being selected to participate in the current round of federated aggregation.
[0108] After calculating the reduction amount corresponding to each first client, the server can select clients from each second client to participate in the next round of federated aggregation based on the reduction amount corresponding to each first client. In other words, it can select which clients need to upload model parameters in the next round of federated aggregation. In this embodiment, the specific method used by the server to select clients based on the reduction amount is not limited. In the specific implementation, the method can be selected as needed. For the target client, the selected method must satisfy the following requirements: if the reduction amount corresponding to the target client is greater than the average of the reduction amounts corresponding to each second client, the probability of the target client being selected to participate in the next round of federated aggregation is greater than the probability of the target client being selected to participate in the current round of federated aggregation, and if the reduction amount corresponding to the target client is less than or equal to the average, the probability of the target client being selected to participate in the next round of federated aggregation is less than the probability of the target client being selected to participate in the current round of federated aggregation.
[0109] It should be noted that the second clients other than the first client are referred to as third clients; the probability of the third client being selected to participate in the next round of federation aggregation can be the same as the probability of the third client being selected to participate in the current round of federation aggregation, which is not limited in this embodiment.
[0110] It should be noted that during the previous round of federated aggregation, the server also calculated the corresponding reduction amount of each second client according to the model loss sent by each client participating in the previous round of federated aggregation, and then selected the clients participating in the current round of federated aggregation based on the reduction amount. Therefore, the target client is selected to participate in the current round of federated aggregation with a certain probability. In the current round of federated aggregation, the probability of the target client being selected to participate in the current round of federated aggregation can be adjusted based on the comparison result between the reduction amount corresponding to the target client and the average value, so as to change the probability of the target client being selected in the next round of federated aggregation.
[0111] In this embodiment, the reduction in model loss of each first client in this round of federated aggregation compared to the model loss in the previous round of federated aggregation is used as the basis for selecting clients to participate in the next round of federated aggregation. When the loss reduction of a certain client is greater than the average loss reduction of each client, the probability of the client being selected is increased. Conversely, when the loss reduction of the client is smaller than the average, the probability of the client being selected is decreased. In this way, while the communication efficiency of horizontal federated learning is improved by selecting some clients to participate in federated aggregation, as the number of rounds of federated aggregation increases, the selected clients can increasingly represent all clients, thereby ensuring the modeling effect of horizontal federated learning.
[0112] In one feasible implementation, step S60 includes:
[0113] Step S601: When the reduction amount corresponding to the target client is greater than the average value of the reduction amounts corresponding to each of the first clients, the α parameter value of the beta distribution corresponding to the target client is increased to update the beta distribution corresponding to the target client, wherein, before the first round of federal aggregation, the same beta distribution is initialized for each of the second clients, and the initialization values of the α parameter and the β parameter in the beta distribution are equal.
[0114] The method of selecting clients based on random strategies in related technologies will lead to biased sampling, as shown in the following formula (1), and some clients with unique data distribution may be difficult to be selected, thereby affecting the convergence of the global model.
[0115] p1w1+p2w2+…+p n w n ≠w1+w2+…+w n (1)
[0116] The left formula represents the expected value based on client sampling, and the right formula represents the expected value of model aggregation based on unsampled data. i represents the probability that client i is selected.
[0117] In this embodiment, in order to solve the problems existing in the related art, a feasible implementation method is proposed for selecting clients to participate in the next round of federation aggregation according to the reduction amount of the first client.
[0118] Before starting horizontal federated learning, that is, before performing the first round of federated aggregation, the server can initialize a beta distribution for each second client. The beta distributions of each client are the same, and the initialization values of the α parameter and the β parameter in each beta distribution are equal, for example, both are 1.
[0119] During each round of federated aggregation, the server adjusts the parameter values of the α and β parameters in the beta distribution of each second client based on the comparison of the corresponding reduction amount of the clients participating in the federated aggregation with the average value, so as to update the beta distribution of each second client, thereby adjusting the probability of each second client being selected to participate in each round of federated aggregation.
[0120] During a round of federated aggregation, after the server obtains the reduction corresponding to the target client, if the reduction corresponding to the target client is greater than the average reduction corresponding to each first client, the α parameter value of the Beta distribution corresponding to the target client is increased, and the β parameter value may remain unchanged. For example, in a feasible implementation, the server may increase the α parameter value of the Beta distribution corresponding to the target client by adding 1, that is, it can be expressed as αi+1 =α i +1, β i+1 =β i .
[0121] Step S602 : When the reduction amount corresponding to the target client is less than or equal to the average value, the β parameter value of the beta distribution corresponding to the target client is increased to update the beta distribution corresponding to the target client.
[0122] After the server obtains the reduction amount corresponding to the target client, if the reduction amount corresponding to the target client is less than or equal to the average of the reduction amounts corresponding to the first clients, the β parameter value of the beta distribution corresponding to the target client is increased, and the α parameter value may remain unchanged. For example, in a feasible implementation, the server may increase the β parameter value of the beta distribution corresponding to the target client by adding 1, that is, it can be expressed as β i+1 =β i +1, α i+1 =α i .
[0123] Step S603: After updating the beta distribution of each of the first clients, a random number is generated using the beta distribution of each of the second clients. The second clients are sorted in descending order according to the random number, and the second clients that are ranked in the front by a preset number are selected as the clients participating in the next round of federated aggregation.
[0124] The beta distribution corresponding to the second clients except the first client may remain unchanged.
[0125] After updating the Beta distribution for each first client, the server can generate a random number using the Beta distribution for each second client. The second clients are sorted by their corresponding random numbers, with clients with larger random numbers being ranked higher. A preset number of second clients can then be selected as clients participating in the next round of federated aggregation. The preset number is less than the total number of second clients and can be set as needed without limitation in this embodiment.
[0126] It should be noted that when the α parameter value of the target client's Beta distribution increases, the probability of generating a larger random number according to the Beta distribution increases, thereby making the target client more likely to be ranked higher in the ranking, thereby increasing the probability of the target client being selected. Conversely, when the β parameter value of the target client's Beta distribution increases, the probability of generating a larger random number according to the Beta distribution increases, thereby making the target client more likely to be ranked lower in the ranking, thereby reducing the probability of the target client being selected.
[0127] In this embodiment, in order to ensure that each client can participate in the model iteration process during the horizontal federated learning process, the server can rely on the number of times each client has been selected historically as prior information when selecting the client, and combine the model loss to obtain the posterior distribution. As the number of iterations increases, the selected client can represent all clients, ensuring that the client model weights after sampling meet unbiased estimation.
[0128] When there are N clients, the server needs to maintain 2*N Beta parameters. It should be noted that when the initial Beta distribution is used, for any client i∈{1,…,N}, the Beta distribution is a uniform distribution in [0,1]. Therefore, the prior information does not bring any preference in the initial stage. As the model continues to iterate, α i +β i As the value of gets larger, the random numbers generated using it are essentially in the center of the Beta distribution, close to the average return. The client selection method proposed in this embodiment utilizes the client's prior information and the model loss value, ensuring that more valuable clients are selected. This method is also unbiased.
[0129] Based on the above first and / or second embodiments, a third embodiment of the federated learning communication method of the present invention is proposed. In this embodiment, referring to Figure 3 The quantized model parameters sent by the target client are obtained by quantizing the original model parameters of the target client using the target client's quantizer and then encoding them. Step S20 includes:
[0130] Step S202 : decoding the encoded quantization model parameters corresponding to the target client to obtain a decoding result, and then dequantizing the decoding result using a dequantizer corresponding to the target client to obtain a restored model parameter.
[0131] In this embodiment, it is proposed to encode the quantized model parameters before the target client sends them to the server, and send the encoded quantized model parameters to the server to further compress the amount of data uploaded by the client to the server, thereby further improving the communication efficiency of horizontal federated learning.
[0132] Accordingly, after receiving the encoded quantization model parameters corresponding to the target client, the server first decodes the encoded quantization model parameters, and then dequantizes the decoded results using the dequantizer corresponding to the target client to obtain the restored model parameters.
[0133] In this embodiment, the encoding method used by the client is not limited and can be set as needed. Correspondingly, the decoding method of the server is also not limited and can correspond to the encoding method used by the client.
[0134] In a feasible implementation, a client encoding method and a server decoding method are proposed.
[0135] The pre-encoded quantization model parameters include multiple quantized values, each corresponding to a parameter value after quantization of the model parameters. The client can use each quantized value in the pre-encoded quantization model parameters as a source symbol and calculate the cumulative probabilities corresponding to each source symbol when arranged in a preset order. The preset order can be pre-set as needed and is not limited herein. For example, it can be arranged in ascending order of numerical value. The occurrence probabilities of various source symbols can be pre-set, either as needed or by calculating the proportion of the number of occurrences of a particular source symbol in the quantization model parameters to the total number of source symbols in the quantization model parameters as the occurrence probability of that source symbol. The cumulative probabilities of various source symbols can be calculated based on the occurrence probabilities of the various source symbols. In one feasible embodiment, the cumulative probability of the source symbol can be calculated by summing the occurrence probabilities of the various source symbols that precede the source symbol when the various source symbols are arranged in the preset order. The cumulative probability of the source symbol that is ranked first is equal to 0.
[0136] After obtaining the cumulative probabilities of various source symbols, a coding interval can be initialized. In one feasible implementation, the lower limit of the coding interval can be 0, and the upper limit can be 1. The first source symbol in the pre-encoded quantization model parameters is used as the target source symbol. The lower limit of the current coding interval is added to the product of the length of the coding interval and the cumulative probability of the next source symbol. The upper limit of the coding interval is updated using the calculated result. The next source symbol is the source symbol that follows the target source symbol when the various source symbols are arranged in a preset order.
[0137] The lower limit of the current coding interval is added to the product of the length of the coding interval and the cumulative probability of the target source symbol, and the calculation result is used to update the lower limit of the coding interval.
[0138] At this point, the upper and lower limits of the coding interval have been updated, resulting in a new coding interval. The next source symbol after the target source symbol in the pre-coding quantization model parameters is used as the new target source symbol. Based on the updated lower and upper limits of the coding interval, the above steps are repeated, adding the lower limit of the coding interval to the product of the coding interval length and the cumulative probability of the next source symbol. The result is used to update the upper limit of the coding interval. This loop then continues to update the coding interval.
[0139] After obtaining an updated coding interval based on the last source symbol in the quantization model parameters, a number is selected from the last updated coding interval. The encoding result of the quantization model parameters is obtained based on the selected number, which is then used as the encoded quantization model parameters. Specifically, the selected number can be used directly as the encoding result, or it can be processed and used as the encoding result, without limitation. The client sends the encoding result and the dequantizer to the server participating in horizontal federated learning, which decodes the encoding result to obtain the quantization model parameters.
[0140] The server decodes the encoded quantization model parameters in a manner that is the inverse of the encoding process. The step of decoding the encoded quantization model parameters corresponding to the target client to obtain a decoding result in step S202 includes:
[0141] In step S2021, each quantized value of the quantized model parameter before encoding corresponding to the target client is used as a source symbol, and the cumulative probabilities corresponding to the various source symbols when arranged in a preset order are obtained.
[0142] For the target client, the server also uses the quantized values of the pre-encoded quantization model parameters corresponding to the target client as source symbols and obtains the cumulative probabilities corresponding to the various source symbols when arranged in a preset order. The preset order and cumulative probabilities can be obtained from the target client.
[0143] Step S2022: Use the encoded quantization model parameters corresponding to the target client as the target decoding object, determine the target source symbol from various source symbols, and satisfy that the target decoding object is in the interval formed by the cumulative probability of the target source symbol and the cumulative probability of the next source symbol, wherein the next source symbol is the next source symbol among the target source symbols when various source symbols are arranged in the preset order.
[0144] The server uses the encoded quantization model parameters corresponding to the target client as the target decoding object. A source symbol (referred to as a target source symbol) is determined from the various source symbols. The target source symbol must satisfy the following conditions: the target decoding object is within the interval formed by the cumulative probability of the target source symbol and the cumulative probability of the next source symbol, where the next source symbol is the next source symbol in the target source symbol when the various source symbols are arranged in a preset order.
[0145] Step S2023: Use the parameter value corresponding to the target signal source symbol as a parameter value obtained by decoding, subtract the cumulative probability of the target signal source symbol from the target decoding object, and then divide it by the preset occurrence probability of the target signal source symbol, update the calculation result to the target decoding object, and return to execute the step of determining the target signal source symbol from various signal source symbols in step S2022.
[0146] After finding the target source symbol, the parameter value corresponding to that target source symbol can be used as one of the parameters in the decoded quantization model. The server subtracts the cumulative probability of the target source symbol from the target decoding object, then divides it by the preset probability of occurrence of the target source symbol. The result is used as the new target decoding object. Based on this new target decoding object, the server then executes the step of determining the target source symbol from various source symbols, thus entering a loop and decoding the new parameter value.
[0147] Step S2024: After the last parameter value is obtained through decoding, the quantization model parameters before encoding corresponding to the target client are obtained according to the decoded parameter values, and the quantization model parameters before encoding corresponding to the target client are used as the decoding result.
[0148] The client can pre-send the number of parameter values in the pre-encoded quantization model parameters to the server. This allows the server to, after decoding the last parameter value, obtain the pre-encoded quantization model parameters corresponding to the target client based on the decoded parameter values, i.e., obtain the decoding result. The server can then dequantize the decoded result to obtain the restored model parameters. In one feasible embodiment, the server can arrange the parameter values in descending order according to the decoding time, thereby restoring the order of the parameter values in the quantization model parameters.
[0149] In one feasible implementation, during the tth round of federation aggregation, the communication process between the client and the server can refer to Figure 4 shown.
[0150] ① The server selects k clients to send model parameters based on the method for adaptively selecting clients in the second embodiment.
[0151] ②The server will model parameters Using the encoding method in the third embodiment above, the encoding is compressed to Sent to the selected k clients Ci (i∈1,…,k).
[0152] ③k clients download the compressed model parameters from the server
[0153] ④k clients will Decoded use Update W t Get W t+1 .
[0154] ⑤k clients use the quantizer training method in the first embodiment to train the quantizer, and use the quantizer to t+1 Quantized as Q(W t+1 ); using the encoding method in the third embodiment above to Q(W t+1 ) is compressed into Q(W t+1 ) c .
[0155] ⑥k clients upload the compressed model parameters Q(W t+1 ) c .
[0156] ⑦The server will Q(W t+1 ) c Decoding gets Q(W t+1 );
[0157] Use the inverse quantizer to quantize Q(W t+1 ) dequantize to get W t+1 ;
[0158] based on Aggregate model parameters for each client;
[0159] Will Encoded compression Then based on Proceed to the next round of federation aggregation.
[0160] This cycle iterates until the stopping condition is met, and the client obtains the trained federated model.
[0161] The communication loss of horizontal federated learning can be described by formula (2).
[0162] L=N iter *f*|W|*(H(ΔW up / down )+η) (2)
[0163] where N iter represents the number of iterations for each client, f represents the communication frequency, |W| represents the size of the model, H(△W up / down ) represents the entropy of the model parameter updates uploaded and downloaded by the participants during the training process, and η is the efficiency of the encoding.
[0164] In this embodiment, by selecting some clients to participate in federation aggregation in each round of federation aggregation, the communication frequency f is reduced, the data exchange between the participants and the central server is reduced, and the communication efficiency is improved; by quantizing the model parameters, the entropy value H(△W up / down ), and the quantization method in the above-mentioned first embodiment can ensure higher quantization accuracy; by encoding and compressing the model weights through the more efficient encoding method in the above-mentioned third embodiment, the encoding efficiency η can be improved, and the transmission bandwidth utilization rate can be improved.
[0165] Based on the first, second and / or third embodiments above, a fourth embodiment of the federated learning communication method of the present invention is proposed. In this embodiment, the execution subject of the federated learning communication method can be any target client among the first clients participating in horizontal federated learning. The client can be implemented using a personal computer, a smart phone, a server and other devices, which is not limited in this embodiment. Figure 5 , the federated learning communication method comprises the following steps:
[0166] In step A10, during any round of federated aggregation in horizontal federated learning, the original model parameters are used for training to obtain a quantizer and a dequantizer.
[0167] Step A20: quantize the original model parameters using the quantizer to obtain quantized model parameters.
[0168] Step A30: Send the quantization model parameters and the dequantizer to the server participating in horizontal federated learning, so that the server uses the dequantizer corresponding to the target client to dequantize the quantization model parameters corresponding to the target client to obtain the restored model parameters, and after obtaining the restored model parameters corresponding to each of the first clients, aggregate the restored model parameters of each of the first clients to obtain the aggregated model parameters.
[0169] In this embodiment, the specific implementation of the above steps A10 to A30 can refer to the specific implementation of the above steps S10 to S30 in the first embodiment, and will not be described in detail here.
[0170] In one feasible implementation, step A30 includes:
[0171] Step A301: taking each quantized value in the quantized model parameter as a signal source symbol, and calculating the cumulative probabilities corresponding to each signal source symbol when arranged in a preset order.
[0172] Step A302: Initialize the coding interval and use the first source symbol in the quantization model parameters as the target source symbol.
[0173] Step A303, add the lower limit of the coding interval to the product of the length of the coding interval and the cumulative probability of the next source symbol, and use the calculation result to update the upper limit of the coding interval, wherein the next source symbol is the next source symbol in the target source symbols when various source symbols are arranged in the preset order.
[0174] Step A304: Add the product of the length of the coding interval and the cumulative probability of the target signal source symbol to the lower limit of the coding interval, and use the calculation result to update the lower limit of the coding interval.
[0175] Step A305, update the next source symbol of the target source symbol in the quantization model parameters to the target source symbol, and based on the coding interval after updating the lower limit and upper limit, return to execute the step of adding the lower limit of the coding interval to the product of the length of the coding interval and the cumulative probability of the next source symbol, and use the calculation result to update the upper limit of the coding interval.
[0176] Step A306: After obtaining the updated coding interval based on the last source symbol in the quantization model parameter, select a number from the coding interval, and obtain the coding result of the quantization model parameter based on the selected number.
[0177] Step A307: Send the encoding result and the dequantizer to the server participating in horizontal federated learning, so that the server can decode the encoding result to obtain the quantization model parameters.
[0178] In this embodiment, the specific implementation of the above steps A301 to A307 can refer to the specific implementation of step S202 in the above first embodiment, and will not be described in detail here.
[0179] In addition, an embodiment of the present invention further provides a federated learning communication system, wherein the federated learning communication system includes a horizontal federated learning server and multiple clients;
[0180] The client is configured to use the original model parameters to train a quantizer and a dequantizer during any round of federated aggregation in horizontal federated learning;
[0181] The client is further configured to quantize the original model parameters using the quantizer to obtain quantized model parameters;
[0182] The client is further configured to send the quantization model parameters and the dequantizer to a server participating in horizontal federated learning;
[0183] The server is configured to receive, during any round of federated aggregation in horizontal federated learning, quantized model parameters and dequantizers sent by each client participating in the current round of federated aggregation, wherein, for any target client among the clients, the quantized model parameters sent by the target client are obtained by the target client using the quantizer of the target client to quantize the original model parameters of the target client, and the quantizer and dequantizer of the target client are obtained by training the target client using the original model parameters of the target client;
[0184] The server is further configured to use a dequantizer corresponding to the target client to dequantize the quantization model parameters corresponding to the target client to obtain restored model parameters;
[0185] The server is further configured to, after obtaining the restoration model parameters corresponding to each client participating in the current round of federated aggregation, aggregate the restoration model parameters of each client to obtain an aggregate model parameter, and feed the aggregate model parameter back to each client;
[0186] The client is further configured to receive aggregation model parameters fed back by the server, and use the aggregation model parameters to update the local federation model and then enter the next round of federation aggregation.
[0187] The various embodiments of the federated learning communication system of the present invention can refer to the various embodiments of the federated learning communication method of the present invention, and will not be repeated here.
[0188] In addition, the embodiment of the present invention also proposes a federated learning communication device, such as Figure 6 As shown, Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention. It should be noted that the federated learning communication device in the embodiment of the present invention can be a monitoring device, a smart phone, a personal computer, a server, etc., and is not specifically limited here.
[0189] like Figure 6As shown, the federated learning communication device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0190] Those skilled in the art will understand that Figure 6 The device structure shown in the figure does not constitute a limitation on the federated learning communication device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0191] like Figure 6 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a federated learning communication program. The operating system is a program that manages and controls the hardware and software resources of the device and supports the operation of the federated learning communication program and other software or programs. Figure 6 In the device shown, the user interface 1003 is mainly used for data communication with the client; the network interface 1004 is mainly used for establishing a communication connection with the server.
[0192] When the federated learning communication device serves as a server participating in horizontal federated learning, the processor 1001 may be configured to call a federated learning communication program stored in the memory 1005 and perform the following operations:
[0193] In any round of federated aggregation in horizontal federated learning, receiving quantization model parameters and dequantizers sent by each first client participating in the current round of federated aggregation, wherein, for any target client among the first clients, the quantization model parameters sent by the target client are obtained by quantizing the original model parameters of the target client using the quantizer of the target client, and the quantizer and dequantizer of the target client are obtained by training the target client using the original model parameters of the target client;
[0194] Dequantizing the quantization model parameters corresponding to the target client using the dequantizer corresponding to the target client to obtain the restored model parameters;
[0195] After obtaining the restoration model parameters corresponding to each of the first clients, the restoration model parameters of each of the first clients are aggregated to obtain an aggregate model parameter, and the aggregate model parameter is fed back to each of the first clients.
[0196] In one feasible implementation, the quantizer of the target client is obtained by training the original model parameters of the target client using the LQ-Nets algorithm, the quantizer includes a quantization reference vector, the dequantizer is used to divide the object of the dequantization operation by the quantization reference vector to dequantize the object of the dequantization operation, and the operation of dequantizing the quantization model parameters corresponding to the target client using the dequantizer corresponding to the target client to obtain the restored model parameters includes:
[0197] The quantization model parameters corresponding to the target client are divided by the quantization reference vector corresponding to the target client using the inverse quantizer corresponding to the target client to obtain the restored model parameters.
[0198] In one feasible implementation, the processor 1001 may also be configured to call the federated learning communication program stored in the memory 1005 to perform the following operations:
[0199] receiving the model loss sent by each of the first clients;
[0200] Calculate the reduction in the model loss sent by the target client during the current round of federated aggregation compared to the model loss sent by the target client during the previous round of federated aggregation;
[0201] According to the reduction amount corresponding to each first client, each client participating in the next round of federated aggregation is selected from each second client participating in horizontal federated learning; wherein, when the reduction amount corresponding to the target client is greater than the average value of the reduction amounts corresponding to each first client, the probability of the target client being selected to participate in the next round of federated aggregation is greater than the probability of the target client being selected to participate in the current round of federated aggregation; when the reduction amount corresponding to the target client is less than or equal to the average value, the probability of the target client being selected to participate in the next round of federated aggregation is less than the probability of the target client being selected to participate in the current round of federated aggregation.
[0202] In one feasible implementation, the operation of selecting, from the second clients participating in horizontal federated learning, clients to participate in the next round of federated aggregation based on the reduction amount corresponding to each first client includes:
[0203] When the reduction amount corresponding to the target client is greater than the average of the reduction amounts corresponding to the first clients, the α parameter value of the beta distribution corresponding to the target client is increased to update the beta distribution corresponding to the target client, wherein before the first round of federated aggregation, the same beta distribution is initialized for each of the second clients, and the initialization values of the α parameter and the β parameter in the beta distribution are equal;
[0204] When the reduction amount corresponding to the target client is less than or equal to the average value, increasing the β parameter value of the beta distribution corresponding to the target client to update the beta distribution corresponding to the target client;
[0205] After updating the beta distribution of each first client, a random number is generated using the beta distribution of each second client, and each second client is sorted in descending order according to the random number, and the second clients ranked in the front by a preset number are selected as the clients participating in the next round of federal aggregation.
[0206] In one feasible implementation, the quantization model parameters sent by the target client are obtained by quantizing the original model parameters of the target client using a quantizer of the target client and then encoding the quantized model parameters. The operation of dequantizing the quantization model parameters corresponding to the target client using a dequantizer corresponding to the target client to obtain the restored model parameters includes:
[0207] The encoded quantization model parameters corresponding to the target client are decoded to obtain a decoding result, and then the decoding result is dequantized using a dequantizer corresponding to the target client to obtain a restored model parameter.
[0208] In one feasible implementation, the operation of decoding the encoded quantization model parameters corresponding to the target client to obtain a decoding result includes:
[0209] Taking each quantized value of the quantized model parameter before encoding corresponding to the target client as a signal source symbol, and obtaining the cumulative probabilities corresponding to the various signal source symbols when arranged in a preset order;
[0210] The encoded quantization model parameters corresponding to the target client are used as a target decoding object, and a target source symbol is determined from various source symbols, so that the target decoding object is within an interval formed by a cumulative probability of the target source symbol and a cumulative probability of a next source symbol, wherein the next source symbol is a next source symbol among the target source symbols when the various source symbols are arranged in the preset order;
[0211] Using the parameter value corresponding to the target signal source symbol as a parameter value obtained by decoding, subtracting the cumulative probability of the target signal source symbol from the target decoding object and dividing the result by the preset occurrence probability of the target signal source symbol, updating the calculation result as the target decoding object, and returning to perform the operation of determining the target signal source symbol from the various signal source symbols;
[0212] After the last parameter value is obtained through decoding, the quantization model parameters before encoding corresponding to the target client are obtained according to the decoded parameter value.
[0213] In one feasible implementation, when the federated learning communication device serves as any target client among the first clients participating in horizontal federated learning, the processor 1001 may be configured to call the federated learning communication program stored in the memory 1005 and perform the following operations:
[0214] In any round of federated aggregation in horizontal federated learning, the original model parameters are used to train the quantizer and dequantizer;
[0215] quantizing the original model parameters using the quantizer to obtain quantized model parameters;
[0216] The quantization model parameters and the dequantizer are sent to a server participating in horizontal federated learning, so that the server uses the dequantizer corresponding to the target client to dequantize the quantization model parameters corresponding to the target client to obtain the restored model parameters, and after obtaining the restored model parameters corresponding to each of the first clients, the restored model parameters of each of the first clients are aggregated to obtain the aggregated model parameters.
[0217] In one feasible implementation, the operation of sending the quantization model parameters and the dequantizer to a server participating in horizontal federated learning includes:
[0218] Taking each quantized value in the quantized model parameter as a signal source symbol, and calculating the cumulative probabilities corresponding to the various signal source symbols when arranged in a preset order;
[0219] Initializing the coding interval, taking the first source symbol in the quantization model parameters as the target source symbol;
[0220] Adding the product of the length of the coding interval and the cumulative probability of the next source symbol to the lower limit of the coding interval, and using the calculation result to update the upper limit of the coding interval, wherein the next source symbol is the next source symbol among the target source symbols when various source symbols are arranged in the preset order;
[0221] Adding the product of the length of the coding interval and the cumulative probability of the target source symbol to the lower limit of the coding interval, and using the calculation result to update the lower limit of the coding interval;
[0222] Updating the next source symbol of the target source symbol in the quantization model parameters to the target source symbol, and returning to perform the operation of adding the lower limit of the coding interval to the product of the length of the coding interval and the cumulative probability of the next source symbol based on the updated lower limit and upper limit of the coding interval, and updating the upper limit of the coding interval using the calculation result;
[0223] After obtaining an updated coding interval based on the last source symbol in the quantization model parameter, selecting a number from the coding interval, and obtaining a coding result of the quantization model parameter based on the selected number;
[0224] The encoding result and the dequantizer are sent to a server participating in horizontal federated learning, so that the server can decode the encoding result to obtain the quantization model parameters.
[0225] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a federated learning communication program is stored. When the federated learning communication program is executed by a processor, the steps of the federated learning communication method described above are implemented.
[0226] The various embodiments of the federated learning communication device and computer-readable storage medium of the present invention may refer to the various embodiments of the federated learning communication method of the present invention, and will not be repeated here.
[0227] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0228] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0229] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0230] The above are only preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the scope of protection of the present invention.
Claims
1. A federated learning communication method, characterized in that: The federated learning communication method is applied to a server participating in horizontal federated learning, and the federated learning communication method includes the following steps: In any round of federated aggregation in horizontal federated learning, receiving quantization model parameters and dequantizers sent by each first client participating in the current round of federated aggregation, wherein, for any target client among the first clients, the quantization model parameters sent by the target client are obtained by quantizing the original model parameters of the target client using the quantizer of the target client, and the quantizer and dequantizer of the target client are obtained by training the target client using the original model parameters of the target client; Dequantizing the quantization model parameters corresponding to the target client using the dequantizer corresponding to the target client to obtain the restored model parameters; After obtaining the restoration model parameters corresponding to each of the first clients, the restoration model parameters of each of the first clients are aggregated to obtain an aggregate model parameter, and the aggregate model parameter is fed back to each of the first clients.
2. The federated learning communication method according to claim 1, wherein: The quantizer of the target client is obtained by training the original model parameters of the target client using the LQ-Nets algorithm. The quantizer includes a quantization reference vector. The dequantizer is used to divide the object of the dequantization operation by the quantization reference vector to dequantize the object of the dequantization operation. The step of dequantizing the quantization model parameters corresponding to the target client using the dequantizer corresponding to the target client to obtain the restored model parameters includes: The quantization model parameters corresponding to the target client are divided by the quantization reference vector corresponding to the target client using the inverse quantizer corresponding to the target client to obtain the restored model parameters.
3. The federated learning communication method according to claim 1, wherein: The federated learning communication method further includes: receiving the model loss sent by each of the first clients; Calculate the reduction in the model loss sent by the target client during the current round of federated aggregation compared to the model loss sent by the target client during the previous round of federated aggregation; According to the reduction amount corresponding to each first client, each client participating in the next round of federated aggregation is selected from each second client participating in horizontal federated learning; wherein, when the reduction amount corresponding to the target client is greater than the average value of the reduction amounts corresponding to each first client, the probability of the target client being selected to participate in the next round of federated aggregation is greater than the probability of the target client being selected to participate in the current round of federated aggregation; when the reduction amount corresponding to the target client is less than or equal to the average value, the probability of the target client being selected to participate in the next round of federated aggregation is less than the probability of the target client being selected to participate in the current round of federated aggregation.
4. The federated learning communication method according to claim 3, wherein: The step of selecting clients to participate in the next round of federated aggregation from the second clients participating in horizontal federated learning according to the reduction amount corresponding to each first client includes: When the reduction amount corresponding to the target client is greater than the average of the reduction amounts corresponding to the first clients, the α parameter value of the beta distribution corresponding to the target client is increased to update the beta distribution corresponding to the target client, wherein before the first round of federated aggregation, the same beta distribution is initialized for each of the second clients, and the initialization values of the α parameter and the β parameter in the beta distribution are equal; When the reduction amount corresponding to the target client is less than or equal to the average value, increasing the β parameter value of the beta distribution corresponding to the target client to update the beta distribution corresponding to the target client; After updating the beta distribution of each first client, a random number is generated using the beta distribution of each second client, and each second client is sorted in descending order according to the random number, and the second clients ranked in the front by a preset number are selected as the clients participating in the next round of federal aggregation.
5. The federated learning communication method according to claim 1, wherein: The quantization model parameters sent by the target client are obtained by quantizing the original model parameters of the target client using a quantizer of the target client and then encoding them. The step of using a dequantizer corresponding to the target client to dequantize the quantization model parameters corresponding to the target client to obtain the restored model parameters includes: The encoded quantization model parameters corresponding to the target client are decoded to obtain a decoding result, and then the decoding result is dequantized using a dequantizer corresponding to the target client to obtain a restored model parameter.
6. The federated learning communication method according to claim 5, wherein: The step of decoding the encoded quantization model parameters corresponding to the target client to obtain a decoding result includes: Taking each quantized value of the quantized model parameter before encoding corresponding to the target client as a signal source symbol, and obtaining the cumulative probabilities corresponding to the various signal source symbols when arranged in a preset order; The encoded quantization model parameters corresponding to the target client are used as a target decoding object, and a target source symbol is determined from various source symbols, so that the target decoding object is within an interval formed by a cumulative probability of the target source symbol and a cumulative probability of a next source symbol, wherein the next source symbol is a next source symbol among the target source symbols when the various source symbols are arranged in the preset order; Using the parameter value corresponding to the target signal source symbol as a parameter value obtained by decoding, subtracting the cumulative probability of the target signal source symbol from the target decoding object and dividing the result by the preset occurrence probability of the target signal source symbol, updating the calculation result as the target decoding object, and returning to the step of determining the target signal source symbol from the various signal source symbols; After the last parameter value is obtained through decoding, the quantization model parameters before encoding corresponding to the target client are obtained according to the decoded parameter values, and the quantization model parameters before encoding corresponding to the target client are used as the decoding result.
7. A federated learning communication method, characterized in that: The federated learning communication method is applied to any target client among the first clients participating in horizontal federated learning, and the federated learning communication method includes the following steps: In any round of federated aggregation in horizontal federated learning, the original model parameters are used to train the quantizer and dequantizer; quantizing the original model parameters using the quantizer to obtain quantized model parameters; The quantization model parameters and the dequantizer are sent to a server participating in horizontal federated learning, so that the server uses the dequantizer corresponding to the target client to dequantize the quantization model parameters corresponding to the target client to obtain the restored model parameters, and after obtaining the restored model parameters corresponding to each of the first clients, the restored model parameters of each of the first clients are aggregated to obtain the aggregated model parameters.
8. The federated learning communication method according to claim 7, wherein: The step of sending the quantization model parameters and the dequantizer to the server participating in horizontal federated learning includes: Taking each quantized value in the quantized model parameter as a signal source symbol, and calculating the cumulative probabilities corresponding to the various signal source symbols when arranged in a preset order; Initializing the coding interval, taking the first source symbol in the quantization model parameters as the target source symbol; Adding the product of the length of the coding interval and the cumulative probability of the next source symbol to the lower limit of the coding interval, and using the calculation result to update the upper limit of the coding interval, wherein the next source symbol is the next source symbol among the target source symbols when various source symbols are arranged in the preset order; Adding the product of the length of the coding interval and the cumulative probability of the target source symbol to the lower limit of the coding interval, and using the calculation result to update the lower limit of the coding interval; Updating the next signal source symbol of the target signal source symbol in the quantization model parameters to the target signal source symbol, and returning to the step of adding the lower limit of the coding interval to the product of the length of the coding interval and the cumulative probability of the next signal source symbol based on the updated lower limit and upper limit of the coding interval, and updating the upper limit of the coding interval using the calculation result; After obtaining an updated coding interval based on the last source symbol in the quantization model parameter, selecting a number from the coding interval, and obtaining a coding result of the quantization model parameter based on the selected number; The encoding result and the dequantizer are sent to a server participating in horizontal federated learning, so that the server can decode the encoding result to obtain the quantization model parameters.
9. A federated learning communication system, characterized in that: The federated learning communication system includes a horizontal federated learning server and multiple clients; The client is configured to use the original model parameters to train a quantizer and a dequantizer during any round of federated aggregation in horizontal federated learning; The client is further configured to quantize the original model parameters using the quantizer to obtain quantized model parameters; The client is further configured to send the quantization model parameters and the dequantizer to a server participating in horizontal federated learning; The server is configured to receive, during any round of federated aggregation in horizontal federated learning, quantized model parameters and dequantizers sent by each client participating in the current round of federated aggregation, wherein, for any target client among the clients, the quantized model parameters sent by the target client are obtained by the target client using the quantizer of the target client to quantize the original model parameters of the target client, and the quantizer and dequantizer of the target client are obtained by training the target client using the original model parameters of the target client; The server is further configured to use a dequantizer corresponding to the target client to dequantize the quantization model parameters corresponding to the target client to obtain restored model parameters; The server is further configured to, after obtaining the restoration model parameters corresponding to each client participating in the current round of federated aggregation, aggregate the restoration model parameters of each client to obtain an aggregate model parameter, and feed the aggregate model parameter back to each client; The client is further configured to receive aggregation model parameters fed back by the server, and use the aggregation model parameters to update the local federation model and then enter the next round of federation aggregation.
10. A federated learning communication device, characterized in that: The federated learning communication device includes: a memory, a processor, and a federated learning communication program stored in the memory and executable on the processor. When the federated learning communication program is executed by the processor, the steps of the federated learning communication method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Communication cost and model robustness optimization method based on multitask federated learning
CN114219094A
Transverse federal learning method and device and storage medium
CN114742240A