Efficient federated learning method based on adaptive quantization and dynamic gradient coding

By adopting adaptive quantization and dynamic gradient coding technology in federated learning, the problem of inefficient communication between edge terminal devices is solved, the model training process is optimized, and the system performance and scalability are improved.

CN120146225APending Publication Date: 2025-06-13NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510297264.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing federated learning technology in edge terminal devices has inefficient communication efficiency due to differences in computing power and network bandwidth, and the data heterogeneity between devices makes global model training complex, affecting model accuracy.

Method used

Using an efficient federated learning method of adaptive quantization and dynamic gradient coding, the client divides the gradient into important and non-important gradients through the gradient classification mechanism, and performs adaptive quantization and L1 norm sparse processing, and encodes and uploads them through the dynamic encoding protocol.

Benefits of technology

It significantly reduces communication overhead, optimizes the model training process, reduces model accuracy loss, and improves the overall performance and scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146225A_ABST
    Figure CN120146225A_ABST
Patent Text Reader

Abstract

According to the efficient federated learning method based on adaptive quantization and dynamic gradient coding provided by the invention, the gradient is classified, quantized and sparsified, so that the communication overhead is reduced, and the efficiency in the federated learning process is improved. The method comprises the following steps of: firstly, locally training a model by a client, calculating a gradient, dividing the gradient into an important gradient and a non-important gradient by utilizing a gradient classification mechanism, and adjusting the precision of the important gradient by adopting a self-adaptive quantization algorithm; and the non-important gradient is subjected to L1 norm sparsification processing. Then, the client uploads the processed gradient to a server through a dynamic coding protocol; after the server receives the gradients from the multiple clients, the gradients are decoded and aggregated by adopting an aggregation strategy, and the global model is updated; according to the method, the communication efficiency in federated learning is improved, the model training process is optimized, and the pressure of network bandwidth is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an efficient federated learning method based on adaptive quantization and dynamic gradient encoding, belonging to the technical field of federated learning. Background Art

[0002] In recent years, with the rapid development of artificial intelligence technology, data in various fields such as finance, manufacturing, and service industries have been widely used for artificial intelligence model training. Traditional artificial intelligence models need to centralize data to the server side for learning and training, which is prone to user privacy leakage. To solve this problem, Google proposed the concept of Federated Learning (FL) in 2016. Data holders perform model training locally and then upload the model training parameters to the server for aggregation and update, thus avoiding the leakage of private data. However, the edge terminal devices participating in federated learning are different. They have different software and hardware resources and network environments. A large number of mobile devices have limited battery power and network bandwidth, and the data held by different devices is not completely independent and identically distributed. The complex edge environment has caused a decline in the efficiency of the federated learning system. Since the existing network communication technology can gradually no longer meet the network requirements brought by the sharp increase in computing resources and data volume of edge terminal devices, communication has gradually become a bottleneck restricting the development of federated learning. Therefore, federated learning needs to optimize communication to promote the participation of edge terminal devices with limited communication and power resources in joint learning and contribute personal non-private data, which also has practical significance for the implementation of federated learning applications in complex edge environments.

[0003] Existing technologies face multiple technical challenges in the practical application of federated learning, especially in industries such as finance, manufacturing, service, and healthcare. First, the differences in computing power and network bandwidth of edge terminal devices lead to low communication efficiency. Especially in mobile devices and remote areas with limited bandwidth, the communication overhead is too large, affecting the training and update speed of the model. Second, the data heterogeneity between devices makes the training of the global model more complex, and it is impossible to achieve efficient parameter compression and quantization, thereby affecting the model accuracy. In addition, although the existing parameter compression technology can reduce the communication overhead, it usually adopts a unified compression strategy and cannot be dynamically adjusted according to the characteristics of different devices and data. Therefore, how to optimize communication and computing efficiency while ensuring privacy protection remains an important issue faced by federated learning technology.

[0004] At present, the communication optimization strategies of federated learning can be divided into optimization methods based on parameter compression, optimization methods for model update strategies, optimization methods for system architectures, and optimization methods for communication protocols. Among them, the most maturely developed is the optimization method based on parameter compression, which mainly includes techniques such as pruning, quantization, knowledge distillation, and low-rank decomposition. In recent years, researchers have been keen on dynamically adjusting parameter compression methods as federated learning progresses, so as to achieve the goal of maximizing model accuracy while reducing communication overhead. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies in the prior art and provide an efficient federated learning framework based on adaptive quantization and dynamic gradient encoding, which not only reduces the communication overhead in federated learning but also optimizes the model training process, thereby reducing the loss of training model accuracy.

[0006] After the client trains the model locally and calculates the gradients, it uses a gradient classification mechanism to divide the gradients into important gradients and unimportant gradients, and adopts an adaptive quantization algorithm to adjust the accuracy of the important gradients, while the unimportant gradients are sparsified by L1 norm processing. Then, the client encodes the processed gradients through a dynamic encoding protocol and uploads them to the server. After receiving the gradients from multiple clients, the server uses an aggregation strategy to decode, aggregate, and update the global model to complete the training iteration of federated learning.

[0007] An efficient federated learning method based on adaptive quantization and dynamic gradient encoding, the method comprising the following steps:

[0008] Step 1, the client receives the global model from the server and trains it based on the local dataset.

[0009] Step 2, the client locally trains the received global model based on the local dataset to obtain a local model.

[0010] Step 3, using a gradient classification mechanism to classify according to the magnitude and importance of the gradients, dividing the gradients into important gradients and unimportant gradients, and the process is as follows:

[0011] 3.1 Calculate the communication capabilities of each client, and its calculation expression is:

[0012]

[0013] where B i is the network bandwidth of client i, D i is the network latency of client i, and ∈ is a very small constant usually added to avoid errors caused by division by zero, representing the minimum limit of network latency.

[0014] 3.2 Calculate the weights of each client, and its calculation expression is:

[0015]

[0016] Among them, d i represents the size of the dataset of client i, and c i is the communication ability of client i, and K is the total number of clients.

[0017] 3.3 For each client i, sort the gradients according to their absolute values, and select the top k * important gradients, and the remaining are unimportant gradients. The number k * of important gradients is calculated as follows:

[0018]

[0019] Among them, T is the total communication budget set in advance, and ω i represents the weight of client i.

[0020] Compared with the prior art that uniformly operates on all gradients without distinction, this step can minimize the loss of model accuracy to the greatest extent, reduce redundant calculations, and improve the usability of the model.

[0021] Step 4, for important gradients, the client uses an adaptive quantization method to adjust the accuracy of the gradients, and the process is as follows:

[0022] 4.1 Calculate the contribution rate of each client, and its calculation expression is:

[0023]

[0024] Among them, N i is the number of training samples of client i, is the j-th gradient of client i.

[0025] 4.2 Dynamically adjust the quantization bit width according to the communication ability and contribution rate of the client, and its calculation formula is as follows:

[0026] Q i = round(α · C i + β · G i + γ)

[0027] Among them, α, β, and γ are custom adjustment factors, and G i is the contribution rate of client i.

[0028] 4.3 Quantize the important gradients according to the quantization bit width, and its calculation expression is:

[0029]

[0030] Among them, is the important gradient of client i, and Q i represents the quantization bit width of client i.

[0031] Compared with the prior art, the quantization bit width of this step is dynamically adjusted according to the changes in the client communication ability and contribution rate, which can effectively reduce the communication overhead, adapt to the changes in the client and network environment, and find the optimal balance between accuracy and efficiency. This makes the training process more efficient and robust in large-scale distributed or federated learning scenarios, especially in the case of limited resources.

[0032] Step 5, for unimportant gradients, use L1 norm sparsification, and its calculation expression is:

[0033]

[0034] Among them, is the unimportant gradient of client i, and λ is the threshold of L1 regularization.

[0035] Step 6, the client encodes the gradients after quantization and sparsification through the dynamic gradient encoding protocol, and the process is as follows:

[0036] 6.1 The client combines the gradients after quantization and sparsification: Among them, is the quantized important gradient, is the quantized unimportant gradient.

[0037] 6.2 Calculate the displacement factor according to the preset network bandwidth threshold, and its calculation expression is:

[0038]

[0039] Among them, T max is the preset threshold, representing the maximum network bandwidth that can be accepted during the communication process, is the combined gradient.

[0040] 6.3 Displace the combined gradient according to the displacement factor, and its calculation expression is:

[0041]

[0042] Among them, s represents the displacement factor.

[0043] 6.4 Concatenate the sign bit of the combined gradient, the displacement factor s, and the absolute value of the shifted gradient value into a large integer, and its calculation formula is as follows:

[0044]

[0045] Among them, the symbol represents the gradient, that is, whether it is a positive number or a negative number, and & represents the bitwise AND concatenation operation.

[0046] Step 7, each client uploads the encoded large integer to the server.

[0047] Step 8, the server decodes the large integers uploaded by each client, and its calculation expression is:

[0048]

[0049] Step 9, after the server decodes to obtain the gradient, it aggregates the gradient, obtains the global model and returns it to the client, completing the iteration of the current training round.

[0050] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the efficient federated learning method based on adaptive quantization and dynamic gradient encoding.

[0051] A computer-readable storage medium stores computer instructions thereon. When the computer instructions are executed by a processor, they implement the efficient federated learning method based on adaptive quantization and dynamic gradient encoding.

[0052] Compared with the prior art, the present invention has the following advantages:

[0053] After the client completes model training, the present invention uses a gradient classification mechanism to divide the gradient into important gradients and unimportant gradients, and respectively adopts adaptive quantization and L1-norm sparsification methods for processing. The important gradients dynamically adjust the quantization bit width through the adaptive quantization algorithm, effectively reducing the communication volume while ensuring the gradient accuracy; the unimportant gradients reduce redundant information through L1-norm sparsification, thereby further reducing the data volume of gradient upload. These processing measures not only significantly reduce the communication overhead, but also ensure that the pressure on the network bandwidth is reduced while ensuring the model accuracy by adopting a differential compression strategy for different gradients. The client encodes the processed gradient through a dynamic encoding protocol and uploads it to the server, further optimizing the data transmission efficiency. Generally speaking, the present invention not only improves the communication efficiency in federated learning, optimizes the model training process, but also effectively reduces the pressure on the network bandwidth, improving the overall performance and scalability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is a framework diagram of an efficient federated learning framework based on adaptive quantization and dynamic gradient encoding of the present invention;

[0055] Figure 2It is a flowchart for the client in the present invention to perform gradient partitioning and compress it using different strategies;

[0056] Figure 3 It is a flowchart for the client in the present invention to process the compressed gradient and encode it using a dynamic coding protocol. Detailed implementation manners

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0058] Embodiment: An efficient federated learning method based on adaptive quantization and dynamic gradient coding, the method includes the following steps:

[0059] Step 1, the client receives the global model from the server and trains it based on the local dataset,

[0060] Step 2, the client performs local training on the received global model based on the local dataset to obtain a local model,

[0061] Step 3, adopt a gradient classification mechanism to classify according to the magnitude and importance of the gradient, and divide the gradient into important gradients and unimportant gradients. See Figure 2 , and the process is as follows:

[0062] 3.1 Calculate the communication capabilities of each client, and its calculation expression is:

[0063]

[0064] Among them, B i is the network bandwidth of client i, D i is the network latency of client i. To avoid errors caused by dividing by zero, a very small constant is usually added, representing the minimum limit of network latency.

[0065] 3.2 Calculate the weights of each client, and its calculation expression is:

[0066]

[0067] Among them, d i represents the size of the dataset of client i, c i is the communication capability of client i, and K is the total number of clients.

[0068] 3.3 For each client i, sort the gradients according to their absolute values, and select the top k * important gradients, and the remaining are unimportant gradients. The number k * of important gradients is calculated as follows:

[0069]

[0070] where T is the total communication budget set in advance, and ω i represents the weight of client i.

[0071] Step 4. For important gradients, the client uses an adaptive quantization method to adjust the precision of the gradients, and the process is as follows:

[0072] 4.1 Calculate the contribution rate of each client, and its calculation expression is:

[0073]

[0074] where N i is the number of training samples of client i, is the j-th gradient of client i.

[0075] 4.2 Dynamically adjust the quantization bit width according to the communication ability and contribution rate of the client, and its calculation formula is as follows:

[0076] Q i = round(α·C i + β·G i + γ)

[0077] where α, β, and γ are self-defined adjustment factors, and G i is the contribution rate of client i.

[0078] 4.3 Quantize the important gradients according to the quantization bit width, and its calculation expression is:

[0079]

[0080] where is the important gradient of client i, and Q i represents the quantization bit width of client i.

[0081] Step 5. For unimportant gradients, use L1 norm sparsification, and its calculation expression is:

[0082]

[0083] where is the unimportant gradient of client i, and λ is the threshold of L1 regularization.

[0084] Step 6: The client encodes the quantized and sparsified gradients through the dynamic gradient encoding protocol. See Figure 3 , and the process is as follows:

[0085] 6.1 The client combines the quantized and sparsified gradients:

[0086] Among them, is the quantized important gradient, is the quantized unimportant gradient.

[0087] 6.2 Calculate the displacement factor according to the preset network bandwidth threshold. The calculation expression is:

[0088]

[0089] Among them, T max is the preset threshold, representing the maximum network bandwidth acceptable during the communication process, is the combined gradient.

[0090] 6.3 Displace the combined gradient according to the displacement factor. The calculation expression is:

[0091]

[0092] Among them, s represents the displacement factor.

[0093] 6.4 Concatenate the sign bit of the combined gradient, the displacement factor s, and the absolute value of the shifted gradient value into a large integer. The calculation formula is as follows:

[0094]

[0095] Among them, represents the sign of the gradient, that is, whether it is a positive number or a negative number, and & represents the bitwise AND concatenation operation.

[0096] Step 7: Each client uploads the encoded large integer to the server.

[0097] Step 8: The server decodes the large integers uploaded by each client. The calculation expression is:

[0098]

[0099] Step 9: After the server decodes to obtain the gradients, it aggregates them to obtain the global model and returns it to the client, completing the iteration of the current training round.

[0100] It should be noted that the above embodiments are not intended to limit the protection scope of the present invention. Equivalent transformations or substitutions made on the basis of the above technical solutions all fall within the protection scope of the claims of the present invention.

Claims

1. An efficient federated learning method based on adaptive quantization and dynamic gradient coding, characterized by: The method comprises the following steps: Step 1: The client receives the global model from the server. Step 2: Based on the local data set, the client ensures that the data has been preprocessed and split into batches suitable for training, inputs a batch of local data into the model for forward propagation, and calculates the prediction results. Compare the difference between the predicted result and the true label, calculate the loss value, calculate the gradient of each model parameter with respect to the loss function through the back-propagation algorithm, and use the selected optimization algorithm (such as SGD) Update the model parameters according to the calculated gradients, and repeat the training until the predetermined training round is reached or the loss is no longer significantly reduced. Step 3: The client uses a gradient classification mechanism to classify the local gradients obtained according to their size and importance, and divides the gradients into important gradients and unimportant gradients. Step 4: For important gradients, the client uses an adaptive quantization method to adjust the accuracy of the gradient. Step 5: For non-significant gradients, L1 norm sparsification is used. Step 6: The client encodes the quantized and sparsified gradients using the dynamic gradient coding protocol. Step 7: Each client uploads the encoded large integer to the server. Step 8: The server decodes the large integer uploaded by each client. Step 9: The server decodes and obtains the gradients, aggregates them, obtains the global model and returns it to the client. Complete the iteration of the current training round.

2. The efficient federated learning method based on adaptive quantization and dynamic gradient coding according to claim 1, characterized in that: Step 3: Use the gradient classification mechanism to classify the gradients according to their size and importance, and divide the gradients into important gradients and unimportant gradients. The process is as follows: 3.1 Calculate the communication capability of each client. The calculation expression is: Among them, B i is the network bandwidth of client i, D i is the network delay of client i. In order to avoid errors caused by division by zero, a very small constant is usually added. ∈ represents the minimum limit of network delay. 3.2 Calculate the weight of each client, and the calculation expression is: Among them, d i represents the dataset size of client i, c i is the communication capability of client i, K is the total number of clients, 3.3 For each client i, sort the gradients by absolute value and select the top k * An important gradient, The remaining are non-important gradients, and the number of important gradients k * The calculation formula is as follows: Among them, T is the total communication budget set in advance, ω i represents the weight of client i.

3. The efficient federated learning method based on adaptive quantization and dynamic gradient coding according to claim 1, characterized in that: Step 4: For important gradients, the client uses an adaptive quantization method to adjust the accuracy of the gradient. The process is as follows: 4.1 Calculate the contribution rate of each client, the calculation expression is: Among them, N i is the number of training samples of client i, is the j-th gradient of client i, 4.2 Dynamically adjust the quantization bit width according to the client's communication capability and contribution rate. The calculation formula is as follows: Q i =round(α·C i +β·G i +c) in, α, β and γ are custom adjustment factors, G i The contribution rate of client i, 4.3 Quantize the important gradient according to the quantization bit width, and its calculation expression is: in, is the important gradient of client i, Q i Indicates the quantization bit width of client i.

4. The efficient federated learning method based on adaptive quantization and dynamic gradient coding according to claim 1, characterized in that: Step 5: For non-significant gradients, L1 norm sparsification is used, and its calculation expression is: in, is the non-important gradient of client i, and λ is the threshold of L1 regularization.

5. The efficient federated learning method based on adaptive quantization and dynamic gradient coding according to claim 1, characterized in that: Step 6: The client encodes the quantized and sparse gradients through the dynamic gradient encoding protocol. The process is as follows: 6.1 The client combines the quantized and sparsified gradients: in, is the quantized important gradient, is the quantized non-significant gradient, 6.2 Calculate the displacement factor according to the preset network bandwidth threshold, and its calculation expression is: in, T max It is a preset threshold, indicating the maximum value of the network bandwidth that can be accepted during the communication process. is the combined gradient, 6.3 The combined gradient is displaced according to the displacement factor, and its calculation expression is: Where s is the displacement factor, 6.4 The sign bit of the combined gradient, the shift factor s, and the absolute value of the shifted gradient value are concatenated into a large integer, and the calculation formula is as follows: in, Denoted as the combined gradient later, Indicates the sign of the gradient, that is, whether it is a positive or negative number, and & indicates a bitwise AND concatenation operation.

6. The efficient federated learning method based on adaptive quantization and dynamic gradient coding according to claim 1, characterized in that: Step 8: The server decodes the large integer uploaded by each client, and the calculation expression is:

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the efficient federated learning method based on adaptive quantization and dynamic gradient coding as described in any one of claims 1 to 6 above is implemented.

8. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instruction is executed by the processor, an efficient federated learning method based on adaptive quantization and dynamic gradient coding as described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Gradient compressor based on information consistency driving and gradient compression method and equipment

    CN120911525A

  • Gradient compressor and gradient compression method and device based on information consistency driving

    CN120911525B

  • Self-adaption-based layered compression federated learning method and device, and medium

    CN121352073A