An adaptive gradient compression and asynchronous aggregation federated learning method
Patent Information
- Application Number
- CN202610838554.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-09-29
AI Technical Summary
[0007]为解决现有技术异步联邦学习中聚合间隔与梯度压缩率无法协同优化、梯度陈旧度高、异构客户端适配性差、资源约束下训练效率低的技术问题,本发明提供了一种自适应梯度压缩与异步聚合联邦学习方法,本发明采用的技术方案是:
本发明通过将聚合间隔与客户端梯度压缩率作为耦合变量,构建以最大化低陈旧度有效梯度信息总量为目标的非线性规划问题,突破了单独优化或简单组合配置的局限,实现训练效率与通信开销之间的最优权衡。本发明采用信赖域牛顿法对联合优化问题进行稳定迭代求解,保证算法在有限步内快速收敛至满足精度要求的全局最优解,避免陷入局部最优。本发明采用双曲正切函数对符号函数等不连续、不可微的判断逻辑进行连续可导近似,将原始非光滑优化问题转化为可稳定求解的标准形式,确保优化算法正常运行。本发明以低频更新全局聚合间隔保证训练稳定性,以高频更新客户端梯度压缩率快速适配设备状态变化,降低计算复杂度的同时提升系统对异构环境的适应性,避免模型震荡。本发明根据梯度上传延迟计算陈旧度,对超时梯度权重置零,对有效梯度赋予与陈旧度成反比的平方根衰减权重,有效抑制陈旧更新对模型精度的破坏,加速全局模型收敛。本发明以客户端本地计算时间和通信时间为聚类特征,将性能相近的客户端归为同一簇,同一簇内共享相同的压缩率优化变量,将优化变量维度从客户端数量降低至聚类数量,避免维度灾难,支持大规模异构设备高效扩展。本发明根据求解得到的最优压缩率,按公式计算保留梯度参数数量,执行绝对值最大的梯度稀疏化压缩,在保证有效梯度信息的前提下显著减少单轮通信数据量,适配窄带边缘网络。
Smart Images

Figure CN122840176A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to an adaptive gradient compression and asynchronous aggregated federated learning method. Background Technology
[0002] Federated learning, a privacy-preserving distributed machine learning framework, consists of a central server and multiple distributed clients. The server is responsible for aggregating and updating the global model, while clients train the model locally using private data. They do not need to upload the original data, only transmitting model gradients or parameter updates, thus ensuring user privacy from the data source. Therefore, it is widely used in scenarios with high data privacy requirements, such as the Internet of Things (IoT), mobile internet, and industrial internet. In practical deployments, federated learning systems generally face challenges such as limited computing power of client devices, unstable network bandwidth, and strained communication resources. Especially in edge computing scenarios, uploading complete gradients by clients incurs huge communication overhead, easily causing network congestion, increased training latency, slow model convergence, and even training failure due to communication interruptions. To alleviate this problem, gradient compression technology has been introduced into the federated learning system. Through sparsity, quantization, and low-rank decomposition, the amount of gradient data uploaded by clients is reduced. Among these methods, Top-K sparsity is the most commonly used gradient compression method in federated learning due to its simplicity, stable compression effect, and minimal loss of model accuracy, effectively reducing communication burden and improving overall training efficiency.
[0003] Asynchronous federated learning is an improvement on traditional synchronous federated learning, which requires waiting for all clients to complete uploading, resulting in high overall training latency. Instead of waiting for all clients to upload, the server actively triggers global model aggregation at preset time intervals, effectively reducing synchronous waiting time. This is more suitable for real-world applications with uneven computing power and unstable network conditions on edge devices. However, the asynchronous mechanism also brings a series of core technical challenges, such as high model staleness, low communication efficiency, and poor client heterogeneity adaptability. Combining gradient compression with asynchronous federated learning can simultaneously achieve communication optimization and reduced training latency, but the coordination and adaptation issues between the two become prominent. The aggregation interval determines the server's gradient collection rhythm, while the gradient compression rate determines the amount of data uploaded by clients and communication time. These two are strongly coupled, jointly affecting system training efficiency, communication overhead, and model convergence performance. Therefore, how to reasonably configure the aggregation interval and compression rate becomes the key to communication optimization and training acceleration in asynchronous federated learning.
[0004] Currently, various optimization methods for asynchronous federated learning have been proposed in the industry. The existing technologies closest to this invention mainly revolve around adaptive aggregation intervals, adaptive gradient compression, and limited collaborative strategies. One typical approach focuses on the adaptive adjustment of the asynchronous aggregation interval. By monitoring the client's upload speed, gradient staleness changes, and model convergence trends in real time, it dynamically lengthens or shortens the server's aggregation time window. Its core objective is to reduce the impact of stale updates on model accuracy and accelerate model convergence. However, this type of method only treats gradient compression as a fixed pre-operation and does not consider the impact of compression ratio changes on transmission time, upload success rate, and effective gradient information. It cannot match the optimal aggregation rhythm according to the client's actual communication capabilities, resulting in a large number of gradient timeouts or resource idleness in highly heterogeneous scenarios. Another mainstream technology focuses on gradient compression adaptive optimization in federated learning. Based on the client's local gradient distribution, device computing power, and uplink bandwidth, it dynamically allocates different sparsity ratios or quantization bits to minimize the amount of communication data while ensuring the integrity of gradient information. This type of method can significantly reduce the communication overhead per round, but it is completely unaware of the server-side aggregation interval setting. The compression strategy and aggregation rhythm are independent of each other. When the compression ratio is too high, resulting in slower transmission, or the compression ratio is too low, resulting in wasted traffic, it cannot be compensated by the aggregation interval. It is also difficult to achieve overall optimization under the dual constraints of global training time and communication budget. Some improvement schemes attempt to simply combine and configure the aggregation interval and compression strategy, usually using a fixed interval with a fixed compression ratio, or making segmented adjustments based on empirical rules. Although it can achieve certain results in small-scale homogeneous environments, it does not establish a coupling relationship model between the two, and does not form a unified optimization goal and constraint system. It cannot be personalized for heterogeneous clients, nor can it handle the non-convex and non-smooth characteristics in optimization problems. Its stability and performance are significantly insufficient in large-scale edge deployment scenarios.
[0005] Overall, existing asynchronous federated learning optimization methods suffer from several significant drawbacks. First, they generally ignore the strong coupling between the aggregation margin and the gradient compression rate. These two factors jointly determine the system's training efficiency and communication overhead. Optimizing either one individually or in simple combinations fails to achieve the optimal trade-off between training efficiency and communication overhead, often resulting in an imbalance where training speed increases but communication costs rise dramatically, or communication overhead decreases but model accuracy drops significantly. Second, jointly optimizing the aggregation margin and gradient compression rate creates a complex, high-dimensional, non-convex, and non-smooth optimization problem. Traditional optimization techniques such as gradient descent, greedy algorithms, and heuristics struggle to solve this efficiently. These methods either exhibit poor convergence for non-convex problems, are prone to getting trapped in local optima, or cannot handle the non-differentiability of sign functions and piecewise constraints, ultimately failing to yield a stable and reliable parameter configuration. Furthermore, optimizing the compression rate independently for each client causes the dimensionality of the optimization variable to grow linearly with the number of clients, leading to the curse of dimensionality. In large-scale client deployments, the computational overhead becomes unbearable, severely impacting system scalability. Furthermore, most existing technologies lack a gradient weight decay mechanism based on staleness, making it impossible to effectively distinguish the contributions of fresh and old gradients. This results in a large number of low-value, outdated updates participating in model updates, significantly slowing down model convergence and reducing final test accuracy. Moreover, existing solutions generally lack the ability to manage both global training time and communication traffic budget constraints. This either exceeds resource limits, preventing training from being implemented, or leads to low resource utilization and wasted computing power and bandwidth. Additionally, it is difficult to provide personalized, fine-grained compression strategies for clients with highly heterogeneous computing capabilities and network conditions, making it difficult for the system to operate stably and efficiently in real edge environments.
[0006] Chinese invention application No. 202311529748.6 discloses "A Secure Aggregation Method for Asynchronous Federated Learning Based on Error-Based Learning," which includes: a client downloading a global model for local training, randomly selecting secret vectors and noise, masking the locally trained gradients using the secret vectors and noise, and uploading the masked gradients; a server receiving the masked gradients uploaded by the clients, placing them into a buffer K, and when K is full, preparing to aggregate the secret vectors of the clients in K, and broadcasting the data to the clients in K; and the clients in K performing secure aggregation to obtain the aggregated secret vector s. sum The server then transmits the set of clients participating in the secure aggregation process to the server; the server sums the mask gradients uploaded by the clients in the client set and uses s sum Remove the mask of the mask gradient to obtain the updated gradient of the global model. Summary of the Invention
[0007] To address the technical problems in existing asynchronous federated learning technologies, such as the inability to coordinate the optimization of aggregation interval and gradient compression rate, high gradient staleness, poor adaptability to heterogeneous clients, and low training efficiency under resource constraints, this invention provides an adaptive gradient compression and asynchronous aggregation federated learning method. The technical solution adopted by this invention is as follows: The first aspect of this invention provides an adaptive gradient compression and asynchronous aggregation federated learning method, comprising: Pre-training optimization phase: The central server collects device status information from each client, clusters the clients based on the device status information, establishes a joint optimization problem for each client class, and finally solves the joint optimization problem to obtain the optimal global aggregation interval and the optimal gradient compression rate for each client class. Collaborative training phase: The server asynchronously triggers global model aggregation according to the optimal global aggregation interval; after receiving the global model from the server, each client performs local training using its local private data, compresses the calculated gradients according to the optimal gradient compression rate corresponding to its category, and uploads the compressed gradients to the server; when triggering aggregation, the server assigns weights to each gradient according to the staleness of the gradients uploaded by each client, and performs weighted aggregation of all received valid gradients based on the weights to update the global model.
[0008] As a preferred embodiment, the pre-training optimization phase specifically includes: S1: Each client performs a benchmark test, obtains the pure computation time required to complete one local training session and the communication time required to upload one complete gradient, and reports the device status information containing the pure computation time and communication time to the central server. S2: Based on the device status information reported by all clients, the server performs a clustering operation on all clients using the pure computing time and communication time as clustering features, and groups clients with similar performance into the same category; S3: The server aims to maximize the total amount of effective gradient information with low staleness, and sets at least a global training time limit and a total communication traffic budget as constraints to establish a joint optimization problem for each type of client. S4: The server uses the trust region Newton method to solve the joint optimization problem, and obtains the optimal global aggregation interval and the optimal gradient compression rate for each type of client. S5: The server configures the optimal global aggregation interval in its own aggregation module, and sends the optimal gradient compression rate and its initial global model corresponding to each type of client to that type of client.
[0009] As a preferred embodiment, the collaborative training phase specifically includes: Step T1: The server starts the global clock and periodically triggers global aggregation events according to the optimal global aggregation interval; Step T2: After receiving the global model from the server, each client uses its local private data to perform a specified number of rounds of local training and calculates the original gradient. Step T3: Each client performs Top-K sparsity compression on the original gradient according to the received optimal gradient compression rate, retaining only the K gradient parameters with the largest absolute values, and uploads the compressed gradient to the server; Step T4: Each time global aggregation is triggered, the server calculates the staleness of the gradient based on the difference between the gradient uploaded by each client and the round in which it successfully participated in the aggregation last time, and assigns a weight inversely proportional to the staleness to each gradient. Step T5: The server determines whether the total time taken for the client from the most recent successful upload to the current time exceeds the sum of the preset aggregation intervals. If so, the weight of the gradient is reset to zero, and it is determined to be an invalid gradient. Step T6: The server performs a weighted summation of the valid gradients with non-zero weights to complete the aggregation and update of the global model, and then sends the updated global model to the clients participating in this aggregation. Step T7: Repeat steps T2 to T6 until the global training time limit is reached or the model converges.
[0010] As a preferred embodiment, the expression for the total amount of effective gradient information with low obsolescence is:
[0011] in, The total amount of effective gradient information with low obsolescence The aggregate weight of the i-th client in round t. For gradient compression ratio, This represents the original dimension of the gradient.
[0012] As a preferred embodiment, the server assigns differentiated aggregation weights to different clients based on the staleness of the gradient; the lower the staleness of the gradient, the larger the aggregation weight. If gradient transmission times out, the aggregation weight is reset to zero, and the gradient does not participate in global aggregation.
[0013] in, From the last successful transmission by the client to the [number]th round t Total aggregation time of the wheel, For the first i The client in the first tBefore the global aggregation round, the last global round in which gradient transfer was successfully completed, initialized to a value of ,Right now: .
[0014] As a preferred embodiment, the joint optimization problem aims to maximize the total amount of effective gradient information with low staleness, and includes at least global training time constraints, communication traffic resource constraints, convergence constraints, and variable range constraints. The global training time constraint is defined as follows: the sum of the aggregation intervals of all rounds does not exceed the preset total global training time. T ,Right now:
[0015] The communication traffic resource constraint is defined as follows: Let b be the traffic consumption for transmitting a complete gradient. Then, the total traffic consumption of effective gradients across all rounds shall not exceed the preset traffic resource budget B, i.e.:
[0016] in, This represents the bandwidth consumption for transmitting the compressed gradient. For symbolic functions, express , No. t The gradient transfer is completed in one round; express , No. t Gradient transfer was not completed in the first round; The convergence constraint is defined as follows: the training loss of the global model needs to converge to a preset threshold, let the convergence threshold be... ( ),but:
[0017] in, Let be the loss value of the global model in round t. This represents the loss value of the optimal model. The variable range constraint is defined as follows: the aggregation interval and compression ratio are positive real numbers, the aggregation interval is within a preset range, and the compression ratio is between 0 and 1, that is:
[0018] in, and These are the minimum and maximum values of the aggregation interval, respectively; The complete formal representation of the joint optimization problem is as follows: .
[0019] As a preferred embodiment, the method for solving the joint optimization problem includes: A smooth approximation strategy is adopted to replace the discontinuous and non-differentiable judgment logic in the joint optimization problem with a continuously differentiable hyperbolic tangent function. The expression of the hyperbolic tangent function is as follows: tanh (γ x) Where γ is the steepness coefficient; A multi-timescale decomposition strategy is adopted to update the global aggregation interval at a first preset frequency and the gradient compression rate at a second preset frequency, wherein the first preset frequency is lower than the second preset frequency. The trust region Newton method is used to iteratively solve the optimization problem after smooth approximation and multi-time scale decomposition until the preset convergence condition is met. As a preferred embodiment, the trust region radius Δ in the trust region Newton method is dynamically updated according to the following rules: Let the ratio be ,but: like Then the radius is increased as follows: Δ = min (2 Δ, Δ_max); like Then, the radius is reduced as follows: Δ = max(0.25) Δ, Δ_min); like If so, then the current radius remains unchanged; Where Δ_max = 10.0 is the upper limit of the radius, Δ_min = 1e-4 is the lower limit of the radius, and the initial radius Δ0 = 1.0.
[0020] As a preferred embodiment, the preset convergence condition includes any one of the following: Parameter variation convergence: The maximum absolute difference between the aggregation interval and the compression ratio obtained from two adjacent iterations is less than 1e-5; Objective function convergence: The relative change in the total amount of low-stale effective gradient information between two adjacent iterations is less than 1e-6; Maximum iteration count convergence: When the number of iterations reaches the preset upper limit of 200 rounds, the iteration is forcibly stopped.
[0021] A second aspect of the present invention provides a computer device including a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor, wherein the computer program, when executed by the processor, implements the steps of the aforementioned adaptive gradient compression and asynchronous aggregate federated learning method.
[0022] Compared with the prior art, the beneficial effects of this invention are: This invention constructs a nonlinear programming problem with the goal of maximizing the total amount of effective gradient information with low staleness by using the aggregation interval and client gradient compression rate as coupling variables. This overcomes the limitations of individual optimization or simple combination configurations, achieving an optimal trade-off between training efficiency and communication overhead. The invention employs the trust region Newton method to stably iteratively solve the joint optimization problem, ensuring that the algorithm converges quickly to the global optimum that meets the accuracy requirements within a finite number of steps, avoiding getting trapped in local optima. The invention uses the hyperbolic tangent function to make the judgment logic, such as the sign function, continuously differentiable, transforming the original non-smooth optimization problem into a stable, solvable standard form, ensuring the normal operation of the optimization algorithm. This invention uses low-frequency updates to the global aggregation interval to ensure training stability and high-frequency updates to the client gradient compression rate to quickly adapt to changes in device state, reducing computational complexity while improving the system's adaptability to heterogeneous environments and avoiding model oscillations. This invention calculates staleness based on gradient upload delay, resets the weights of timed-out gradients to zero, and assigns square-root decay weights inversely proportional to staleness to effective gradients, effectively suppressing the damage of stale updates to model accuracy and accelerating global model convergence. This invention uses client-side local computation time and communication time as clustering features to group clients with similar performance into the same cluster. Within the same cluster, clients share the same compression ratio optimization variable, reducing the dimensionality of the optimization variable from the number of clients to the number of clusters, thus avoiding the curse of dimensionality and supporting efficient scaling on large-scale heterogeneous devices. Based on the optimal compression ratio obtained from the solution, this invention calculates the number of gradient parameters to retain according to a formula and performs gradient sparsity compression with the largest absolute value. This significantly reduces the amount of data per round of communication while ensuring effective gradient information, making it suitable for narrowband edge networks. Attached Figure Description
[0023] Figure 1 This embodiment provides a flowchart of an adaptive gradient compression and asynchronous aggregation federated learning method; Figure 2 This is a schematic diagram of the pre-training optimization stage process provided in this embodiment; Figure 3 This is a schematic diagram of the collaborative training phase process provided in this embodiment. Detailed Implementation
[0024] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the invention. It should be understood that the described embodiments are merely some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of the embodiments of this application.
[0025] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the embodiments of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0026] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims. In the description of this application, it should be understood that the terms "first," "second," "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0027] Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship. The invention will be further described below with reference to the accompanying drawings and embodiments.
[0028] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0029] Example 1 Please refer to Figure 1 This embodiment provides an adaptive gradient compression and asynchronous aggregation federated learning method, including: Step A1: Pre-training optimization stage: The central server collects device status information from each client, clusters the clients based on the device status information, establishes a joint optimization problem for each client class, and finally solves the joint optimization problem to obtain the optimal global aggregation interval and the optimal gradient compression rate for each client class. In one specific embodiment, please refer to Figure 2 The pre-training optimization phase specifically includes: S1: Each client performs a benchmark test, obtains the pure computation time required to complete one local training session and the communication time required to upload one complete gradient, and reports the device status information containing the pure computation time and communication time to the central server. S2: Based on the device status information reported by all clients, the server performs a clustering operation on all clients using the pure computing time and communication time as clustering features, and groups clients with similar performance into the same category; S3: The server aims to maximize the total amount of effective gradient information with low staleness, and sets at least a global training time limit and a total communication traffic budget as constraints to establish a joint optimization problem for each type of client. In a specific embodiment, the expression for the total amount of low-staleness effective gradient information is:
[0030] in, The total amount of effective gradient information with low obsolescence The aggregate weight of the i-th client in round t. For gradient compression ratio, This represents the original dimension of the gradient.
[0031] In one specific embodiment, the server assigns differentiated aggregation weights to different clients based on the staleness of the gradient; the lower the staleness of the gradient, the larger the aggregation weight. If gradient transmission times out, the aggregation weight is reset to zero, and the gradient does not participate in global aggregation.
[0032] in, From the last successful transmission by the client to the [number]th round t Total aggregation time of the wheel, For the first i The client in the first t Before the global aggregation round, the last global round in which gradient transfer was successfully completed, initialized to a value of ,Right now: .
[0033] In one specific embodiment, the joint optimization problem aims to maximize the total amount of effective gradient information with low staleness, and includes at least global training time constraints, communication traffic resource constraints, convergence constraints, and variable range constraints. The global training time constraint is defined as follows: the sum of the aggregation intervals of all rounds does not exceed the preset total global training time. T ,Right now:
[0034] The communication traffic resource constraint is defined as follows: Let b be the traffic consumption for transmitting a complete gradient. Then, the total traffic consumption of effective gradients across all rounds shall not exceed the preset traffic resource budget B, i.e.:
[0035] in, This represents the bandwidth consumption for transmitting the compressed gradient. For symbolic functions, express , No. t The gradient transfer is completed in one round; express , No. t Gradient transfer was not completed in the first round; The convergence constraint is defined as follows: the training loss of the global model needs to converge to a preset threshold, let the convergence threshold be... ( ),but:
[0036] in, Let be the loss value of the global model in round t. This represents the loss value of the optimal model. The variable range constraint is defined as follows: the aggregation interval and compression ratio are positive real numbers, the aggregation interval is within a preset range, and the compression ratio is between 0 and 1, that is:
[0037] in, and These are the minimum and maximum values of the aggregation interval, respectively; The complete formal representation of the joint optimization problem is as follows: .
[0038] S4: The server uses the trust region Newton method to solve the joint optimization problem, and obtains the optimal global aggregation interval and the optimal gradient compression rate for each type of client. In one specific embodiment, the method for solving the joint optimization problem includes: A smooth approximation strategy is adopted to replace the discontinuous and non-differentiable judgment logic in the joint optimization problem with a continuously differentiable hyperbolic tangent function. The expression of the hyperbolic tangent function is as follows: tanh (γ x) Where γ is the steepness coefficient; A multi-timescale decomposition strategy is adopted to update the global aggregation interval at a first preset frequency and the gradient compression rate at a second preset frequency, wherein the first preset frequency is lower than the second preset frequency. The trust region Newton method is used to iteratively solve the optimization problem after smooth approximation and multi-time scale decomposition until the preset convergence condition is met.
[0039] In a specific embodiment, the trust region radius Δ in the trust region Newton method is dynamically updated according to the following rules: Let the ratio be ,but: like Then the radius is increased as follows: Δ = min (2 Δ, Δ_max); like Then, the radius is reduced as follows: Δ = max(0.25) Δ, Δ_min); like If so, then the current radius remains unchanged; Where Δ_max = 10.0 is the upper limit of the radius, Δ_min = 1e-4 is the lower limit of the radius, and the initial radius Δ0 = 1.0.
[0040] In one specific embodiment, the preset convergence condition includes any one of the following: Parameter variation convergence: The maximum absolute difference between the aggregation interval and the compression ratio obtained from two adjacent iterations is less than 1e-5; Objective function convergence: The relative change in the total amount of low-stale effective gradient information between two adjacent iterations is less than 1e-6; Maximum iteration count convergence: When the number of iterations reaches the preset upper limit of 200 rounds, the iteration is forcibly stopped.
[0041] S5: The server configures the optimal global aggregation interval in its own aggregation module, and sends the optimal gradient compression rate and its initial global model corresponding to each type of client to that type of client.
[0042] Step A2: Collaborative Training Phase: The server asynchronously triggers global model aggregation according to the optimal global aggregation interval; after receiving the global model from the server, each client performs local training using its local private data, compresses the calculated gradients according to the optimal gradient compression rate corresponding to its category, and uploads the compressed gradients to the server; when triggering aggregation, the server assigns weights to each gradient according to the staleness of the gradients uploaded by each client, and performs weighted aggregation of all received valid gradients based on the weights to update the global model; In one specific embodiment, please refer to Figure 3 The collaborative training phase specifically includes: Step T1: The server starts the global clock and periodically triggers global aggregation events according to the optimal global aggregation interval; Step T2: After receiving the global model from the server, each client uses its local private data to perform a specified number of rounds of local training and calculates the original gradient. Step T3: Each client performs Top-K sparsity compression on the original gradient according to the received optimal gradient compression rate, retaining only the K gradient parameters with the largest absolute values, and uploads the compressed gradient to the server; Step T4: Each time global aggregation is triggered, the server calculates the staleness of the gradient based on the difference between the gradient uploaded by each client and the round in which it successfully participated in the aggregation last time, and assigns a weight inversely proportional to the staleness to each gradient. Step T5: The server determines whether the total time taken for the client from the most recent successful upload to the current time exceeds the sum of the preset aggregation intervals. If so, the weight of the gradient is reset to zero, and it is determined to be an invalid gradient. Step T6: The server performs a weighted summation of the valid gradients with non-zero weights to complete the aggregation and update of the global model, and then sends the updated global model to the clients participating in this aggregation. Step T7: Repeat steps T2 to T6 until the global training time limit is reached or the model converges.
[0043] Example 2 Please refer to Figure 1 This embodiment provides an adaptive gradient compression and asynchronous aggregation federated learning method, including: Step A1: Pre-training optimization stage: The central server collects device status information from each client, clusters the clients based on the device status information, establishes a joint optimization problem for each client class, and finally solves the joint optimization problem to obtain the optimal global aggregation interval and the optimal gradient compression rate for each client class. Step A2: Collaborative Training Phase: The server asynchronously triggers global model aggregation according to the optimal global aggregation interval; after receiving the global model from the server, each client performs local training using its local private data, compresses the calculated gradients according to the optimal gradient compression rate corresponding to its category, and uploads the compressed gradients to the server; when triggering aggregation, the server assigns weights to each gradient according to the staleness of the gradients uploaded by each client, and performs weighted aggregation of all received valid gradients based on the weights to update the global model.
[0044] It should be noted that this invention adopts a centralized federated learning architecture, consisting of a central server and multiple distributed clients. The server is responsible for policy computation, model management, and aggregation updates, and does not store the original data from the clients; the clients only perform local training and gradient compression upload, which does not involve privacy data and has low requirements for device performance.
[0045] In this invention, device status includes the pure computation time for the client to complete one local training iteration and the communication time for uploading the complete gradient. During the pre-training phase, each client automatically performs a mini-benchmark test to obtain the aforementioned metrics and reports them to the server. The data collection frequency is a one-time collection during the pre-training phase; during training, retesting is only performed as needed in cases of anomalies or long periods to reduce overhead.
[0046] During the pre-training phase, the client automatically performs micro-benchmark tests: first, it performs local simulation training with a fixed number of steps using a randomly initialized model, measuring the local computation time; then, it constructs a virtual gradient of the same dimension and uploads it to the server, measuring the communication time. The test results are automatically packaged and reported. The process is lightweight, requires no manual intervention, and does not affect normal business operations.
[0047] A strategy of one-time data collection during the pre-training phase and lightweight updates on demand during training is adopted: all clients complete state collection before training for clustering and optimization calculations; during training, data collection is not repeated by default to reduce communication overhead, and retesting and state updates are only triggered when clients experience consecutive timeouts, network fluctuations, or reach a preset period. This design minimizes additional overhead while ensuring state accuracy.
[0048] This invention centralizes complex optimization calculations on the server side, while the client only performs lightweight local training and uploading. The server and client interact asynchronously, aggregating data at fixed intervals to avoid slowing down the overall process with individual slow devices, thus achieving friendly compatibility with heterogeneous devices. Simultaneously, the trade-off between training efficiency and communication overhead is transformed into a joint optimization problem of aggregation interval and gradient compression rate. These two factors are mutually restrictive and must be planned in a unified manner to achieve overall optimality.
[0049] This invention aims to maximize the total amount of effective gradient information with low staleness. Mathematically, this is defined as the sum of the products of the staleness weights and gradient compression rates of all participating clients, multiplied by the total gradient dimension. To quantify gradient freshness, a staleness weight function is designed: it adopts a square root decay form inversely proportional to staleness, with lower weights for higher upload latency; if the total time exceeds the allowable window, the weights are reset to zero, and the gradient does not participate in aggregation. This design both suppresses the negative impact of stale gradients on the model and ensures computational differentiability.
[0050] In this invention, "timeout" is dynamically determined: a timeout occurs when the total training and communication time of the client exceeds the sum of the aggregation intervals of its last successful upload to the current round. This rule adapts to changes in the aggregation rhythm. The aggregation interval is constrained within a reasonable range (e.g., 2-20 seconds), and the gradient compression rate is between 0 and 1 (typically 0.05-0.5). The system also incorporates constraints such as global training time, communication traffic budget, and model convergence, forming a complete and solvable optimization problem.
[0051] This invention transforms the fuzzy goal of "improving efficiency and reducing overhead" into a precisely solvable optimization problem by incorporating multiple constraints such as convergence accuracy, training time, communication traffic, and variable range. The model assigns weights based on the timing of gradient uploads: the more timely the upload, the higher the weight; gradients that time out are directly filtered. Simultaneously, it strictly controls the communication traffic in each round, maximizing effective information transmission within the budget, avoiding resource waste and network congestion, and ensuring stable system operation in constrained scenarios.
[0052] It should be noted that this invention addresses the joint optimization problem of high-dimensional, non-convex, and non-differentiable problems, and proposes three core strategies: The first key strategy is the smooth approximation strategy, which transforms the rigid "yes / no" judgments during optimization into a continuously computable smooth function, ensuring stable, uninterrupted, and stable algorithm operation. In the original problem, it's necessary to determine whether the client successfully uploaded gradients and participated in the current aggregation round. These judgments are implemented using a sign function, which involves discontinuous and non-differentiable computations, causing the optimization algorithm to fail. This invention employs the hyperbolic tangent function tanh(γ) x), as a continuous approximation of the symbolic function, is a recognized smooth approximation tool in mathematics, capable of making the entire calculation process continuously differentiable while maintaining the judgment logic unchanged. The key parameter γ is the steepness coefficient; this invention explicitly provides an engineering value of γ = 10^6. This value ensures sufficiently high approximation accuracy, making the judgment result close to the original logic, while avoiding computational overflow due to excessive value; it is the optimal stable value verified through extensive experiments. Through this smoothing process, the originally unsolvable non-smooth constraints are transformed into a standard computable form, laying the foundation for subsequent efficient optimization.
[0053] The second key strategy is a multi-timescale decomposition strategy. The core of this strategy is to separate the aggregation interval and gradient compression rate, optimizing them alternately at different frequencies. This avoids system chaos, model oscillation, and slow convergence caused by adjusting too many variables simultaneously. This invention provides clear engineering definitions for "slower frequency" and "faster frequency": the outer slow-scale optimization updates the global aggregation interval, with an update frequency set to once every 10 rounds of global aggregation. This frequency ensures that the aggregation rhythm does not change frequently, keeping the overall training process stable and preventing model fluctuations caused by frequent adjustments to the aggregation interval. The inner fast-scale optimization updates the client-side gradient compression rate, with an update frequency set to once every 1 round of global aggregation. This allows for rapid response to real-time conditions such as client network fluctuations, changes in computing power, and upload latency, ensuring that the compression strategy always matches the current device conditions. This layered design of "slow updates for stability, fast adjustments to adapt to differences" reduces computational complexity and achieves an optimal balance between stability and adaptability, demonstrating strong robustness in actual deployments.
[0054] The third key strategy is a compression ratio clustering strategy, designed to address the problem of excessive computation and inability to solve problems in real time due to a large number of clients. This invention provides a clear and directly usable method for setting clustering parameters, ensuring that anyone can complete the configuration step by step. Clustering adopts the standard K-means algorithm, and the clustering features are fixed using two device status indicators of the client: the pure computation time of a single local training session and the communication time for uploading the complete gradient. These two indicators can fully reflect the heterogeneity of computing power and network. The number of clusters K adopts an engineering-based setting rule, with the formula K=min (5,N / 20), where N is the total number of clients. When the number of clients is less than 100, K is 5, and it increases proportionally when the number of clients is greater. This setting maintains an optimal balance between optimization accuracy and computational efficiency. Clustering is executed only once during the pre-training phase, and no re-clustering is performed during training to avoid additional overhead. In this way, clients within the same cluster share the same compression ratio, significantly reducing the number of variables that need to be optimized. This allows the server to complete all optimal strategy calculations in a short time, even when facing hundreds of heterogeneous clients, enabling efficient deployment in large-scale scenarios.
[0055] It should be noted that this invention employs the Trust Region Newton's Method (TRON) to solve optimization problems. This algorithm is designed for high-dimensional, non-convex, constrained nonlinear programming. By limiting the trust region at each step, constructing a second-order approximation model, and dynamically adjusting the search radius, it ensures stable convergence. The algorithm provides explicit convergence criteria, radius update rules, subproblem-solving methods, and iteration stopping conditions, and can be fully reproduced, making it suitable for the joint optimization scenarios described in this invention.
[0056] I. Convergence Criterion (A clear standard for determining when the result is stable and no longer changes) This invention employs a triple convergence criterion. If any one of these criteria is met, the result is deemed "stable and no longer changing," the iteration stops, and the current optimal solution is output. All three criteria are quantifiable standards that can be directly used in engineering. 1. Parameter variation convergence: The maximum absolute difference between the aggregation interval and the compression ratio obtained from two adjacent iterations is less than 1e-5, indicating that the parameters have basically stopped changing and have reached a stable state.
[0057] 2. Objective function convergence: The relative change in the objective function value (total amount of effective gradient information with low staleness) between two adjacent iterations is less than 1e-6, indicating that the optimization objective can no longer be improved.
[0058] 3. Maximum iteration count convergence: When the number of iterations reaches the preset upper limit of 200 rounds, the iteration is forcibly stopped to avoid infinite loops and ensure that training can be completed within a limited time.
[0059] The above convergence conditions balance accuracy and efficiency, ensuring the quality of the solution in actual deployment without wasting computing resources due to excessive iteration.
[0060] II. Key parameters of the trust region (initial values, update rules) To ensure the algorithm can be directly implemented, this invention provides fixed values and dynamic update rules for all key parameters, eliminating the need for manual parameter tuning: 1. Initial value of trust region radius: Set to Δ0 = 1.0. This value is a general safety initial value and is applicable to most federated learning optimization scenarios.
[0061] 2. Upper limit of trust region radius: set to Δ_max = 10.0 to prevent the search range from being too large and causing unstable iteration.
[0062] 3. Trust region radius lower limit: set to Δ_min = 1e-4 to ensure that the search range does not become too small and stagnate.
[0063] 4. Radius update rules: • If the ratio of the actual target descent to the predicted descent is >0.75: this indicates that the current step is effective; increase the radius, Δ = min (2 Δ, Δ_max).
[0064] • If the ratio is <0.25: This indicates poor performance in the current step. Reduce the radius, Δ = max(0.25) Δ, Δ_min).
[0065] • If the ratio is between 0.25 and 0.75: keep the trust region radius unchanged.
[0066] III. Subproblem Solving Methods The trust region Newton method requires solving a trust region subproblem in each iteration, i.e., finding the search direction that maximizes the decrease in the objective function within a defined radius. This invention employs the conjugate gradient (CG) method to solve the subproblem. This method has low computational cost, low memory usage, and is suitable for high-dimensional variable scenarios, making it a standard choice for large-scale optimization problems. The maximum number of iterations for the subproblem is set to 30 steps to ensure rapid completion of a single solution.
[0067] It should be noted that, in the actual solution process, the server will run the trust region Newton method in an alternating iterative manner across multiple time scales: first, fix the compression ratio and perform TRON iteration on the aggregation interval; then fix the aggregation interval and perform TRON iteration on the compression ratio, alternating until overall convergence.
[0068] This invention enables the trust region Newton method to converge rapidly within 20 to 50 iterations through smooth approximation, cluster dimensionality reduction, and multi-timescale decoupling, without the need for manual parameter tuning, and strictly adheres to the dual constraints of training time and communication traffic.
[0069] The overall process is divided into two stages: During the pre-training optimization phase, the system first completes all policy calculations to prepare for subsequent training. First, each client only needs to perform a very brief local test to measure its local computation speed and upload communication speed, and then report these two privacy-neutral performance data points to the server. Specifically, upon receiving the server's test command, each client automatically loads a randomly initialized model with the same structure and parameter dimensions as the actual training model. It does not load any real local business data, but only uses virtual samples to complete five rounds of fixed-step local forward computation and backward gradient derivation, recording the pure time consumed from start to finish; this value is the client's local training time. After completing the computation test, the client generates virtual gradient data with dimensions completely identical to the real gradient, containing random values and no privacy information. This virtual gradient is uploaded to the server, and the total time consumed from sending to receiving the acknowledgment is recorded; this value is the client's communication time. After the test, the client immediately deletes the virtual data and temporary model, reporting only the local training time and communication time to the server. The entire test process typically takes less than one second, with extremely low computational and communication loads, having virtually no impact on the device. Upon receiving performance information from all clients, the server first divides the clients into several groups according to a clustering strategy, then initiates the aforementioned joint optimization calculation to quickly determine the optimal aggregation interval for each round and the optimal compression ratio for each group of clients. After completing the calculation, the server distributes the corresponding compression strategy to each group of clients and simultaneously distributes the initial global model to all devices. This stage is extremely short and involves minimal communication, placing no burden on the system.
[0070] During the collaborative training phase, the system enters the formal asynchronous training process. The server starts a global clock, triggering model updates at pre-calculated aggregation intervals and continuously monitoring compressed gradient information uploaded by clients. After receiving the model from the server, the client completes a specified number of training iterations locally using private data, calculates gradient information, and simplifies it according to a specified compression ratio before quickly uploading the compressed, small amount of gradient data. When the aggregation time point is reached, the server performs staleness assessment and weighted aggregation on the collected gradients, discards invalid gradients that have timed out, updates the global model with valid information, and distributes the new model to the clients participating in this aggregation round to continue the next round of training. When the total training time reaches the preset upper limit, the server automatically sends a stop command to all clients, the training process ends, and a high-precision global model that has converged is finally output.
[0071] The two-stage system features a high degree of automation and minimal human intervention. The client performs only simple operations while the server makes complex decisions, balancing efficiency, accuracy, and communication overhead. It is suitable for scenarios such as smart cities, the Internet of Things, and the Industrial Internet.
[0072] Example 3 Please refer to Figure 1 This embodiment provides an adaptive gradient compression and asynchronous aggregation federated learning method, including: Step A1: Pre-training optimization stage: The central server collects device status information from each client, clusters the clients based on the device status information, establishes a joint optimization problem for each client class, and finally solves the joint optimization problem to obtain the optimal global aggregation interval and the optimal gradient compression rate for each client class. Step A2: Collaborative Training Phase: The server asynchronously triggers global model aggregation according to the optimal global aggregation interval; after receiving the global model from the server, each client performs local training using its local private data, compresses the calculated gradients according to the optimal gradient compression rate corresponding to its category, and uploads the compressed gradients to the server; when triggering aggregation, the server assigns weights to each gradient according to the staleness of the gradients uploaded by each client, and performs weighted aggregation of all received valid gradients based on the weights to update the global model.
[0073] This embodiment is applied to a collaborative intelligent training scenario for heterogeneous terminals in urban IoT. The scenario includes a large number of edge acquisition devices of different brands, configurations, and network conditions. These devices are geographically dispersed, have limited uplink bandwidth, and vary significantly in computing power. Furthermore, they cannot centrally upload raw data, necessitating the use of federated learning to train the model. This scenario places strict requirements on training time, communication traffic consumption, and the final accuracy of the model, making it a typical environment for verifying the technical effectiveness of this invention.
[0074] The hardware environment used in this embodiment consists of a high-performance central server and 100 heterogeneous client terminals. The central server has a multi-core processor and large memory, and is used to handle core tasks such as global model maintenance, joint optimization calculation, and asynchronous gradient aggregation. The clients include high-performance edge gateways, mid-range acquisition devices, and low-performance sensing terminals. The computing power, uplink bandwidth, and storage capacity of the devices vary significantly, fully simulating the heterogeneous characteristics of a real-world scenario. The software environment is built on the Python programming language and the PyTorch deep learning framework, and is equipped with a lightweight distributed communication module. No additional dedicated plugins are required, and it supports cross-platform and cross-device operation.
[0075] This embodiment uses the MNIST handwritten digit dataset as training and testing data. This dataset is the most commonly used and representative standard image classification dataset in the field of federated learning. It contains handwritten digit images of 10 categories, with uniform image size and moderate data scale, making it suitable for training on edge devices. To align with real-world federated learning scenarios, this embodiment employs a label-biased non-independent identically distributed (Non-IID) partitioning method, distributing the dataset to different clients based on label categories: the 10 digit labels are divided into 5 groups, each containing 2 digit categories. Each client is assigned samples corresponding to only one set of labels, resulting in significant differences in data distribution between clients and no category overlap. This aligns with the characteristics of real-world scenarios where data collected from different terminals is of a single category and unevenly distributed. This partitioning method maximizes the impact of heterogeneous environments on training and more realistically verifies the adaptability of this invention.
[0076] This embodiment uses a lightweight convolutional neural network (CNN) as the training model. Its simple structure, low computational cost, and few parameters make it perfectly suited to the computing power limitations of edge devices. The specific structure and parameters are as follows: The model consists of 2 convolutional layers, 2 pooling layers, 1 flattening layer, and 2 fully connected layers. The input is a 28×28 grayscale image. The first convolutional layer uses 16 3×3 convolutional kernels with ReLU activation. The second convolutional layer uses 32 3×3 convolutional kernels with ReLU activation. Both pooling layers use 2×2 max pooling. The first fully connected layer has an output dimension of 128, and the second layer has an output dimension of 10, corresponding to 10 digit categories. The total number of parameters is approximately 167,000. With a small parameter count and fast computation speed, it can be smoothly trained locally on low-power edge devices while possessing sufficient feature extraction capabilities to ensure classification accuracy.
[0077] This embodiment sets explicit resource constraints to be consistent with the real engineering environment, including the global maximum training time, total communication traffic budget, aggregation interval range, gradient compression rate range, etc. All optimization processes are executed within the constraints to ensure that the technical solution of this invention can be directly deployed, rather than just remaining at the theoretical verification level.
[0078] It should be noted that before officially starting training, system environment configuration, basic parameter settings, and initial information collection need to be completed to provide a stable foundation for subsequent optimization and training. First, the federated learning scheduling module, joint optimization solution module, asynchronous aggregation module, and communication management module are configured on the central server. On the client side, a lightweight local training program, gradient calculation module, Top-K gradient sparsity compression module, and status reporting module are deployed. The client modules are small in size and lightweight in operation, and will not interfere with normal device operations.
[0079] This embodiment standardizes and reproducibly sets explicit numerical values for all core parameters, eliminating vague descriptions and ensuring highly consistent experimental results across multiple runs in the same environment. Specifically, the number of local iterations on the client is set to 5. This value ensures sufficient local gradient updates on the client while avoiding excessive computation time and over-updating of gradients, which could lead to model oscillations. It is an optimal and moderate value verified through extensive experiments in edge-heterogeneous scenarios. The client-side local training learning rate is set to 0.01, and the server-side global aggregation learning rate is set to 0.1. These two sets of values are stable combinations widely used in federated learning, allowing the model to converge smoothly during asynchronous training without drastic fluctuations in loss, non-convergence, or excessively slow convergence. The steepness coefficient γ in the smooth approximation strategy is explicitly set to 10^6. This value ensures sufficiently high approximation accuracy of the sign function, making the judgment logic close to the original discrete form, while avoiding numerical overflow due to excessive numerical values. It is a relatively large standard value that balances stability and accuracy.
[0080] This embodiment uses Top-K sparsity as the gradient compression method. There is a clear engineering conversion formula between the compression ratio δ and the number of parameters K retained: K = round(δ × d), where d is the total dimension of the model gradient. The gradient dimension d of the lightweight CNN model used in this embodiment is 167210, and round() represents rounding to the nearest integer. For example, when the client compression ratio δ = 0.1, K = round(0.1 × 167210) = 16721, meaning that only the 16721 parameters with the largest absolute values are retained in the gradient, and all other parameters are set to 0 before uploading. This conversion method is simple, intuitive, and computationally efficient. The client can quickly calculate the K value and perform gradient compression based on the compression ratio sent by the server, without any additional complex operations.
[0081] Meanwhile, the server pre-sets resource constraint thresholds consistent with real-world edge scenarios. The global total training time is capped at 3600 seconds (60 minutes), the global communication traffic budget B is set to 1000MB, the aggregation interval is limited to [2 seconds, 20 seconds], and the gradient compression rate δ is set to a reasonable engineering range of (0, 1). All parameters conform to the actual resource limitations of IoT terminals, edge gateways, and other devices, ensuring that the optimized strategy configuration is fully usable and will not exceed hardware and network capacity. The client automatically loads the above fixed parameter configuration upon startup, requiring no manual input or parameter tuning throughout the entire process. This achieves fully automated operation from startup to training completion, significantly reducing deployment difficulty and operational costs.
[0082] It should be noted that this embodiment strictly follows the two-stage execution process of the present invention, first completing pre-training optimization, and then entering asynchronous collaborative training. No manual intervention is required throughout the process, and the operation is stable and the logic is clear.
[0083] During the pre-training optimization phase, all clients first perform a standardized, lightweight benchmark test without privacy data to accurately obtain device performance metrics. It is clearly defined here that the pure computation time reported by the client refers to the total computation time required to complete one full local training cycle (i.e., 5 fixed local iterations), not the time of a single iteration. This metric more accurately reflects the device's computing power to complete a full training cycle. During the test, the client does not use any real private data; it only loads a virtual model with the same structure and dimensions as the actual training model, executes 5 rounds of standard local stochastic gradient descent iterations, and records the pure computation time from the start of the iteration to its completion, obtaining the local computation time. Subsequently, it generates a virtual gradient with the same dimensions as the real gradient, uploads it to the server, and records the complete transmission time, obtaining the communication time. The client only reports these two performance metrics to the server. The entire testing process is lightweight, fast, and secure, without leaking any privacy data or affecting the device's normal operations.
[0084] After receiving performance data from all clients, the server executes the K-means clustering algorithm, using local computation time and communication time as two-dimensional features, to divide the 100 clients into 5 clusters with similar performance. This embodiment explicitly specifies the key implementation details of K-means: the initial cluster centers are selected by randomly sampling from the client features, i.e., 5 are randomly selected from the feature vectors of the 100 clients as the initial cluster centers, ensuring a uniform and unbiased initial distribution; the iteration stopping condition for clustering uses a dual standard: first, the maximum offset of the cluster center positions between two adjacent iterations is less than 1e-6, indicating that the cluster structure is stable; second, the iteration count reaches the upper limit of 100, forcibly stopping to avoid infinite loops. Clustering is completed when either condition is met. Through this clustering process, clients with similar performance are grouped into the same cluster, and clients within the same cluster share the same compression ratio strategy, significantly reducing the variable dimensionality of subsequent optimizations.
[0085] After clustering is completed, the server initiates the trust-region Newton method solution process based on the joint optimization model constructed in this invention. First, a smooth approximation of the sign function is made using the hyperbolic tangent function. Then, it iterates alternately across multiple time scales: the outer layer updates the global aggregation interval every 10 rounds to ensure training stability; the inner layer updates the compression ratio of each client cluster every round to quickly adapt to device conditions. The iteration process converges and stops when the parameter change is less than 1e-5, the relative change of the objective function is less than 1e-6, or the maximum number of iterations is reached (200). Finally, the optimal interval for each round of global aggregation and the optimal gradient compression ratio for each client cluster are output. The server then distributes the corresponding compression ratio configuration to the clients and simultaneously distributes the initial global model, completing the pre-training phase.
[0086] During the collaborative training phase, the server starts the global clock and initializes the aggregation timer, entering the asynchronous listening and aggregation process. After receiving the global model, the client performs 5 local iterations starting from that model, calculates the local total gradient, and then calculates the Top-K retention quantity according to the compression ratio δ using the formula K=round (δ×d). The gradient is then sparsified and compressed, retaining only the K parameters with the largest absolute values, and then uploaded as compressed gradient. The server continuously collects gradient information. When the difference between the global clock and the aggregation time of the previous round reaches the optimal aggregation interval for the current round, global aggregation is triggered: the server first calculates the gradient staleness (i.e., the difference between the current round and the round most recently participated in aggregation). The validity criterion is that the total training communication time of the client does not exceed the sum of the aggregation intervals from its last successful upload to the current round. Gradients that time out are reset to 0 and do not participate in aggregation. Valid gradients are assigned weights according to the weight calculation formula in Section 4.2.2. After the server weights and aggregates the valid gradients, it updates the global model and distributes the new model to the clients participating in this round of aggregation, resetting the aggregation timer to enter the next round. When the global clock reaches the preset total training time limit of 1800 seconds, the server broadcasts a stop command to all clients, and the clients terminate training and uploading, thus officially ending the entire federated learning process.
[0087] It should be noted that this embodiment verifies the technical effect of the present invention through multiple sets of repeated experiments, and makes a comprehensive comparison with existing mainstream asynchronous federated learning methods, asynchronous federated learning combined with fixed gradient compression methods, single adaptive aggregation interval methods, and single adaptive gradient compression methods. All results are the average of multiple experiments to ensure objectivity and persuasiveness.
[0088] In terms of training speed, existing technologies are affected by problems such as synchronous waiting, individual optimization, and heterogeneous mismatch, resulting in a long total training time required to achieve the same convergence goal. However, this invention significantly reduces invalid waiting and communication blockage by jointly optimizing the aggregation interval and gradient compression rate, thus significantly shortening the training completion time by 30% to 50% compared to the existing best method. It has obvious advantages in time-sensitive edge scenarios.
[0089] In terms of communication overhead, existing fixed compression methods consume a lot of traffic, and standalone adaptive compression methods do not combine with aggregation rhythm, resulting in low utilization. However, this invention allocates personalized compression rates according to the heterogeneous characteristics of the client, minimizing the amount of transmitted data while ensuring effective gradient information. The overall communication traffic consumption is reduced by 40% to 60% compared with existing methods, significantly saving uplink bandwidth resources. It is particularly suitable for scenarios with poor network conditions and high traffic costs.
[0090] Regarding model accuracy, existing technologies suffer from slow model convergence and low final accuracy due to the large number of old gradients involved in aggregation. In contrast, this invention strengthens the contribution of fresh gradients and weakens the influence of old gradients through a weighted aggregation mechanism based on staleness, resulting in more stable global model convergence and higher final test accuracy. This is 2% to 5% higher than existing methods, and a higher quality model is obtained with the same amount of data and number of iterations.
[0091] Regarding system scalability, existing methods suffer from problems such as computational explosion and slow optimization when the number of clients increases. However, this invention significantly reduces the dimensionality of optimization variables through a client clustering strategy. Even when the number of clients expands to hundreds or thousands, the server can still quickly complete the optimization calculations, and the system runs stably without lag, meeting the deployment needs of large-scale IoT and industrial internet scenarios.
[0092] Based on a comprehensive evaluation of various indicators, this invention outperforms existing technologies in terms of training speed, communication overhead, model accuracy, device compatibility, and scalability. It can achieve efficient, stable, and high-precision asynchronous federated learning training under strict resource constraints, fully achieving the invention's objectives and possessing extremely high practical application value.
[0093] Example 4 This embodiment provides a computer device, including a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor. When the computer program is executed by the processor, it implements the steps of an adaptive gradient compression and asynchronous aggregation federated learning method as described in Embodiment 1, Embodiment 2, or Embodiment 3.
Claims
1. An adaptive gradient compression and asynchronous aggregation federated learning method, characterized in that, include: Pre-training optimization phase: The central server collects device status information from each client, clusters the clients based on the device status information, establishes a joint optimization problem for each client class, and finally solves the joint optimization problem to obtain the optimal global aggregation interval and the optimal gradient compression rate for each client class. Collaborative training phase: The server asynchronously triggers global model aggregation according to the optimal global aggregation interval; After receiving the global model from the server, each client performs local training using its local private data, compresses the calculated gradients according to the optimal gradient compression rate corresponding to its category, and uploads the compressed gradients to the server. When the server triggers aggregation, it assigns weights to each gradient according to the staleness of the gradients uploaded by each client, and performs weighted aggregation of all received valid gradients based on the weights to update the global model.
2. The adaptive gradient compression and asynchronous aggregation federated learning method according to claim 1, characterized in that, The pre-training optimization phase specifically includes: S1: Each client performs a benchmark test, obtains the pure computation time required to complete one local training session and the communication time required to upload one complete gradient, and reports the device status information containing the pure computation time and communication time to the central server. S2: Based on the device status information reported by all clients, the server performs a clustering operation on all clients using the pure computing time and communication time as clustering features, and groups clients with similar performance into the same category; S3: The server aims to maximize the total amount of effective gradient information with low staleness, and sets at least a global training time limit and a total communication traffic budget as constraints to establish a joint optimization problem for each type of client. S4: The server uses the trust region Newton method to solve the joint optimization problem, and obtains the optimal global aggregation interval and the optimal gradient compression rate for each type of client. S5: The server configures the optimal global aggregation interval in its own aggregation module, and sends the optimal gradient compression rate and its initial global model corresponding to each type of client to that type of client.
3. The adaptive gradient compression and asynchronous aggregation federated learning method according to claim 1, characterized in that, The collaborative training phase specifically includes: Step T1: The server starts the global clock and periodically triggers global aggregation events according to the optimal global aggregation interval; Step T2: After receiving the global model from the server, each client uses its local private data to perform a specified number of rounds of local training and calculates the original gradient. Step T3: Each client performs Top-K sparsity compression on the original gradient according to the received optimal gradient compression rate, retaining only the K gradient parameters with the largest absolute values, and uploads the compressed gradient to the server; Step T4: Each time global aggregation is triggered, the server calculates the staleness of the gradient based on the difference between the gradient uploaded by each client and the round in which it successfully participated in the aggregation last time, and assigns a weight inversely proportional to the staleness to each gradient. Step T5: The server determines whether the total time taken for the client from the most recent successful upload to the current time exceeds the sum of the preset aggregation intervals. If so, the weight of the gradient is reset to zero, and it is determined to be an invalid gradient. Step T6: The server performs a weighted summation of the valid gradients with non-zero weights to complete the aggregation and update of the global model, and then sends the updated global model to the clients participating in this aggregation. Step T7: Repeat steps T2 to T6 until the global training time limit is reached or the model converges.
4. The adaptive gradient compression and asynchronous aggregation federated learning method according to claim 2, characterized in that, The expression for the total amount of low-obsolescence effective gradient information is: in, The total amount of effective gradient information with low obsolescence The aggregate weight of the i-th client in round t. For gradient compression ratio, This represents the original dimension of the gradient.
5. The adaptive gradient compression and asynchronous aggregation federated learning method according to claim 4, characterized in that, The server assigns differentiated aggregation weights to different clients based on the staleness of the gradient; the lower the staleness of the gradient, the larger the aggregation weight. If gradient transmission times out, the aggregation weight is reset to zero, and the gradient does not participate in global aggregation. in, From the last successful transmission by the client to the [number]th round t Total aggregation time of the wheel, For the first i The client in the first t Before the global aggregation round, the last global round in which gradient transfer was successfully completed, initialized to a value of ,Right now: .
6. The adaptive gradient compression and asynchronous aggregation federated learning method according to claim 1, characterized in that, The joint optimization problem aims to maximize the total amount of effective gradient information with low staleness, and includes at least global training time constraints, communication traffic resource constraints, convergence constraints, and variable range constraints. The global training time constraint is defined as follows: the sum of the aggregation intervals of all rounds does not exceed the preset total global training time. T ,Right now: The communication traffic resource constraint is defined as follows: Let b be the traffic consumption for transmitting a complete gradient. Then, the total traffic consumption of effective gradients across all rounds shall not exceed the preset traffic resource budget B, i.e.: in, This represents the bandwidth consumption for transmitting the compressed gradient. For symbolic functions, express , No. t The gradient transfer is completed in one round; express , No. t Gradient transfer was not completed in the first round; The convergence constraint is defined as follows: the training loss of the global model needs to converge to a preset threshold, let the convergence threshold be... ( ),but: in, Let be the loss value of the global model in round t. This represents the loss value of the optimal model. The variable range constraint is defined as follows: the aggregation interval and compression ratio are positive real numbers, the aggregation interval is within a preset range, and the compression ratio is between 0 and 1, that is: in, and These are the minimum and maximum values of the aggregation interval, respectively; The complete formal representation of the joint optimization problem is as follows: 。 7. The adaptive gradient compression and asynchronous aggregation federated learning method according to claim 1, characterized in that, The method for solving the joint optimization problem includes: A smooth approximation strategy is adopted, replacing the discontinuous and non-differentiable judgment logic in the joint optimization problem with a continuously differentiable hyperbolic tangent function. The expression of the hyperbolic tangent function is: tanh (γ x); where γ is the steepness coefficient; A multi-timescale decomposition strategy is adopted to update the global aggregation interval at a first preset frequency and the gradient compression rate at a second preset frequency, wherein the first preset frequency is lower than the second preset frequency. The trust region Newton method is used to iteratively solve the optimization problem after smooth approximation and multi-time scale decomposition until the preset convergence condition is met.
8. The adaptive gradient compression and asynchronous aggregation federated learning method according to claim 7, characterized in that, The trust region radius Δ in the trust region Newton method is dynamically updated according to the following rules: Let the ratio be ,but: like Then the radius is increased as follows: Δ = min (2 Δ, Δ_max); like Then, the radius is reduced as follows: Δ = max(0.25) Δ, Δ_min); like If so, then the current radius remains unchanged; Where Δ_max = 10.0 is the upper limit of the radius, Δ_min = 1e-4 is the lower limit of the radius, and the initial radius Δ0 = 1.
0.
9. The adaptive gradient compression and asynchronous aggregation federated learning method according to claim 7, characterized in that, The preset convergence condition includes any one of the following: Parameter variation convergence: The maximum absolute difference between the aggregation interval and the compression ratio obtained from two adjacent iterations is less than 1e-5; Objective function convergence: The relative change in the total amount of low-stale effective gradient information between two adjacent iterations is less than 1e-6; Maximum iteration count convergence: When the number of iterations reaches the preset upper limit of 200 rounds, the iteration is forcibly stopped.
10. A computer device, characterized in that: It includes a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor, wherein the computer program, when executed by the processor, implements the steps of an adaptive gradient compression and asynchronous aggregation federated learning method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Safe aggregation method based on asynchronous federated learning with error learning
CN117494794A